A data quality improvement system and method driven by business cognition in a data middle platform
By building a business knowledge base and local neighborhood analysis, identifying and managing the logical differences between new and old businesses, the data consistency and traceability problems of the data middle platform in complex business scenarios are solved, and dynamic adaptation and accurate evaluation of data quality are achieved.
Patent Information
- Application Number
- CN202510535082.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-27
AI Technical Summary
The existing data middle platform is difficult to adapt to business logic changes in complex business scenarios, which makes it difficult to ensure data traceability ambiguity and consistency, affecting the reliability of data quality evaluation and decision-making.
Build a unified business knowledge base, integrate cognitive rules in multiple business fields, identify the logic differences between new and old business through differential analysis, carry out version management of data fields and calculation methods, and use local neighborhood analysis to calculate structural entropy to ensure data consistency.
It realizes accurate mapping configuration and versioning management of historical data, ensures that the data has accurate traceability and consistency in the process of business logic changes, dynamically adapts to business rules changes, and improves the accuracy and reliability of data quality evaluation.
Smart Images

Figure CN120067091B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data quality management. More specifically, the present invention relates to a data quality improvement system and method driven by business cognition in a data middle platform. Background Art
[0002] When the existing data middle platform supports data quality improvement driven by business cognition, it usually relies on predefined business rules and data quality standards. However, in complex business scenarios, the business logic will evolve continuously with market demands, policy adjustments, or enterprise strategic upgrades, resulting in changes in data processing rules, field definitions, calculation methods, etc.
[0003] When the generation, storage, and processing logics of data are adjusted, historical data may not be directly adaptable to the new rules, and thus data traceability ambiguity may occur; for example, a certain business indicator may be generated based on different calculation logics at different time points, resulting in inconsistent meanings of the same indicator in different historical versions, making it difficult for data quality verification to accurately adapt to business requirements. The prior art is difficult to effectively identify, record, and manage such business logic changes, which may cause difficulties in data traceability, distortion of quality assessment, and even affect the reliability of data-driven decision-making. Summary of the Invention
[0004] In order to overcome the above-mentioned defects of the prior art, embodiments of the present invention provide a data quality improvement system and method driven by business cognition in a data middle platform to solve the problems raised in the above background art.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A data quality improvement method driven by business cognition in a data middle platform, comprising the following steps:
[0007] S1: Obtain business cognition rules in multiple business domains and integrate the business cognition rules into a unified business knowledge base;
[0008] S2: Perform differential analysis on the newly generated business logic based on the business knowledge base to identify the specific differences between the new business logic and the existing business logic in data fields, data processing processes, and data verification rules;
[0009] S3: Perform mapping configuration on the business indicator data in the historical data according to the differential analysis results, and implement version management for data field definitions and calculation methods;
[0010] S4: Compare the mapped and configured business indicator data in the historical data with the business indicator data set generated according to the new business logic based on local neighborhood analysis, calculate the structural entropy within the local data distribution area divided by neighborhoods, and generate a data consistency evaluation result according to the calculation result;
[0011] S5: Write the data with normal data consistency assessment results and meeting the business cognition rules into the data middle platform, and associate and save the data traceability information with the business logic version information.
[0012] In a preferred embodiment, S1 includes:
[0013] Extract the original business information related to business processes, data metric definitions, data operation rules, and data verification criteria from each business system within the enterprise respectively;
[0014] Preprocess the original business information to form structured business cognition rules. The structured business cognition rules include rule identification, rule name, rule description, applicable business fields, and corresponding rule parameters;
[0015] Classify, associate, and merge the structured business cognition rules, and uniformly store the processed business cognition rules in a business knowledge base with a predefined data table structure and index design.
[0016] In a preferred embodiment, S2 includes:
[0017] Obtain the structured business cognition rules reflecting the existing business logic, including data fields, data processing flows, and data verification rules, from the business knowledge base;
[0018] Extract the business rule information in the newly generated business logic according to the same data fields, data processing flows, and data verification rules as the existing business logic to form a new business logic rule set;
[0019] Compare each item of the new business logic rule set with the business cognition rules corresponding to the existing business logic stored in the aforementioned business knowledge base:
[0020] Analyze the specific differences in data type, data format, data length, and data meaning of each data field; compare each link of the data processing flow item by item to identify the specific differences in the data input, data conversion, and data output processes; compare each item of the data verification rules to identify the specific differences in each verification standard, verification condition, and rule parameters;
[0021] Record the specific differences in a predefined format to form the difference analysis result between the new business logic and the existing business logic.
[0022] In a preferred embodiment, S3 includes:
[0023] Perform mapping configuration on the business metric data in the historical data;
[0024] Implement version management for the data field definitions included in the data related to business indicators in historical data. The data field definitions include the name, data type, data format, data length, and data meaning of the data fields. Version management records, identifies, and subsequently updates each data field definition according to predefined version control rules.
[0025] Implement version management for the calculation methods of business indicator data in historical data. The calculation methods include the formulas, calculation parameters, and data processing rules for business indicator calculations. Version management records and manages different versions of calculation methods based on predefined version identifiers and version update rules.
[0026] In a preferred embodiment, the mapping configuration includes:
[0027] a. Compare item by item the data reflecting business indicators in historical data with the business rule information involved in the new business logic according to predefined mapping criteria.
[0028] b. Determine the corresponding relationships between each business indicator data in historical data and the corresponding business rules in the new business logic.
[0029] c. Store the corresponding relationships in the form of mapping records.
[0030] In a preferred embodiment, S4 includes:
[0031] Obtain the historical business indicator data after mapping configuration and the business indicator data set generated according to the new business logic, and perform preliminary merging processing on the two.
[0032] Perform neighborhood partitioning on the merged data according to the local neighborhood analysis method. Neighborhood partitioning groups different records according to the data distribution characteristics.
[0033] Within the local data distribution area corresponding to each neighborhood, calculate the entropy value of the data based on a pre-set structural entropy calculation method. The structural entropy calculation method measures the distribution characteristics of each record in the same neighborhood in multiple field dimensions.
[0034] Compare the entropy value differences between the historical business indicator data and the new business logic business indicator data in each neighborhood, and use the entropy value difference results as the basis for local consistency determination.
[0035] Summarize the data consistency evaluation results based on the entropy value difference results of each neighborhood, and store the data consistency evaluation results in association with the previous mapping configuration information.
[0036] In a preferred embodiment, calculate the entropy value of the data based on a pre-set structural entropy calculation method, specifically:
[0037] Within each local neighborhood, the historical business metric data after mapping configuration and the data generated according to the new business logic are respectively classified and statistically analyzed to obtain their respective probability distributions. For a certain local neighborhood, let the probabilities of various categories of historical data be ; where represents the Shannon entropy of the historical business metric data within the current neighborhood, represents the occurrence probability of the th category in the historical business metric data, is the total number of categories statistically counted within this neighborhood; the Shannon entropy of the new business logic data is denoted as ; where represents the Shannon entropy of the data generated according to the new business logic within the current neighborhood, represents the occurrence probability of the th category in the new business logic data.
[0038] In a preferred embodiment, S5 includes:
[0039] Receiving the generated data consistency evaluation results, and screening out the data with normal data consistency evaluation results and meeting the business cognition rules;
[0040] Writing the screened data into the data middle platform. The writing process includes recording the data according to the predefined field definitions and data formats, and attaching a unique identifier to each piece of data;
[0041] Obtaining the traceability information of the screened data. The traceability information includes the historical version numbers and corresponding processing descriptions recorded during the acquisition, processing, and evaluation of the data;
[0042] Associating and saving the data traceability information with the business logic version information. The business logic version information includes the business logic release time, the evolution history of business rules, and the corresponding version identifiers;
[0043] Performing complete retention and traceable management on the final data written into the data middle platform according to the data traceability records associated and saved with the business logic version information.
[0044] On the other hand, the present invention provides a data quality improvement system driven by data middle platform business cognition, including a business rule integration module, a logical difference analysis module, a data mapping version module, a local entropy value evaluation module, and a data traceability archiving module;
[0045] Business rule integration module: Obtaining the business cognition rules of multiple business domains, and integrating the business cognition rules into a unified business knowledge base;
[0046] Logical difference analysis module: performs difference analysis on the newly generated business logic based on the business knowledge base, and identifies the specific differences between the new business logic and the existing business logic in data fields, data processing flow and data verification rules;
[0047] Data mapping version module: maps and configures the business indicator data in the historical data according to the difference analysis results, and implements version management for the data field definition and calculation method;
[0048] Local entropy value evaluation module: Based on local neighborhood analysis, the business indicator data that has been mapped and configured in the historical data is compared with the business indicator data set generated according to the new business logic, the structural entropy is calculated in the local data distribution area based on the neighborhood division, and the data consistency evaluation result is generated according to the calculation result;
[0049] Data traceability and archiving module: writes data with normal data consistency assessment results and that meets business cognition rules into the data center, and associates and saves data traceability information with business logic version information.
[0050] The technical effects and advantages of a data quality improvement system and method driven by business cognition of a data middle platform of the present invention are as follows:
[0051] 1. The data quality improvement method driven by business cognition of the data middle platform provided by the present invention can effectively solve the problems of data traceability ambiguity and difficulty in ensuring data consistency caused by the dynamic evolution of business logic in the prior art. By integrating the business cognition rules of multiple business fields, a unified business knowledge base is constructed, and the new business logic is analyzed based on the knowledge base to accurately identify the specific differences between the new and old business rules in data fields, data processing procedures and data verification rules, so as to achieve precise mapping configuration and versioning management of business indicator data in historical data; using the local neighborhood analysis method, the structural entropy is calculated in the local data distribution area, so as to quantitatively evaluate the consistency of new and old data, and ensure that the data has accurate traceability during the update process.
[0052] 2. Further write the data that has passed the consistency assessment into the data middle platform, and associate the data traceability information with the business logic version information to build a complete data retention and traceability management system. This method can dynamically adapt to changes in business rules in complex business scenarios, avoiding the neglect of local data anomalies by global statistical methods, and overcoming the defects of untimely data updates and distorted data quality assessments in traditional mapping configuration technologies, thereby providing more accurate and comprehensive technical guarantees for enterprise data-driven decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 This is a schematic diagram of a data quality improvement method driven by business cognition of a data middle platform of the present invention;
[0054] Figure 2 This is a schematic structural diagram of a data quality improvement system driven by business cognition in a data middle platform according to the present invention. Specific implementation manners
[0055] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0056] Embodiment 1: Figure 1 A data quality improvement method driven by business cognition in a data middle platform according to the present invention is provided, which includes the following steps:
[0057] S1: Obtain business cognition rules in multiple business domains and integrate the business cognition rules into a unified business knowledge base.
[0058] S2: Based on the business knowledge base, perform differential analysis on the newly generated business logics, and identify the specific differences between the new business logics and the existing business logics in data fields, data processing processes, and data verification rules.
[0059] S3: Perform mapping configuration on the business indicator data in the historical data according to the differential analysis results, and implement version management on the data field definitions and calculation methods.
[0060] S4: Based on local neighborhood analysis, compare the mapped and configured business indicator data in the historical data with the business indicator data sets generated according to the new business logics, calculate the structural entropy within the local data distribution area divided by the neighborhood, and generate a data consistency evaluation result according to the calculation result.
[0061] S5: Write the data with normal data consistency evaluation results and meeting the business cognition rules into the data middle platform, and associate and save the data traceability information and the business logic version information.
[0062] S1 includes:
[0063] Respectively extract the original business information related to business processes, data indicator definitions, data operation rules, and data verification standards from each business system within the enterprise.
[0064] First, extract the original business information related to business processes, data metric definitions, data operation rules, and data verification standards from each business system within the enterprise. The original business information includes records from the transaction processing system, business reporting system, and rules management system, which detail the execution steps of business processes, the specific definitions of data metrics, the implementation rules of data operations, and the standards for data verification.
[0065] To ensure the accuracy and integrity of the extracted information, filter the information in each business system through preset data screening criteria to ensure that only information relevant to business cognition is extracted.
[0066] Preprocess the original business information to form structured business cognition rules. The structured business cognition rules include rule identification, rule name, rule description, applicable business area, and corresponding rule parameters.
[0067] Perform processing operations such as format conversion, data cleaning, and duplicate data elimination on the original business information to eliminate noise and redundant information in the data, and generate structured business cognition rules according to the predefined data structure requirements.
[0068] Among them, the predefined data structure requirements include rule identification, rule name, rule description, applicable business area, and rule parameters, and specify data formats, lengths, and indexing methods to ensure the stability and consistency of data quality.
[0069] Among them, the rule identification is used to uniquely identify each business cognition rule; the rule name and rule description respectively provide concise and detailed descriptions of the rule; the applicable business area clearly indicates the specific business segment within the enterprise to which the rule applies; and the corresponding rule parameters include the key data elements for implementing the business rule.
[0070] Classify, associate, and merge the structured business cognition rules, and store the processed business cognition rules uniformly in a business knowledge base with a predefined data table structure and index design.
[0071] Classify the structured rules according to preset criteria, and classify the rules into corresponding categories according to criteria such as business process stages or data domains; secondly, perform an association process on the classified rules to determine the internal connections and dependencies between the rules in different categories; finally, perform a merging process on the rules that are similar or repetitive to each other to form a unified and non-redundant set of business cognition rules.
[0072] The processed business cognition rules are then uniformly stored in the business knowledge base according to the predefined data table structure and index design, thus providing a consistent and accurate rule basis for subsequent analysis of business logic differences based on the business knowledge base.
[0073] Among them, the business knowledge base adopts a predefined data table structure, including a business cognition rule table, a rule association table, and a rule version management table. Each table contains a unique identifier, rule content, scope of application, and update time. The index design adopts a multi-level index structure based on rule identifiers, business domains, and rule versions to improve query and management efficiency.
[0074] S2 includes:
[0075] Obtain the structured business cognition rules reflecting the existing business logic, including data fields, data processing flows, and data verification rules, from the business knowledge base.
[0076] First, obtain the structured business cognition rules reflecting the existing business logic from the business knowledge base pre-constructed within the enterprise. The structured business cognition rule records generated through S1 are stored in this business knowledge base, and each record contains field information such as "rule identifier", "rule name", "rule description", "applicable business domain", and "rule parameters".
[0077] To ensure the accuracy of the acquisition process, predefined retrieval conditions and data matching rules are adopted to extract all information related to the existing business logic from the business knowledge base. Specifically, the system sets conditions, such as predefined retrieval items like "data field identifier", "processing flow version number", "verification rule code", etc., to load all target records into memory for subsequent comparison. Here, the "data field" refers to the data items defined in the existing business logic, including descriptive information such as data type, format, length, and data meaning; the "data processing flow" refers to the complete process of data transmission, conversion, processing, etc. in each business process; the "data verification rule" is the rule standard, conditions, and parameters used for verifying the accuracy of data.
[0078] Extract the business rule information in the newly generated business logic according to the same data fields, data processing flows, and data verification rules as the existing business logic to form a new business logic rule set.
[0079] The newly generated business logic comes from the new business rules introduced by the system according to market demands or business adjustments during the enterprise operation process, and its content includes data field definitions, data processing flow descriptions, and data verification rule settings.
[0080] To ensure that the extracted content is consistent with the existing business logic, the same fields, processes, and verification rule standards in the business knowledge base are adopted to match and extract the business rules involved in the new business logic. During the extraction process, the system converts the data items, processing steps, and verification criteria described in the new business logic into structured rule records according to the predefined mapping rules, and forms a "new business logic rule set". Here, each record in the new business logic rule set should have the same field structure as the existing business cognitive rules stored in the business knowledge base for subsequent comparison.
[0081] Compare each item in the new business logic rule set with the business cognitive rules corresponding to the existing business logic stored in the aforementioned business knowledge base item by item:
[0082] Analyze the specific differences in data type, data format, data length, and data meaning of each data field; compare each link in the data processing process item by item to identify the specific differences in the data input, data conversion, and data output processes of each link; compare each data verification rule item by item to identify the specific differences in each verification standard, verification condition, and rule parameter.
[0083] This comparison process uses the method of item-by-item comparison, that is, for each new business logic rule record, compare its data fields, data processing process, and data verification rules with the content described in the corresponding existing business logic rule record in turn. For data fields, the content compared by the system includes data type (such as character type, numeric type, etc.), data format (such as date format, currency format, etc.), data length (referring to the maximum number of characters or number of digits allowed for a data item), and data meaning (such as the role and meaning of each field in the business process); for the data processing process, the system compares the specific descriptions of the operation steps, order, and execution conditions in the data input, data conversion, and data output processes of each link; for the data verification rules, compare the verification standards, verification conditions, and rule parameters adopted by each rule.
[0084] To quantitatively describe each difference, for example, the following formula is used to represent the overall difference value of data fields: ; where represents the overall difference value of data fields, which is used to quantify the data field difference degree between the new business logic and the existing business logic; is the numerical representation of each data field in the existing business logic, including but not limited to data type, data format, data length, and data meaning, etc.; is the numerical representation of the corresponding data field in the new business logic, and the numericalization rule is consistent with the numerical representation of each data field in the existing business logic; and Numerical representations of data types in the existing business logic and the new business logic respectively, for example, using integers to encode different data types (such as strings, integers, floating-point numbers, etc.); and Numerical representations of data formats in the existing business logic and the new business logic respectively, such as date format, currency format, etc., which can be represented by predefined encodings; and Numerical representations of data lengths in the existing business logic and the new business logic respectively, where the specific values correspond to the maximum number of characters or digits allowed for the data item; and Numerical representations of data meanings in the existing business logic and the new business logic respectively, using the encoding method in the predefined dictionary to make the business meanings of different fields comparable; Represents the total number of data fields to be compared between the new business logic and the existing business logic.
[0085] The above formula is only an example. In actual implementation, various information can be converted into standardized numerical values according to predefined rules for difference calculation. After adopting this formula, if the overall difference value of the calculated result data fields is greater than its corresponding preset threshold, it is determined that there are significant differences in the corresponding data fields. Similarly, for the comparison of data processing flows and data verification rules, corresponding quantitative calculation methods can also be adopted. However, in this embodiment, the main method is to compare item by item and record item by item, manually or automatically identify the specific operation descriptions of each link, and mark the existing differences.
[0086] Record the specific differences in the predefined format to form the difference analysis result between the new business logic and the existing business logic.
[0087] During the comparative analysis process, record all identified specific difference information according to the predefined comparison format; specifically, this predefined format includes difference categories (data fields, data processing flows, or data verification rules), difference items (such as data types, data formats, verification conditions, etc.), existing rule values, new rule values, and descriptions of the degree of difference. Each record is attached with a unique identifier for subsequent tracking and statistics. Here, each difference between the new business logic and the existing business logic should be recorded according to the same standard to ensure the continuity and consistency of the data. After sorting out all the difference record information, it constitutes the "difference analysis result", which serves as an important basis for subsequent data quality assessment and version management.
[0088] During the process of differential analysis, it is necessary to ensure the continuity of data transfer between sub-steps. For example, first, the structured business cognitive rules extracted from the business knowledge base maintain consistent field information when new business logic is extracted; second, each new business logic rule in the comparison process clearly corresponds to the specific records of the existing business logic in the business knowledge base, ensuring the consistency of the comparison basis; finally, the differential data obtained from the comparison is stored in a unified record format for subsequent system calls and data queries.
[0089] S3 includes:
[0090] Perform mapping configuration on the business metric data in the historical data. The mapping configuration includes:
[0091] a. Compare the data records reflecting business metrics in the historical data with the business rule information involved in the new business logic item by item according to the predefined mapping criteria;
[0092] b. Determine the corresponding relationships between the business metric data in the historical data and the corresponding business rules in the new business logic;
[0093] c. Store the corresponding relationships in the form of mapping records.
[0094] Among them, first, extract the data records reflecting business metrics from the historical data storage system. These data records contain various data field information corresponding to each business metric, such as field name, data type, data format, data length, and business meaning. At the same time, obtain the structured business cognitive rules of the existing business logic stored in the S2 step from the business knowledge base. The rule record details the predefined standards of the data fields and the business rule information involved in the new business logic.
[0095] To ensure data consistency, the system preliminarily sorts out the extracted historical data fields and the business rule information in the new business logic according to the predefined mapping criteria. The mapping criteria clearly stipulate the matching requirements of each data field in terms of name, data type, data format, data length, and business meaning.
[0096] Specifically, the field name, data type, data format, data length, and data meaning of the historical data fields are compared in turn according to predefined mapping criteria. For example, if the field name of a certain business indicator in the historical data is "transaction amount", and the data type is numeric, the data format is currency format, the data length is 12 digits, and the business meaning is "record the transaction amount", then the corresponding business rule information in the new business logic should also include a field named "transaction amount" or a synonymous description, and have similar numeric type, currency format, 12-digit length, and similar business meaning. By comparing item by item, the system automatically determines the corresponding relationship between each business indicator field in the historical data and the corresponding business rules in the new business logic, thus forming a mapping correspondence. This correspondence not only includes the one-to-one matching situation between fields, but also records the subtle differences found during the matching process for reference during subsequent data quality assessment.
[0097] The above mapping correspondence is stored in the form of a mapping record. The mapping record details the matching information between each pair of corresponding historical data fields and new business logic fields, including field name, data type, data format, data length, business meaning, and matching status. The mapping records are sorted according to predefined data formats and stored uniformly in the mapping configuration database within the system for subsequent data processing and querying. The mapping configuration database adopts a predefined data table structure and index design to ensure efficient retrieval and management of the mapping records.
[0098] Implement version management for the data field definitions included in the data related to business indicators in the historical data. The data field definitions include the name, data type, data format, data length, and data meaning of the data fields. The version management records, identifies, and subsequently updates each data field definition according to predefined version control rules.
[0099] The original data field definitions in the historical data are recorded as the initial version according to the predefined version control rules, for example, identified as version V1. The predefined version control rules clearly stipulate that when the business logic changes or the relevant standards of the data fields are updated, the system should generate a new version record according to the changed content. Specifically, when it is detected that there are differences between the description of a certain data field in the new business logic and the existing description, the system will retain the original version record in the historical data and update the data field definition according to the new business logic to generate a new version, for example, identified as version V2. All versioned records include the version number, version generation time, and an explanation of the version update reason, and are stored in the version management database for subsequent data traceability and version comparison. This version management process ensures that the historical data can accurately trace the corresponding data field definition version during subsequent data processing and consistency assessment, thus guaranteeing the consistency and compatibility of data processing.
[0100] Implement version management for the calculation methods of business metric data in historical data. The calculation methods include the formulas for business metric calculations, calculation parameters, and data processing rules. Version management records and manages different versions of the calculation methods based on predefined version identifiers and version update rules.
[0101] The calculation methods mainly include the formulas for business metric calculations, calculation parameters, and data processing rules. Initially, the system records the calculation methods of business metrics used in historical data to form an initial version record, which details the calculation formula (such as each operator, operation order, and each data item participating in the calculation), calculation parameters (such as proportional coefficients, constant values, etc.), and data processing rules (such as data preprocessing, outlier handling methods, etc.). When new business logic introduces changes that result in adjustments to the business metric calculation formula or related parameters, the system, based on the predefined version update rules, records the changed content as a new version of the calculation method, generates a new version record, and assigns a unique version identifier to this record. The predefined version update rules clearly stipulate that each change must retain the complete record of the previous version and record the comparison between the new version and the previous version to ensure that historical data can trace back to the original calculation logic when applying the new calculation method. All calculation method version records are stored in a dedicated version management system using a unified record format to ensure that subsequent data consistency evaluation and data processing operations can be performed based on the correct calculation version.
[0102] S4 includes:
[0103] Obtain the historical business metric data after mapping configuration and the business metric data set generated according to the new business logic, and perform preliminary merging processing on the two.
[0104] Obtain the historical business metric data after mapping configuration from the aforementioned S3 step. This data records the detailed definitions and related mapping relationships of each business metric in historical data under the existing business logic. At the same time, the business metric data set generated according to the new business logic has also been formed, reflecting the calculation results of each business metric under the new business logic. For subsequent comparative analysis, first perform preliminary merging processing on the above two parts of data. The merging processing includes unifying the data format, aligning the field names, and handling missing data items, so that the merged data set contains both the original attributes of historical data and incorporates the rule information of each item in the new business logic, forming a comprehensive data set as the input data for subsequent local analysis.
[0105] Divide the data after merging processing into neighborhoods according to the local neighborhood analysis method. The neighborhood division groups different records according to the data distribution characteristics.
[0106] After the initial data merging is completed, a local neighborhood analysis method is used to divide the merged data into neighborhoods. Specifically, according to the distribution characteristics of data records on key fields (such as business metric values, timestamps, category identifiers, etc.), a predefined neighborhood division rule is used to group the data set. The division rule can be based on the statistical distribution of data records, such as density or similarity measure, to group records with similar distribution characteristics into the same neighborhood. Each neighborhood is defined here as a local data distribution area, and each neighborhood is assigned a unique neighborhood identifier. Through this neighborhood division process, it is ensured that when calculating the structural entropy in each local area subsequently, the data has a clear attribution range and statistical basis.
[0107] Within the local data distribution area corresponding to each neighborhood, the entropy value of the data is calculated based on a pre-set structural entropy calculation method, and the structural entropy calculation method measures the distribution characteristics of each record in the same neighborhood on multiple field dimensions.
[0108] After the neighborhood division is completed, the structural entropy of the data is calculated within each neighborhood to quantitatively describe the consistency of the distribution of old and new business metric data on each key field dimension within the same local data distribution area. A measurement method based on information entropy is adopted, that is, by calculating the Shannon entropy of the distribution of old and new data within each neighborhood, and using their absolute difference as the structural entropy difference index.
[0109] Within each local neighborhood, first, the historical business metric data after mapping configuration and the data generated according to the new business logic are respectively classified and statistically analyzed to obtain their respective probability distributions. Specifically, for a certain local neighborhood, let the probability of each category of historical data be ; among them, represents the Shannon entropy of the historical business metric data within the current neighborhood, describing the uncertainty of this data distribution; represents the occurrence probability of the th category in the historical business metric data, is the total number of categories statistically counted within this neighborhood; similarly, the Shannon entropy of the new business logic data is denoted as ; among them, represents the Shannon entropy of the data generated according to the new business logic within the current neighborhood; represents the occurrence probability of the th category in the new business logic data. To ensure the consistency of the comparison, the category division here is kept consistent according to the predefined rules, and both take values from the same set.
[0110] Subsequently, the local structural entropy difference is defined as ; among them, As a quantitative indicator of the consistency of data distribution in the local neighborhood, it is used to reflect the degree of difference in statistical distribution between historical data and new business logic data.
[0111] It is the difference in local structural entropy, which indicates the difference in structural entropy between two sets of data in the current neighborhood. The smaller the value, the more consistent the distribution of new and old data is, and vice versa.
[0112] Compare the entropy value differences between historical business indicator data and new business logic business indicator data in each neighborhood, and use the entropy value difference results as the basis for local consistency judgment.
[0113] After calculating the local structural entropy difference in each neighborhood, the data consistency in each neighborhood is judged according to the preset consistency judgment standard.
[0114] Specifically, the preset consistency judgment standard includes one or more thresholds. When the difference in local structural entropy within a local neighborhood is less than or equal to the threshold, the data consistency assessment result is judged to be normal, that is, the distribution of new and old business indicator data in the neighborhood is consistent; conversely, if the difference in local structural entropy exceeds the threshold, it is considered that there is a large distribution difference in the neighborhood and the consistency requirements are not met.
[0115] Among them, the preset consistency judgment standard includes a threshold range set based on the structural entropy difference. When the structural entropy difference value calculated in the local neighborhood is lower than or equal to the threshold, the data distribution of the new and old business indicators in the neighborhood is judged to be consistent. Otherwise, it is determined that there are significant differences in the data, and the corresponding judgment results are recorded.
[0116] The judgment result is output in the form of a record, and detailed entropy data and comparison results are attached in each neighborhood to form a local consistency judgment report.
[0117] The data consistency evaluation result is formed based on the entropy value difference results of each neighborhood, and the data consistency evaluation result is associated with the previous mapping configuration information and stored.
[0118] The judgment results in all local neighborhoods are aggregated to form an overall data consistency assessment result. The aggregation process includes counting the values of the local structural entropy differences in each neighborhood, calculating the overall consistency score, and classifying and integrating the local judgment results. The overall data consistency assessment result formed includes the entropy value differences, judgment conclusions, and corresponding neighborhood identifiers of each local data distribution area. To facilitate subsequent data tracing and application, the aggregated overall data consistency assessment results are associated and stored with the mapping records previously generated in the S3 step. The storage format follows the predefined data table structure and index design requirements to ensure that the data consistency assessment results can be quickly retrieved and called in the data center.
[0119] The local neighborhood analysis method adopted in this embodiment calculates the structural entropy within the local distribution area of the historical service metric data after mapping configuration and the data set generated according to the new service logic, realizes the quantitative evaluation of the local data distribution difference, can finely capture the local data abnormal changes caused by the dynamic evolution of the service logic, divides the data set into neighborhoods according to predefined criteria, and calculates the entropy value difference based on local statistical features, thereby generating a data consistency evaluation result. Combined with the dynamic adjustment of business rules, it organically connects mapping configuration, version management and local data distribution analysis, avoiding both the problem that the global average value masks local fluctuations and realizing real-time, accurate and regional management of data quality monitoring.
[0120] S5 includes:
[0121] Receiving the generated data consistency evaluation result, and screening out the data with normal data consistency evaluation result and meeting the business cognitive rules.
[0122] Among them, the evaluation result records the structural entropy difference calculated by comparing the historical service metric data after mapping configuration with the service metric data set generated according to the new service logic within the local neighborhood, and makes a consistency determination for each local data area according to the preset consistency determination criteria. According to the predefined business cognitive rules, the evaluation results are screened to eliminate the data with abnormal consistency evaluation results or not meeting the preset consistency requirements. The screening process strictly follows the requirements of the data field definition, data processing flow and data verification rules in the business cognitive rules to ensure that only the records that match the old and new service logics and have consistent data distributions are retained. The screened data not only represents the qualified part of the data consistency evaluation result, but also provides an accurate basis for subsequent data writing and data traceability.
[0123] Writing the screened data into the data middle platform. The writing process includes recording the data according to the pre-set field definitions and data formats, and attaching a unique identifier to each piece of data.
[0124] According to the pre-set field definitions and data formats, write the screened qualified data records into the data middle platform. The writing process includes format conversion, field alignment and data completion for each data record to ensure that the written data matches the predefined database table structure in the data middle platform.
[0125] For the convenience of subsequent data management and query, a unique identifier is attached to each written data record. The unique identifier is generated by the system according to the preset rules and can ensure no duplication in the entire data middle platform. The writing operation strictly follows the predefined record specifications, making the data content, data field definitions consistent with the previous mapping configuration and consistency evaluation results, ensuring data integrity and consistency.
[0126] Obtain the traceability information of the filtered data. The traceability information includes the historical version numbers and corresponding processing descriptions recorded during the acquisition, processing, and evaluation of the data.
[0127] Obtain data traceability information for the filtered data records. The data traceability information includes all historical records generated during the processes of data collection, preprocessing, mapping configuration, difference analysis, local neighborhood analysis, and data consistency evaluation. Specifically, it includes the historical version numbers, timestamps, descriptions of processing steps, and corresponding processing descriptions recorded at each stage.
[0128] This traceability information is generated through the system's automatic recording and log management functions and stored in a dedicated traceability database in a structured record format. The traceability information ensures that each piece of written data can be traced back to the entire process of its generation and processing, providing detailed basis for subsequent data auditing and version updates.
[0129] Associate and save the data traceability information with the business logic version information. The business logic version information includes the release time of the business logic, the evolution history of the business rules, and the corresponding version identifiers.
[0130] Associate and save the data traceability information with the business logic version information. The business logic version information includes the release time of the business logic, the evolution history of the business rules, and the corresponding version identifiers. The system matches the historical version numbers in the data traceability information with the version identifiers in the business logic version information according to the pre-set association rules and records the matching results in the data center.
[0131] The association and saving process adopts a structured data record format. The recorded content includes the unique identifier of the data record, the corresponding traceability information, the business logic version identifier, the release time of the business logic, etc., ensuring that each piece of written data in the data center is closely associated with the corresponding business logic version information, facilitating version comparison during subsequent data tracing and business rule adjustment.
[0132] Based on the data traceability records associated with the business logic version information, perform complete retention and traceable management on the final data written to the data center.
[0133] This management process includes regular backups, data integrity verification, and log archiving to ensure that the data can be traced and verified at each stage of its life cycle. The system establishes a unified query interface that can quickly retrieve the historical processing records of the data based on the data unique identifier, traceability information, or business logic version information.
[0134] Through complete retention and traceable management, the full process supervision after data writing is realized, ensuring that the data can provide sufficient historical basis during future business adjustments and meet the requirements of data consistency and version management.
[0135] Embodiment 2: The difference between Embodiment 2 and Embodiment 1 of the present invention is that this embodiment introduces a data quality improvement system driven by business cognition in the data middle platform.
[0136] Figure 2 The structural schematic diagram of a data quality improvement system driven by business cognition in the data middle platform of the present invention is given. A data quality improvement system driven by business cognition in the data middle platform includes a business rule integration module, a logical difference analysis module, a data mapping version module, a local entropy value evaluation module, and a data traceability and archiving module.
[0137] Business rule integration module: Obtain business cognition rules in multiple business domains and integrate the business cognition rules into a unified business knowledge base.
[0138] Logical difference analysis module: Based on the business knowledge base, perform difference analysis on the newly generated business logic, and identify the specific differences between the new business logic and the existing business logic in terms of data fields, data processing processes, and data verification rules.
[0139] Data mapping version module: Perform mapping configuration on the business metric data in the historical data according to the difference analysis results, and implement version management on the data field definitions and calculation methods.
[0140] Local entropy value evaluation module: Based on local neighborhood analysis, compare the business metric data with mapped configuration in the historical data with the business metric data set generated according to the new business logic, calculate the structural entropy within the local data distribution area divided by the neighborhood, and generate a data consistency evaluation result according to the calculation result.
[0141] Data traceability and archiving module: Write the data with normal data consistency evaluation results and meeting the business cognition rules into the data middle platform, and associate and save the data traceability information with the business logic version information.
[0142] The above formulas are all dimensionless and take their numerical values for calculation. The formula is obtained by collecting a large amount of data and performing software simulation to obtain a formula that is closest to the actual situation. The preset parameters and threshold selection in the formula are set by those skilled in the art according to the actual situation.
[0143] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0144] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and modules described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0145] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings, direct couplings, or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or modules can be in electrical, mechanical, or other forms.
[0146] The modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules. They can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0147] In addition, in each embodiment of the present application, each functional module can be integrated into a processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.
[0148] If the above-mentioned function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0149] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0150] Finally: The above is only the preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for improving data quality driven by business cognition of a data middle platform, characterized in that It includes the following steps: S1: Obtain business recognition rules in multiple business domains and integrate the business recognition rules into a unified business knowledge base; S2: Based on the business knowledge base, perform differential analysis on the newly generated business logic, and identify the specific differences between the new business logic and the existing business logic in terms of data fields, data processing flows, and data verification rules; S3: According to the results of the differential analysis, perform mapping configuration on the business indicator data in the historical data, and implement version management for the data field definitions and calculation methods, including: Perform mapping configuration on the business indicator data in the historical data; Implement version management for the data field definitions included in the data related to business indicators in the historical data. The data field definitions include the name, data type, data format, data length, and data meaning of the data fields. The version management records, identifies, and subsequently updates each data field definition according to predefined version control rules; Implement version management for the calculation methods of the business indicator data in the historical data. The calculation methods include the formulas, calculation parameters, and data processing rules for business indicator calculations. The version management records and manages different versions of the calculation methods according to predefined version identifiers and version update rules; Among them, the mapping configuration includes: a. Compare item by item the data reflecting business indicators in the historical data with the business rule information involved in the new business logic according to predefined mapping criteria; b. Determine the corresponding relationships between the business indicator data in the historical data and the corresponding business rules in the new business logic; c. Store the corresponding relationships in the form of mapping records; S4: Based on local neighborhood analysis, compare the business indicator data in the historical data that has undergone mapping configuration with the business indicator data set generated according to the new business logic. Calculate the structural entropy within the local data distribution area divided by the neighborhood, and generate a data consistency evaluation result according to the calculation result; S5: Write the data with normal data consistency evaluation results and meeting the business recognition rules into the data middle platform, and associate and save the data traceability information and business logic version information.
2. The data quality improvement method driven by business cognition of the data middle platform according to claim 1, characterized in that S1 includes: Extract the original business information involving business processes, data indicator definitions, data operation rules, and data verification standards from each business system within the enterprise respectively; Preprocess the original business information to form structured business recognition rules. The structured business recognition rules include rule identifiers, rule names, rule descriptions, applicable business domains, and corresponding rule parameters; Classify, associate, and merge the structured business recognition rules, and uniformly store the processed business recognition rules in a business knowledge base with a predefined data table structure and index design.
3. A method for improving data quality driven by business cognition of a data middle platform according to claim 1, characterized in that, S2 includes: Obtain the structured business recognition rules of data fields, data processing flows, and data verification rules reflecting the existing business logic from the business knowledge base; Extract the business rule information in the newly generated business logic according to the same data fields, data processing flows, and data verification rules as the existing business logic to form a set of new business logic rules; Compare the new business logic rule set item by item with the business recognition rules corresponding to the existing business logic stored in the aforementioned business knowledge base: Analyze the specific differences in data type, data format, data length, and data meaning of each data field; compare each link in the data processing process item by item to identify the specific differences in the data input, data conversion, and data output processes of each link; compare each data verification rule item by item to identify the specific differences in each verification standard, verification condition, and rule parameter; Record the specific differences in a predefined format to form the difference analysis result between the new business logic and the existing business logic.
4. A method for improving data quality driven by business cognition of a data middle platform according to claim 1, characterized in that, S4 includes: Obtain the historical business metric data after mapping configuration and the business metric data set generated according to the new business logic, and perform preliminary merging processing on the two; Divide the merged data into neighborhoods according to the local neighborhood analysis method. The neighborhood division groups different records according to the data distribution characteristics; Within the local data distribution area corresponding to each neighborhood, calculate the entropy value of the data based on a pre-set structural entropy calculation method. The structural entropy calculation method measures the distribution characteristics of each record in multiple field dimensions within the same neighborhood; Compare the entropy value differences between the historical business metric data and the new business logic business metric data in each neighborhood, and use the entropy value difference result as the basis for local consistency determination; Summarize the data consistency evaluation results based on the entropy value difference results of each neighborhood, and store the data consistency evaluation results in association with the previous mapping configuration information.
5. A method for improving data quality driven by business cognition of a data middle platform according to claim 4, characterized in that, Calculate the entropy value of the data based on a pre-set structural entropy calculation method, specifically: In each local neighborhood, the historical business metric data after mapping configuration and the data generated according to the new business logic are respectively classified and statistically analyzed to obtain their respective probability distributions. For a certain local neighborhood, let the probabilities of various categories of historical data be ; where represents the Shannon entropy of the historical business metric data in the current neighborhood, represents the occurrence probability of the -th category in the historical business metric data, is the total number of categories statistically analyzed in this neighborhood; the Shannon entropy of the new business logic data is denoted as ; where represents the Shannon entropy of the data generated according to the new business logic in the current neighborhood, represents the occurrence probability of the -th category in the new business logic data.
6. A method for improving data quality driven by business cognition of a data middle platform according to claim 1, characterized in that S5 Include: Receive the generated data consistency evaluation results, and filter out the data with normal data consistency evaluation results and meeting the business recognition rules; Write the filtered data into the data middle platform. The writing process includes recording the data according to the pre-set field definition and data format, and attaching a unique identifier to each piece of data; Obtain the traceability information of the filtered data. The traceability information includes the historical version number and corresponding processing description recorded in the data acquisition, processing, and evaluation links; Associate and save the data traceability information with the business logic version information. The business logic version information includes the business logic release time, the evolution process of business rules, and the corresponding version identifier; Perform complete retention and traceable management on the final data written into the data middle platform based on the data traceability record associated with the business logic version information.
7. A data quality improvement system driven by business cognition of a data middle platform, used to implement a data quality improvement method driven by business cognition of a data middle platform as described in any one of claims 1-6, characterized in that: Include a business rule integration module, a logical difference analysis module, a data mapping version module, a local entropy value evaluation module, and a data traceability archiving module; Business rule integration module: Obtain the business recognition rules of multiple business domains, and integrate the business recognition rules into a unified business knowledge base; Logical difference analysis module: Based on the business knowledge base, perform difference analysis on the newly generated business logic, and identify the specific differences between the new business logic and the existing business logic in data fields, data processing processes, and data verification rules; Data mapping version module: Perform mapping configuration on the business metric data in the historical data according to the difference analysis result, and implement version management on the data field definition and calculation method, including: Perform mapping configuration on the business metric data in historical data; Implement version management for the data field definitions included in the data related to business metrics in historical data. The data field definitions include the name, data type, data format, data length, and data meaning of the data fields. Version management records, identifies, and subsequently updates each data field definition according to predefined version control rules; Implement version management for the calculation methods of business metric data in historical data. The calculation methods include the formulas, calculation parameters, and data processing rules for business metric calculations. Version management records and manages different versions of calculation methods based on predefined version identifiers and version update rules; Among them, the mapping configuration includes: a. Compare item by item the data reflecting business metrics in historical data with the business rule information involved in the new business logic according to predefined mapping criteria; b. Determine the corresponding relationships between the business metric data in historical data and the corresponding business rules in the new business logic; c. Store the corresponding relationships in the form of mapping records; Local entropy value evaluation module: Compare the business metric data in historical data that has undergone mapping configuration with the business metric data set generated based on the new business logic through local neighborhood analysis. Calculate the structural entropy within the local data distribution area divided by neighborhoods, and generate a data consistency evaluation result based on the calculation result; Data traceability and archiving module: Write the data with normal data consistency evaluation results that meet the business cognition rules into the data middle platform, and associate and save the data traceability information with the business logic version information.
Citation Information
Patent Citations
A rule-based automated data governance system and method
CN119782715A