Data quality improvement system and method driven by service cognition of data-in-data station

By integrating technical means such as business cognitive rules, difference analysis, data mapping and local entropy value evaluation in Taichung in the data, the data traceability ambiguity and consistency problems of data middle platform when business logic changes are solved, and the data quality and reliability are improved, providing stronger support for enterprise decision-making.

CN120067091AActive Publication Date: 2025-05-30CHANGSHA DILU DIGITAL TECH

Patent Information

Application Number
CN202510535082.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-05-30
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

When the existing data middle platform supports business cognitive-driven data quality improvement, it is difficult to effectively identify, record and manage business logic changes, resulting in data traceability ambiguity and data consistency difficult to ensure, affecting the reliability of data-driven decision-making.

Method used

By obtaining business cognitive rules in multiple business fields, integrating them into a unified business knowledge base, and analyzing the differences in new business logic based on this knowledge base, identifying the differences between new and old business logic, and mapping, configuration and versioning management of business indicator data in historical data. The local neighborhood analysis method is used to calculate structural entropy, evaluate data consistency, and correlate the evaluation results with the data traceability information.

Benefits of technology

It realizes effective management of the dynamic evolution of business logic, ensures data consistency and traceability, improves data quality and reliability, and provides more accurate and comprehensive technical guarantees for enterprise data-driven decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067091A_ABST
    Figure CN120067091A_ABST
Patent Text Reader

Abstract

The invention discloses a data quality improvement method driven by service cognition in data, particularly relates to the technical field of data quality management, and is used for solving the problems of poor data traceability ambiguity consistency and evaluation distortion caused by existing service evolution. According to the method, a unified business knowledge base is constructed by integrating business cognition rules of multiple business fields, the difference between new business logic and existing business logic is analyzed based on the knowledge base, and business indexes in historical data are subjected to mapping configuration and version management; a local neighborhood analysis method is utilized to calculate structure entropy in a local data distribution area, a data consistency evaluation result is generated, finally, data meeting a business cognition rule is written into a data table, and meanwhile data traceability information and business logic version information are stored in an associated mode. Therefore, the problems of data traceability ambiguity and quality evaluation distortion in the dynamic evolution of the service logic of the existing data medium station are effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data quality management. More specifically, the present invention relates to a data quality improvement system and method driven by business cognition in a data middle platform. Background Art

[0002] When existing data middle platforms support data quality improvement driven by business cognition, they usually rely on predefined business rules and data quality standards. However, in complex business scenarios, business logics will evolve continuously with market demands, policy adjustments, or enterprise strategic upgrades, resulting in changes in data processing rules, field definitions, calculation methods, etc.

[0003] When the generation, storage, and processing logics of data are adjusted, historical data may not be directly adaptable to the new rules, thus leading to ambiguity in data traceability. For example, a certain business metric may be generated based on different calculation logics at different time points, resulting in inconsistent meanings of the same metric in different historical versions, making it difficult for data quality verification to accurately adapt to business requirements. Existing technologies are difficult to effectively identify, record, and manage such business logic changes, which may cause difficulties in data traceability, distortion of quality assessment, and even affect the reliability of data-driven decision-making. Summary of the Invention

[0004] To overcome the above-mentioned defects of the prior art, embodiments of the present invention provide a data quality improvement system and method driven by business cognition in a data middle platform to solve the problems raised in the above background art.

[0005] To achieve the above object, the present invention provides the following technical solutions: A data quality improvement method driven by business cognition in a data middle platform, comprising the following steps: S1: Obtain business cognition rules in multiple business domains and integrate the business cognition rules into a unified business knowledge base; S2: Perform differential analysis on the newly generated business logic based on the business knowledge base to identify the specific differences between the new business logic and the existing business logic in data fields, data processing processes, and data verification rules; S3: Perform mapping configuration on the business metric data in historical data according to the differential analysis results, and implement version management for data field definitions and calculation methods; S4: Compare the mapped business metric data in historical data with the business metric data set generated according to the new business logic based on local neighborhood analysis, calculate the structural entropy within the local data distribution area divided by neighborhoods, and generate a data consistency evaluation result according to the calculation result; S5: Write the data with normal data consistency evaluation results that meet the business cognition rules into the data middle platform, and associate and save the data traceability information with the business logic version information.

[0006] In a preferred embodiment, S1 includes: Extract the original business information related to business processes, data metric definitions, data operation rules, and data verification criteria from each business system within the enterprise respectively; Preprocess the original business information to form structured business cognition rules. The structured business cognition rules include rule identification, rule name, rule description, applicable business fields, and corresponding rule parameters; Classify, associate, and merge the structured business cognition rules, and uniformly store the processed business cognition rules in a business knowledge base with a predefined data table structure and index design.

[0007] In a preferred embodiment, S2 includes: Obtain the structured business cognition rules reflecting the existing business logic, including data fields, data processing flows, and data verification rules, from the business knowledge base; Extract the business rule information in the newly generated business logic according to the same data fields, data processing flows, and data verification rules as the existing business logic to form a new business logic rule set; Compare each item of the new business logic rule set with the business cognition rules corresponding to the existing business logic stored in the aforementioned business knowledge base item by item: Analyze the specific differences in data type, data format, data length, and data meaning of each data field; compare each link of the data processing flow item by item to identify the specific differences in the processes of data input, data conversion, and data output; compare each data verification rule item by item to identify the specific differences in each verification standard, verification condition, and rule parameter; Record the specific differences in a predefined format to form the difference analysis result between the new business logic and the existing business logic.

[0008] In a preferred embodiment, S3 includes: Perform mapping configuration on the business metric data in the historical data; Implement version management for the data field definitions of the data related to business metrics in the historical data. The data field definitions include the name, data type, data format, data length, and data meaning of the data fields. The version management records, identifies, and subsequently updates each data field definition according to the predefined version control rules. Implement version management for the calculation methods of business indicator data in historical data. The calculation methods include formulas for calculating business indicators, calculation parameters, and data processing rules. Version management records and manages different versions of calculation methods based on predefined version identifiers and version update rules.

[0009] In a preferred embodiment, the mapping configuration includes: a. Compare item by item the data reflecting business indicators in historical data with the business rule information involved in the new business logic according to predefined mapping criteria; b. Determine the correspondence between each business indicator data in historical data and the corresponding business rules in the new business logic; c. Store the correspondence in the form of mapping records.

[0010] In a preferred embodiment, S4 includes: Obtain the historical business indicator data after mapping configuration and the business indicator data set generated according to the new business logic, and perform preliminary merging processing on the two; Divide the merged data into neighborhoods according to the local neighborhood analysis method. Neighborhood division groups different records according to the data distribution characteristics; Within the local data distribution area corresponding to each neighborhood, calculate the entropy value of the data based on a preset structural entropy calculation method. The structural entropy calculation method measures the distribution characteristics of each record in multiple field dimensions within the same neighborhood; Compare the entropy value differences between the historical business indicator data and the new business logic business indicator data in each neighborhood, and use the entropy value difference results as the basis for local consistency determination; Summarize the data consistency evaluation results according to the entropy value difference results of each neighborhood, and store the data consistency evaluation results in association with the previous mapping configuration information.

[0011] In a preferred embodiment, calculate the entropy value of the data based on a preset structural entropy calculation method, specifically: Within each local neighborhood, classify and count the historical business indicator data after mapping configuration and the data generated according to the new business logic respectively to obtain their respective probability distributions; for a certain local neighborhood, denote the probabilities of each category of historical data as ; where represents the Shannon entropy of the historical business indicator data in the current neighborhood, represents the occurrence probability of the th category in the historical business indicator data, is the total number of categories counted in this neighborhood; the Shannon entropy of the new business logic data is denoted as ; where represents the Shannon entropy of the data generated according to the new business logic within the current neighborhood, represents the occurrence probability of the th category in the new business logic data.

[0012] In a preferred embodiment, S5 includes: Receiving the generated data consistency evaluation result, and screening out the data with normal data consistency evaluation result and meeting the business cognition rules; Writing the screened data into the data middle platform, and the writing process includes recording the data according to the pre-set field definitions and data formats, and attaching a unique identifier to each piece of data; Obtaining the traceability information of the screened data, where the traceability information includes the historical version numbers and corresponding processing descriptions recorded in the data acquisition, processing, and evaluation links; Associating and saving the data traceability information with the business logic version information, where the business logic version information includes the business logic release time, the evolution process of business rules, and the corresponding version identifiers; Based on the data traceability records associated and saved with the business logic version information, performing complete retention and traceable management on the final data written into the data middle platform.

[0013] On the other hand, the present invention provides a data quality improvement system driven by business cognition of the data middle platform, including a business rule integration module, a logic difference analysis module, a data mapping version module, a local entropy value evaluation module, and a data traceability archiving module; Business rule integration module: Obtaining the business cognition rules of multiple business domains, and integrating the business cognition rules into a unified business knowledge base; Logic difference analysis module: Based on the business knowledge base, performing difference analysis on the newly generated business logic, and identifying the specific differences between the new business logic and the existing business logic in data fields, data processing processes, and data verification rules; Data mapping version module: According to the difference analysis results, performing mapping configuration on the business indicator data in the historical data, and implementing version management on the data field definitions and calculation methods; Local entropy value evaluation module: Based on local neighborhood analysis, comparing the business indicator data configured by mapping in the historical data with the business indicator data set generated according to the new business logic, calculating the structural entropy within the local data distribution area divided by the neighborhood, and generating a data consistency evaluation result according to the calculation result; Data traceability archiving module: Writing the data with normal data consistency evaluation result and meeting the business cognition rules into the data middle platform, and associating and saving the data traceability information with the business logic version information.

[0014] Technical effects and advantages of a data quality improvement system and method driven by business cognition in a data middle platform according to the present invention: 1. The data quality improvement method driven by business cognition in the data middle platform provided by the present invention can effectively solve problems such as data traceability ambiguity and difficulty in ensuring data consistency caused by the dynamic evolution of business logic in the prior art. By integrating business cognition rules in multiple business domains, a unified business knowledge base is constructed, and based on this knowledge base, differential analysis of new business logics is carried out to accurately identify the specific differences between old and new business rules in data fields, data processing flows, and data verification rules, realizing precise mapping configuration and version management of business indicator data in historical data; using the local neighborhood analysis method, structural entropy is calculated within the local data distribution area to quantitatively evaluate the consistency of old and new data, ensuring accurate traceability of data during the update process.

[0015] 2. Further, the data that has passed the consistency evaluation is written into the data middle platform, and the data traceability information is associated and saved with the business logic version information, constructing a complete data retention and traceability management system. This method can dynamically adapt to business rule changes in complex business scenarios, avoiding both the neglect of local data anomalies by global statistical methods and the defects of untimely data updates and distorted data quality evaluation in traditional mapping configuration technologies, thus providing more accurate and comprehensive technical support for enterprise data-driven decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a schematic diagram of a data quality improvement method driven by business cognition in a data middle platform according to the present invention; Figure 2 It is a schematic structural diagram of a data quality improvement system driven by business cognition in a data middle platform according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0018] Embodiment 1: Figure 1 A data quality improvement method driven by business cognition in a data middle platform according to the present invention is given, which includes the following steps: S1: Obtain business cognition rules in multiple business domains and integrate the business cognition rules into a unified business knowledge base.

[0019] S2: Based on the business knowledge base, perform differential analysis on the newly generated business logic to identify the specific differences between the new business logic and the existing business logic in terms of data fields, data processing flows, and data verification rules.

[0020] S3: According to the results of the differential analysis, perform mapping configuration on the business metric data in the historical data, and implement version management for data field definitions and calculation methods.

[0021] S4: Based on local neighborhood analysis, compare the business metric data with mapped configuration in the historical data and the business metric data set generated according to the new business logic, calculate the structural entropy within the local data distribution area divided by the neighborhood, and generate a data consistency evaluation result according to the calculation result.

[0022] S5: Write the data with normal data consistency evaluation results that meet the business cognitive rules into the data middle platform, and associate and save the data traceability information and business logic version information.

[0023] S1 includes: Extract the original business information related to business processes, data metric definitions, data operation rules, and data verification standards from each business system within the enterprise respectively.

[0024] First, extract the original business information related to business processes, data metric definitions, data operation rules, and data verification standards from each business system within the enterprise respectively. The original business information includes records from the transaction processing system, business report system, and rule management system, which describe in detail the execution steps of business processes, the specific definitions of data metrics, the implementation rules of data operations, and the standards based on which data verification is performed.

[0025] To ensure the accuracy and integrity of the extracted information, filter the information in each business system through preset data screening criteria to ensure that only the information related to business cognition is extracted.

[0026] Preprocess the original business information to form structured business cognitive rules. The structured business cognitive rules include rule identifiers, rule names, rule descriptions, applicable business domains, and corresponding rule parameters.

[0027] Perform processing operations such as format conversion, data cleaning, and duplicate data elimination on the original business information to eliminate noise and redundant information in the data, and generate structured business cognitive rules according to the predefined data structure requirements.

[0028] Among them, the predefined data structure requirements include rule identifiers, rule names, rule descriptions, applicable business domains, and rule parameters, and specify data formats, lengths, and indexing methods to ensure the stability and consistency of data quality.

[0029] Among them, the rule identifier is used to uniquely identify each business cognitive rule; the rule name and rule description respectively provide a concise and detailed description of the rule; the applicable business area clearly indicates the specific business segment within the enterprise to which the rule applies; and the corresponding rule parameters include the key data elements used to implement the business rule.

[0030] Classify, associate, and merge the structured business cognitive rules, and uniformly store the processed business cognitive rules in a business knowledge base with a predefined data table structure and index design.

[0031] Classify the structured rules according to predefined criteria, and classify the rules into corresponding categories according to criteria such as business process stages or data domains; secondly, perform an association process on the classified rules to determine the internal connections and dependencies between the rules in different categories; finally, perform a merging process on the rules that are similar or repetitive to each other to form a unified and non-redundant set of business cognitive rules.

[0032] The processed business cognitive rules are then uniformly stored in the business knowledge base according to the predefined data table structure and index design, thereby providing a consistent and accurate rule basis for subsequent analysis of business logic differences based on the business knowledge base.

[0033] Among them, the business knowledge base adopts a predefined data table structure, including a business cognitive rule table, a rule association table, and a rule version management table. Each table contains a unique identifier, rule content, applicable scope, and update time. The index design adopts a multi-level index structure based on the rule identifier, business area, and rule version to improve query and management efficiency.

[0034] S2 includes: Obtain structured business cognitive rules reflecting existing business logic, including data fields, data processing flows, and data verification rules, from the business knowledge base.

[0035] First, obtain structured business cognitive rules reflecting existing business logic from the business knowledge base pre-constructed within the enterprise. The business knowledge base stores structured business cognitive rule records generated through S1, and each record contains field information such as "rule identifier", "rule name", "rule description", "applicable business area", and "rule parameters".

[0036] To ensure the accuracy of the acquisition process, all information related to the existing business logic is extracted from the business knowledge base using predefined retrieval conditions and data matching rules. Specifically, the system sets conditions such as "data field identifier", "processing flow version number", "verification rule code", and other predefined retrieval items to load all target records into memory for subsequent comparison. Here, the "data field" refers to the data items defined in the existing business logic, including descriptive information such as data type, format, length, and data meaning; the "data processing flow" refers to the complete process of data transmission, conversion, processing, etc. in each business process; the "data verification rule" is the rule standard, conditions, and parameters used for verifying the accuracy of data.

[0037] Extract the business rule information in the newly generated business logic according to the same data fields, data processing flows, and data verification rules as the existing business logic to form a new business logic rule set.

[0038] The newly generated business logic comes from the new business rules introduced by the system according to market demands or business adjustments during the enterprise operation process. Its content includes data field definitions, data processing flow descriptions, and data verification rule settings.

[0039] To ensure that the extracted content is consistent with the existing business logic, the same fields, processes, and verification rule standards as those in the business knowledge base are used to match and extract the business rules involved in the new business logic. During the extraction process, the system converts the data items, processing steps, and verification standards described in the new business logic into structured rule records according to the predefined mapping rules and forms a "new business logic rule set". Here, each record in the new business logic rule set should have the same field structure as the existing business cognitive rules stored in the business knowledge base for subsequent comparison.

[0040] Compare each item in the new business logic rule set with the business cognitive rules corresponding to the existing business logic stored in the aforementioned business knowledge base: Analyze the specific differences in data type, data format, data length, and data meaning of each data field; compare each link of the data processing flow item by item to identify the specific differences in the data input, data conversion, and data output processes of each link; compare each item of the data verification rule to identify the specific differences in each verification standard, verification condition, and rule parameter.

[0041] The comparison process adopts a method of item-by-item comparison, that is, for each new business logic rule record, its data fields, data processing flow, and data verification rules are compared in turn with the content described in the corresponding existing business logic rule record. For data fields, the content compared by the system includes data type (such as character type, numeric type, etc.), data format (such as date format, currency format, etc.), data length (referring to the maximum number of characters or number of digits allowed for a data item), and data meaning (such as the role and meaning of each field in the business process); for the data processing flow, the system compares the specific descriptions of the operation steps, order, and execution conditions in each link during data input, data conversion, and data output; for the data verification rules, the verification standards, verification conditions, and rule parameters adopted by each rule are compared.

[0042] To quantitatively describe each difference, for example, the following formula is used to represent the overall difference value of data fields: ; where represents the overall difference value of data fields, which is used to quantify the data field difference degree between the new business logic and the existing business logic; is the numerical representation of each data field in the existing business logic, including but not limited to data type, data format, data length, and data meaning, etc.; is the numerical representation of the corresponding data field in the new business logic, and the numericalization rule is consistent with the numerical representation of each data field in the existing business logic; and respectively represent the numerical representations of the data types in the existing business logic and the new business logic. For example, different data types (such as strings, integers, floating-point numbers, etc.) are encoded using integers; and respectively represent the numerical representations of the data formats in the existing business logic and the new business logic. For example, date format, currency format, etc. can be represented using predefined encodings; and respectively represent the numerical representations of the data lengths in the existing business logic and the new business logic. The specific numerical values correspond to the maximum number of characters or number of digits allowed for a data item; and respectively represent the numerical representations of the data meanings in the existing business logic and the new business logic. The encoding method in the predefined dictionary is adopted to make the business meanings of different fields comparable; represents the total number of data fields compared between the new business logic and the existing business logic.

[0043] The above formula is only an example. In actual implementation, various information can be converted into standardized values according to predefined rules for difference calculation. After adopting this formula, if the overall difference value of the calculation result data field is greater than its corresponding preset threshold, it is determined that there is a large difference in the corresponding data field. Similarly, for the comparison of data processing processes and data verification rules, corresponding quantitative calculation methods can also be adopted. However, in this embodiment, a step-by-step comparison and item-by-item recording method is mainly used to manually or automatically identify the specific operation descriptions of each link and mark the differences existing therein.

[0044] Record the specific differences in accordance with the predefined format to form the difference analysis result between the new business logic and the existing business logic.

[0045] During the comparative analysis process, record all identified specific difference information in accordance with the predefined comparison format. Specifically, this predefined format includes difference categories (data fields, data processing processes, or data verification rules), difference items (such as data types, data formats, verification conditions, etc.), existing rule values, new rule values, and descriptions of the degree of difference. Each record is attached with a unique identifier for subsequent tracking and statistics. Here, each difference between the new business logic and the existing business logic should be recorded according to the same standard to ensure the continuity and consistency of the data. After sorting out all the difference record information, it constitutes the "difference analysis result", which serves as an important basis for subsequent data quality assessment and version management.

[0046] During the process of difference analysis, it is necessary to ensure the continuity of data transfer between sub-steps. For example, first, the structured business cognitive rules extracted from the business knowledge base maintain consistent field information when the new business logic is extracted; second, each new business logic rule in the comparison process clearly corresponds to the specific record of the existing business logic in the business knowledge base, ensuring the consistency of the comparison basis; finally, the obtained difference data is stored in a unified record format for subsequent system calls and data queries.

[0047] S3 includes: Perform mapping configuration on the business indicator data in the historical data. The mapping configuration includes: a. Compare item by item the data reflecting business indicators in the historical data with the business rule information involved in the new business logic according to the predefined mapping standard; b. Determine the corresponding relationship between each business indicator data in the historical data and the corresponding business rules in the new business logic; c. Store the corresponding relationship in the form of mapping records.

[0048] Among them, first, data records reflecting business indicators are extracted from the historical data storage system. These data records contain various data field information corresponding to each business indicator, such as field name, data type, data format, data length, and business meaning. At the same time, structured business cognitive rules of the existing business logic stored in the S2 step are obtained from the business knowledge base. The rule record details the predefined standards of data fields and the business rule information involved in the new business logic.

[0049] To ensure data consistency, the system preliminarily sorts out the extracted historical data fields and the business rule information in the new business logic according to the predefined mapping standards. The mapping standards clearly stipulate the matching requirements of each data field in terms of name, data type, data format, data length, and business meaning.

[0050] Specifically, the field name, data type, data format, data length, and data meaning of the historical data fields are compared in turn according to the predefined mapping standards. For example, if the field name of a certain business indicator in the historical data is "transaction amount", and the data type is numeric, the data format is currency format, the data length is 12 digits, and the business meaning is "record the transaction amount", then the corresponding business rule information in the new business logic should also include a field named "transaction amount" or a synonymous description, and have a similar numeric type, currency format, 12-digit length, and similar business meaning. Through item-by-item comparison, the system automatically determines the corresponding relationship between each business indicator field in the historical data and the corresponding business rule in the new business logic, thus forming a mapping correspondence. This correspondence not only includes the one-to-one matching situation between fields, but also records the subtle differences found during the matching process for reference during subsequent data quality assessment.

[0051] The above mapping correspondence is stored in the form of a mapping record. The mapping record details the matching information between each pair of corresponding historical data fields and new business logic fields, including field name, data type, data format, data length, business meaning, and matching status. The mapping record is sorted according to the predefined data format and uniformly stored in the mapping configuration database in the system for subsequent data processing and query. The mapping configuration database adopts a predefined data table structure and index design to ensure the efficient retrieval and management of mapping records.

[0052] Implement version management for the data field definitions of the data related to business indicators in the historical data. The data field definitions include the name, data type, data format, data length, and data meaning of the data fields. The version management records, identifies, and subsequently updates each data field definition according to the predefined version control rules.

[0053] Record the original data field definitions in historical data as the initial version according to predefined version control rules, e.g., labeled as version V1. The predefined version control rules clearly state that when the business logic changes or the relevant standards for data fields are updated, the system should generate new version records based on the changes. Specifically, when it is detected that the description of a certain data field in the new business logic differs from the existing description, the system will retain the original version record in the historical data and update the data field definition according to the new business logic to generate a new version, e.g., labeled as version V2. All versioned records include the version number, the version generation time, and an explanation of the version update reason, and are stored in a versioned management database for subsequent data traceability and version comparison. This versioned management process ensures that historical data can accurately trace the corresponding data field definition versions during subsequent data processing and consistency evaluation, thus guaranteeing the consistency and compatibility of data processing.

[0054] Implement versioned management for the calculation methods of business metric data in historical data. The calculation methods include the formulas for business metric calculations, calculation parameters, and data processing rules. Versioned management records and manages different versions of calculation methods according to predefined version identifiers and version update rules.

[0055] The calculation methods mainly include the formulas for business metric calculations, calculation parameters, and data processing rules. Initially, the system records the business metric calculation methods adopted in historical data to form an initial version record, which details the calculation formula (e.g., each operator, operation order, and each data item participating in the calculation), calculation parameters (such as proportionality coefficients, constant values, etc.), and data processing rules (e.g., data preprocessing, outlier handling methods, etc.). When new business logic introduces changes that result in adjustments to the business metric calculation formula or related parameters, the system, according to the predefined version update rules, records the changed content as a new version of the calculation method, generates a new version record, and assigns a unique version identifier to this record. The predefined version update rules clearly state that each change needs to retain the complete record of the previous version and record the comparison between the new version and the previous version to ensure that the historical data can trace back to the original calculation logic when applying the new calculation method. All calculation method version records are stored in a dedicated versioned management system using a unified record format to ensure that subsequent data consistency evaluation and data processing operations can be carried out based on the correct calculation version.

[0056] S4 includes: Obtain the historical business metric data after mapping configuration and the business metric data set generated according to the new business logic, and perform preliminary merging processing on the two.

[0057] Obtain the historical business metric data after mapping configuration from the aforementioned step S3. This data records the detailed definitions and related mapping relationships of each business metric in the historical data under the existing business logic. At the same time, a dataset of business metrics generated according to the new business logic has also been formed, reflecting the calculation results of each business metric under the new business logic. For the convenience of subsequent comparative analysis, first, perform preliminary merging processing on the above two parts of data. The merging processing includes unifying the data format, aligning the field names, and handling data missing items, so that the merged data set not only contains the original attributes of the historical data but also incorporates the rule information of each item in the new business logic, forming a comprehensive data set as the input data for subsequent local analysis.

[0058] Perform neighborhood partitioning on the data after the merging process according to the local neighborhood analysis method. The neighborhood partitioning groups different records according to the data distribution characteristics.

[0059] After completing the preliminary data merging, use the local neighborhood analysis method to perform neighborhood partitioning on the merged data. Specifically, according to the distribution characteristics of the data records on the key fields (such as business metric values, timestamps, category identifiers, etc.), use the predefined neighborhood partitioning rules to group the data set. The partitioning rules can group the records with similar distribution characteristics into the same neighborhood according to the statistical distribution of the data records, such as density or similarity measure. Each neighborhood is defined here as a local data distribution area, and each neighborhood is assigned a unique neighborhood identifier. Through this neighborhood partitioning process, ensure that when calculating the structural entropy in each local area subsequently, the data has a clear attribution range and statistical basis.

[0060] Within the local data distribution area corresponding to each neighborhood, calculate the entropy value of the data based on the pre-set structural entropy calculation method. The structural entropy calculation method measures the distribution characteristics of each record in the same neighborhood on multiple field dimensions.

[0061] After completing the neighborhood partitioning, calculate the structural entropy of the data within each neighborhood to quantitatively describe the consistency of the distribution of the old and new business metric data on each key field dimension within the same local data distribution area. Adopt a measurement method based on information entropy, that is, calculate the Shannon entropy of the old and new data distributions within each neighborhood, and use their absolute difference as the structural entropy difference index.

[0062] Within each local neighborhood, first classify and statistically analyze the historical business metric data after mapping configuration and the data generated according to the new business logic respectively to obtain their respective probability distributions. Specifically, for a certain local neighborhood, denote the probabilities of each category of historical data as ; among them, represents the Shannon entropy of the historical business metric data within the current neighborhood, describing the uncertainty of this data distribution; represents the occurrence probability of the th category in the historical business metric data, being the total number of categories counted within this neighborhood; similarly, the Shannon entropy of the new business logic data is denoted as ; where represents the Shannon entropy of the data generated according to the new business logic within the current neighborhood; represents the occurrence probability of the th category in the new business logic data. To ensure the consistency of comparison, the category division here is kept consistent according to predefined rules, and both take values from the same set.

[0063] Subsequently, the local structure entropy difference is defined as ; where serves as a quantitative indicator of the data distribution consistency within this local neighborhood, used to reflect the degree of difference in the statistical distributions between the historical data and the new business logic data.

[0064] is the local structure entropy difference, representing the structural entropy difference between the two sets of data within the current neighborhood. The smaller the value, the more consistent the distribution of the old and new data, and vice versa, the lower the consistency.

[0065] By comparing the entropy value differences between the historical business metric data and the new business logic business metric data in each neighborhood, the entropy value difference results are used as the basis for local consistency determination.

[0066] After calculating the local structure entropy difference in each neighborhood, the data consistency in each neighborhood is judged according to the preset consistency determination criteria.

[0067] Specifically, the preset consistency determination criteria include one or more thresholds. When the local structure entropy difference in a certain local neighborhood is less than or equal to this threshold, it is determined that the data consistency evaluation result is normal, that is, the distributions of the old and new business metric data in this neighborhood are consistent; otherwise, if the local structure entropy difference exceeds the threshold, it is considered that there are significant distribution differences in this neighborhood and the consistency requirement is not met.

[0068] Among them, the preset consistency determination criteria include a threshold range set based on the structure entropy difference. When the calculated structure entropy difference value in the local neighborhood is lower than or equal to this threshold, it is determined that the distributions of the old and new business metric data in this neighborhood are consistent, otherwise it is determined that there are significant differences in the data, and the corresponding determination results are recorded.

[0069] This determination result is output in the form of a record, and detailed entropy value data and comparison results are attached within each neighborhood to form a local consistency determination report.

[0070] The data consistency evaluation result is formed based on the entropy value difference results of each neighborhood, and the data consistency evaluation result is associated with the previous mapping configuration information and stored.

[0071] The judgment results in all local neighborhoods are aggregated to form an overall data consistency assessment result. The aggregation process includes counting the numerical values ​​of the local structural entropy differences in each neighborhood, calculating the overall consistency score, and classifying and integrating the local judgment results. The overall data consistency assessment result formed includes the entropy value differences, judgment conclusions, and corresponding neighborhood identifiers of each local data distribution area. To facilitate subsequent data tracing and application, the aggregated overall data consistency assessment results are associated and stored with the mapping records previously generated in the S3 step. The storage format follows the predefined data table structure and index design requirements to ensure that the data consistency assessment results can be quickly retrieved and called in the data center.

[0072] The local neighborhood analysis method adopted in this embodiment realizes quantitative evaluation of local data distribution differences by calculating structural entropy in the local distribution area of ​​the historical business indicator data after mapping configuration and the data set generated according to the new business logic. It can accurately capture the abnormal changes in local data caused by the dynamic evolution of business logic, divide the data set into neighborhoods according to predefined standards, and calculate the entropy value difference based on local statistical characteristics, thereby generating data consistency evaluation results. Combined with the dynamic adjustment of business rules, the mapping configuration, versioning management and local data distribution analysis are organically connected, which not only avoids the problem of global averages covering up local fluctuations, but also realizes real-time, accurate and regional management of data quality monitoring.

[0073] S5 includes: Receive the generated data consistency assessment results and filter out data with normal data consistency assessment results and that meet business cognition rules.

[0074] Among them, the evaluation result records the difference in structural entropy calculated by comparing the historical business indicator data after the mapping configuration with the business indicator data set generated according to the new business logic in the local neighborhood, and makes consistency judgments on each local data area according to the preset consistency judgment criteria. According to the predefined business cognition rules, the evaluation results are screened to eliminate data with abnormal consistency evaluation results or data that does not meet the preset consistency requirements. The screening process strictly follows the requirements of the business cognition rules regarding data field definitions, data processing procedures, and data verification rules to ensure that only records that match the old and new business logics and have consistent data distribution are retained. The screened data not only represents the qualified part of the data consistency evaluation results, but also provides an accurate basis for subsequent data writing and data traceability.

[0075] Write the screened data into the data middle platform. The writing process includes recording the data according to the predefined field definitions and data formats, and attaching a unique identifier to each piece of data.

[0076] Write the qualified data records screened out into the data middle platform according to the predefined field definitions and data formats. The writing process includes format conversion, field alignment, and data completion for each data record to ensure that the written data matches the predefined database table structure in the data middle platform.

[0077] To facilitate subsequent data management and query, a unique identifier is attached to each written data record. This unique identifier is generated by the system according to preset rules and can ensure no duplicates in the entire data middle platform. The writing operation strictly follows the predefined record specifications, making the data content, data field definitions consistent with the previous mapping configuration and consistency evaluation results, ensuring data integrity and consistency.

[0078] Obtain the traceability information of the screened data. The traceability information includes the historical version numbers and corresponding processing descriptions recorded during the acquisition, processing, and evaluation of the data.

[0079] Obtain the data traceability information for the screened data records. The data traceability information includes all historical records generated during the data collection, preprocessing, mapping configuration, difference analysis, local neighborhood analysis, and data consistency evaluation processes, specifically including the historical version numbers, timestamps, processing step descriptions, and corresponding processing descriptions recorded in each link.

[0080] This traceability information is generated through the system's automatic recording and log management functions and stored in a dedicated traceability database in a structured record manner. The traceability information ensures that each piece of written data can be traced back to the entire process of its generation and processing, providing a detailed basis for subsequent data auditing and version updates.

[0081] Associate and save the data traceability information with the business logic version information. The business logic version information includes the business logic release time, the evolution process of business rules, and the corresponding version identifier.

[0082] Associate and save the data traceability information with the business logic version information. The business logic version information includes the release time of the business logic, the evolution process of business rules, and the corresponding version identifier. The system matches the historical version numbers in the data traceability information with the version identifiers in the business logic version information according to the predefined association rules and records the matching results in the data middle platform.

[0083] The associated preservation process adopts a structured data recording method. The recorded content includes the unique identifier of the data record, the corresponding traceability information, the business logic version identifier, and the business logic release time, etc., to ensure that each piece of written data is closely associated with the corresponding business logic version information in the data middle platform, facilitating subsequent data traceability and version comparison during business rule adjustment.

[0084] Based on the data traceability record associated with the business logic version information, perform complete retention and traceable management on the final data written into the data middle platform.

[0085] This management process includes regular backup, data integrity verification, and log archiving to ensure that data can be traced and verified at each stage of its life cycle. The system has established a unified query interface, which can quickly retrieve the historical processing records of data according to the data unique identifier, traceability information, or business logic version information.

[0086] Through complete retention and traceable management, the whole process of data after writing is supervised, ensuring that data can provide sufficient historical basis during future business adjustment, meeting the requirements of data consistency and version management.

[0087] Embodiment 2: The difference between Embodiment 2 and Embodiment 1 of the present invention is that this embodiment introduces a data quality improvement system driven by business cognition in the data middle platform.

[0088] Figure 2 The structural schematic diagram of a data quality improvement system driven by business cognition in the data middle platform of the present invention is given. A data quality improvement system driven by business cognition in the data middle platform includes a business rule integration module, a logical difference analysis module, a data mapping version module, a local entropy value evaluation module, and a data traceability archiving module.

[0089] Business rule integration module: Obtain business cognition rules in multiple business domains and integrate the business cognition rules into a unified business knowledge base.

[0090] Logical difference analysis module: Based on the business knowledge base, perform difference analysis on the newly generated business logic, and identify the specific differences between the new business logic and the existing business logic in data fields, data processing processes, and data verification rules.

[0091] Data mapping version module: Perform mapping configuration on the business metric data in the historical data according to the difference analysis result, and implement version management on the data field definition and calculation method.

[0092] Local entropy value evaluation module: Based on local neighborhood analysis, it compares the mapped and configured service metric data in historical data with the service metric data set generated according to the new service logic, calculates the structural entropy within the local data distribution area divided based on neighborhoods, and generates a data consistency evaluation result according to the calculation result.

[0093] Data traceability and archiving module: Writes the data with normal data consistency evaluation results that meet the service cognition rules into the data middle platform, and associates and saves the data traceability information with the service logic version information.

[0094] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to get a formula closest to the actual situation. The preset parameters and threshold selection in the formulas are set by those skilled in the art according to the actual situation.

[0095] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that the computer can access, or a data storage device such as a server or data center that contains one or more sets of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.

[0096] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, devices, and modules described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0097] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the devices or modules can be in electrical, mechanical, or other forms.

[0098] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules. They can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0099] In addition, in each embodiment of the present application, the functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module.

[0100] If the above functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0101] As described above, the above are only the specific implementation manners of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0102] Finally, the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A data quality improvement method driven by business cognition of a data middle platform, characterized in that: The steps include: S1: Obtain business cognition rules in multiple business fields and integrate the business cognition rules into a unified business knowledge base; S2: Perform difference analysis on the newly generated business logic based on the business knowledge base to identify the specific differences between the new business logic and the existing business logic in data fields, data processing flow and data verification rules; S3: Map and configure the business indicator data in the historical data based on the difference analysis results, and implement version management for the data field definition and calculation method; S4: Based on local neighborhood analysis, the business indicator data after mapping and configuration in the historical data is compared with the business indicator data set generated according to the new business logic, the structural entropy is calculated in the local data distribution area based on the neighborhood division, and the data consistency evaluation result is generated according to the calculation result; S5: Write the data with normal data consistency assessment results and that meet the business cognition rules into the data center, and associate the data traceability information with the business logic version information for storage.

2. According to claim 1, a data quality improvement method driven by business cognition of a data middle platform is characterized in that: S1 includes: Extract original business information involving business processes, data indicator definitions, data operation rules and data verification standards from each business system within the enterprise; Preprocessing the original business information to form structured business cognition rules, which include rule identification, rule name, rule description, applicable business field and corresponding rule parameters; The structured business cognition rules are classified, associated and merged, and the processed business cognition rules are uniformly stored in a business knowledge base with a predefined data table structure and index design.

3. According to claim 1, a data quality improvement method driven by business cognition of a data middle platform is characterized in that: S2 includes: Obtain structured business cognition rules that reflect the data fields, data processing procedures, and data verification rules of existing business logic from the business knowledge base; Extract the business rule information in the newly generated business logic according to the same data fields, data processing flow and data verification rules as the existing business logic to form a new business logic rule set; Compare the new business logic rule set with the business cognition rules corresponding to the existing business logic stored in the aforementioned business knowledge base item by item: Analyze the specific differences in data type, data format, data length and data meaning of each data field; compare the data processing process step by step to identify the specific differences in data input, data conversion and data output in each step; compare the data verification rules item by item to identify the specific differences in each verification standard, verification conditions and rule parameters; The specific differences are recorded in a predefined format to form the difference analysis results between the new business logic and the existing business logic.

4. According to claim 1, a data quality improvement method driven by business cognition of a data middle platform is characterized in that: S3 includes: Map and configure the business indicator data in the historical data; Implement version management for the data field definitions contained in the historical data involving business indicators. The data field definitions include the name, data type, data format, data length and data meaning of the data field. Version management records, identifies and subsequently updates each data field definition according to predefined version control rules. Version management is implemented for the calculation method of business indicator data in historical data. The calculation method includes the formula, calculation parameters and data processing rules for business indicator calculation. Version management records and manages different versions of calculation methods based on predefined version identification and version update rules.

5. According to claim 4, a data quality improvement method driven by business cognition of a data middle platform is characterized in that: The mapping configuration includes: a. Compare the data reflecting business indicators in historical data with the business rule information involved in the new business logic item by item according to the predefined mapping standards; b. Determine the correspondence between each business indicator data in the historical data and the corresponding business rules in the new business logic; c. Store the corresponding relationship in the form of mapping records.

6. According to claim 1, a data quality improvement method driven by business cognition of a data middle platform is characterized in that: S4 includes: Obtain the historical business indicator data after mapping configuration and the business indicator data set generated according to the new business logic, and perform preliminary merging processing on the two; The merged data is divided into neighborhoods according to the local neighborhood analysis method, and the neighborhood division groups different records according to the data distribution characteristics; In the local data distribution area corresponding to each neighborhood, the entropy value of the data is calculated based on the pre-set structural entropy calculation method. The structural entropy calculation method measures the distribution characteristics of each record in the same neighborhood in multiple field dimensions; Compare the entropy value differences between historical business indicator data and new business logic business indicator data in each neighborhood, and use the entropy value difference results as the basis for local consistency judgment; The data consistency evaluation result is formed based on the entropy value difference results of each neighborhood, and the data consistency evaluation result is associated with the previous mapping configuration information and stored.

7. According to claim 6, a data quality improvement method driven by business cognition of a data middle platform is characterized in that: The entropy value of the data is calculated based on the pre-set structural entropy calculation method, specifically: In each local neighborhood, the historical business indicator data after mapping configuration and the data generated according to the new business logic are classified and counted to obtain their respective probability distributions; for a certain local neighborhood, the probability of each category of historical data is recorded as ;in, Represents the Shannon entropy of historical business indicator data in the current neighborhood, Indicates the historical business indicator data The probability of occurrence of a category, is the total number of categories counted in the neighborhood; the Shannon entropy of the new business logic data is recorded as ;in, represents the Shannon entropy of the data generated according to the new business logic in the current neighborhood, Indicates the first The probability of occurrence of a category.

8. According to claim 1, a data quality improvement method driven by business cognition of a data middle platform is characterized in that S5 include: Receive the generated data consistency assessment results, and filter out data with normal data consistency assessment results and that meet business cognition rules; The filtered data is written into the data center. The writing process includes recording the data according to the pre-set field definition and data format, and adding a unique identifier to each piece of data. Obtain the traceability information of the filtered data, which includes the historical version number and corresponding processing description recorded during the acquisition, processing and evaluation of the data; The data traceability information is associated with the business logic version information and saved. The business logic version information includes the business logic release time, business rule evolution history and corresponding version identifier; Based on the data traceability records saved in association with the business logic version information, the final data written into the data center is fully retained and traceable.

9. A data quality improvement system driven by business cognition of a data middle platform, used to implement a data quality improvement method driven by business cognition of a data middle platform as described in any one of claims 1-8, characterized in that: It includes business rule integration module, logic difference analysis module, data mapping version module, local entropy value evaluation module and data traceability archiving module; Business rule integration module: obtains business cognition rules in multiple business fields and integrates the business cognition rules into a unified business knowledge base; Logical difference analysis module: performs difference analysis on the newly generated business logic based on the business knowledge base, and identifies the specific differences between the new business logic and the existing business logic in data fields, data processing flow and data verification rules; Data mapping version module: maps and configures the business indicator data in the historical data according to the difference analysis results, and implements version management for the data field definition and calculation method; Local entropy value evaluation module: Based on local neighborhood analysis, the business indicator data that has been mapped and configured in the historical data is compared with the business indicator data set generated according to the new business logic, the structural entropy is calculated in the local data distribution area based on the neighborhood division, and the data consistency evaluation result is generated according to the calculation result; Data traceability and archiving module: writes data with normal data consistency assessment results and that meets business cognition rules into the data center, and associates and saves data traceability information with business logic version information.

Citation Information

Patent Citations

  • Streaming query semantic map adaptive enhancement method and system based on cognitive calculation

    CN119669298A

  • Data traceability analysis system and method based on big data

    CN119722337A

  • A rule-based automated data governance system and method

    CN119782715A

  • Analytical system for discovery and generation of rules to predict and detect anomalies in data and financial fraud

    US20070027674A1

  • Cognitive rule engine

    US20190065972A1

Cited By

  • Server content updating and rollback method based on static technology

    CN120892234A

  • A server content updating and rollback method based on static technology

    CN120892234B