Data table cleaning method and device, storage medium and electronic equipment

By dynamically assessing trust levels and selecting calibration values ​​in multi-source data tables, the problem of difficulty in assessing the trustworthiness of data sources in existing technologies is solved, achieving efficient and accurate data cleaning results and ensuring high-quality and timely data.

CN121745972APending Publication Date: 2026-03-27CHINA BOND FINANCIAL VALUATION CENT CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-02
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing technologies cannot objectively determine the credibility of data sources when processing multi-source data, leading to the misselection of low-quality data as calibration values, which affects the cleaning effect and makes it difficult to meet the needs of high-timeliness and high-quality data.

Method used

By acquiring multi-source data tables and using a data backtracking chain to trace the original source of each data table, we can dynamically assess its trust level and select the highest level of trust data as the calibration value to replace dirty data. By adopting a trust-first calibration strategy, we can form a database with high accuracy and timeliness.

Benefits of technology

It significantly improves the efficiency and accuracy of data processing, reduces the waste of computing resources, ensures the high quality and timeliness of data, and realizes a leap from passive proofreading to active calibration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121745972A_ABST
    Figure CN121745972A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data table cleaning method and device, a storage medium and electronic device.The method comprises the steps that in response to dirty data existing in the same target resource subject in a multi-source data table of a specified object, the trust degree of each source data table is determined according to a data backtracking chain of each source data table; the data backtracking chain of each source data table refers to a flow path of each source data table; the data backtracking chain of each source data table is used for reversely tracking the original source of each source data table; and according to the credibility of each source data table, selecting a calibration value from the same target resource subject in the multi-source data table, and replacing dirty data of the same target resource subject in the multi-source data table with the calibration value. According to the method and the device, the technical problem that the cleaning effect is poor due to the fact that the reliability of the multi-source data is difficult to evaluate and low-quality data is mistakenly selected for calibration is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of text processing, in particular, relate to a data table cleaning method and device, a storage medium and an electronic device. BACKGROUND

[0002] At present, with the development of the information age, commercial reports are increasingly frequently applied in enterprise operation analysis, market prediction and other fields, and the requirements for data quality (including accuracy, timeliness and completeness) have reached an unprecedented height. At present, enterprises will regularly issue annual reports, audit reports or business profiles containing multi-dimensional information such as business status and market performance, which constitute rich information resources. At present, there are many professional information providers in the market, which are committed to organizing and providing digital abstracts of such documents for analysis by various commercial institutions. However, although these services simplify the information acquisition process to a certain extent, the data often has problems such as insufficient accuracy and delayed error correction, which is difficult to meet the needs of internal business for high timeliness and high quality data.

[0003] In related technologies, when dealing with inconsistency problems of multi-source data, simple rules (such as the latest data first, and the specific system first) are usually used, which cannot objectively judge which data source is more reliable. After identifying dirty data, there is often a lack of scientific calibration value selection mechanism, which may lead to the selection of low-quality data as calibration value, affecting the cleaning effect. SUMMARY

[0004] Embodiments of the present application provide a data table cleaning method, device, storage medium and electronic device to at least solve the technical problem that in related technologies, multi-source data is difficult to evaluate in terms of reliability, low-quality data is selected by mistake for calibration, and the cleaning effect is poor.

[0005] According to an aspect of an embodiment of the present application, a data table cleaning method is provided, comprising: obtaining a multi-source data table of a specified object; the multi-source data table refers to a resource data table collected from different data sources about the specified object; each source data table in the multi-source data table includes resource data of different resource subjects; in response to the existence of dirty data in the same target resource subject in the multi-source data table of the specified object, determining the trust degree of each source data table according to the data traceability chain of each source data table; the data traceability chain of each source data table refers to the flow path of each source data table; the data traceability chain of each source data table is used to trace the original source of each source data table in reverse; according to the trust degree of each source data table, selecting a calibration value from the same target resource subject in the multi-source data table, and replacing the dirty data of the same target resource subject in the multi-source data table with the calibration value.

[0006] According to another aspect of the embodiments of the present application, a data table cleaning device is also provided, comprising: an acquisition module configured to acquire a multi-source data table of a specified object; the multi-source data table refers to a resource data table collected from different data sources about the specified object; each source data table in the multi-source data table includes resource data of different resource subjects; a trust degree confirmation module configured to, in response to the existence of dirty data in a same target resource subject in the multi-source data table of the specified object, determine a trust degree of each source data table according to a data traceability chain of the each source data table; the data traceability chain of the each source data table refers to a flow path of the each source data table; the data traceability chain of the each source data table is used to trace back the original source of the each source data table; and a data cleaning module configured to, according to the trust degree of the each source data table, select a calibration value from the same target resource subject in the multi-source data table, and replace the dirty data of the same target resource subject in the multi-source data table with the calibration value.

[0007] According to still another aspect of the embodiments of the present application, a computer readable storage medium is also provided, and the computer readable storage medium stores a computer program. The computer program is configured to be executed by a processor to perform the steps in any of the method embodiments.

[0008] According to still another aspect of the embodiments of the present application, a computer program product or computer program is provided, and the computer program product or computer program includes computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to enable the computer device to perform the steps in any of the method embodiments.

[0009] According to still another aspect of the embodiments of the present application, an electronic device is also provided, and the electronic device includes a memory and a processor. The memory stores a computer program, and the processor is configured to execute the computer program to perform the steps in any of the method embodiments.

[0010] By the present application, resource data from different data sources is uniformly standardized to form a multi-source data table, solving the problem of difficult horizontal comparison caused by inconsistent fields and chaotic formats among data sources in the related art; when dirty data of a same target resource subject (such as "net profit") in the multi-source data table of the specified object is identified, the trust degree of each source data table is dynamically evaluated according to the data backtracking chain, i.e. the flow path and processing history, that is, the original source of the data is tracked in reverse through the data backtracking chain, the reliability of the source is evaluated, the processing steps of the data in the flow process are considered, the more the flow nodes, the greater the risk of data pollution, the analysis result of the data backtracking chain is converted into a quantifiable trust degree index, solving the problem of difficult evaluation of reliability, which can more accurately reflect the actual credibility of the source data table, solving the problem that the static rating cannot accurately reflect the data timeliness and trust change in the processing process; the highest level of trust degree data of the same target resource subject in the multi-source data table is selected as the calibration value for replacing the dirty data of the same target resource subject in the multi-source data table, through the closed-loop mechanism of "verification driving, accurate analysis and autonomous calibration", not only the dirty data can be automatically found, but also the authoritative source can be traced back, through the triggered cleaning mechanism, only the data identified as dirty data is cleaned, unnecessary calculation resource consumption is avoided, and the processing efficiency is greatly improved; through the calibration value selection strategy of trust degree priority, the credibility of the calibration value is ensured, the calibration error caused by improper selection of data source is avoided, the correct value is used for calibration, thereby forming a database with high accuracy and high timeliness, solving the technical problem of poor cleaning effect caused by difficult evaluation of reliability, misselection of low-quality data calibration in the related art, significantly improving the efficiency and accuracy of data processing, reducing the waste of calculation resources, ensuring the high quality and timeliness of the data, realizing the three great leaps from "passive proofreading" to "active calibration", from "static weighting" to "scene adaptation", and from "simple replacement" to "accurate tracing", providing strong support for resource data analysis and management. BRIEF DESCRIPTION OF DRAWINGS

[0011] Figure 1 is an application scenario of a data table cleaning method according to an embodiment of the present application;

[0012] Figure 2 is a flowchart of an optional data table cleaning method according to an embodiment of the present application;

[0013] Figure 3 is a schematic diagram of another optional data table cleaning method according to an embodiment of the present application;

[0014] Figure 4 is a structural block diagram of an optional data table cleaning device according to an embodiment of the present application;

[0015] Figure 5 This is a computer system architecture block diagram of an optional electronic device according to an embodiment of this application. Detailed Implementation

[0016] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0017] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0018] According to one aspect of the embodiments of this application, a data table cleaning method is provided. Optionally, in this embodiment, the above-described data table cleaning method may be applied to, but is not limited to, [examples of data cleaning methods]. Figure 1 The hardware environment shown includes terminal device 102 and server 104. Server 104 can be connected to terminal device 102 via a network and can be used to provide services (e.g., application services, etc.) to terminal device 102 or clients installed on terminal device 102. A database can be set up on server 104 or independently of server 104 to provide data storage services for server 104.

[0019] The aforementioned network may include, but is not limited to, at least one of the following: wired network and wireless network. The aforementioned wired network may include, but is not limited to, at least one of the following: wide area network (WAN), metropolitan area network (MAN), and local area network (LAN). The aforementioned wireless network may include, but is not limited to, at least one of the following: Wireless Fidelity (WIFI) and Bluetooth. Terminal device 102 may be, but is not limited to, a personal computer (PC), mobile phone, tablet computer, etc. Server 104 may be, but is not limited to, a cloud server, server cluster, or other server types.

[0020] The data table cleaning method of this application embodiment can be executed by server 104, terminal device 102, or jointly by server 104 and terminal device 102. Alternatively, the data table cleaning method of this application embodiment can be executed by a client installed on terminal device 102.

[0021] Taking the data table cleaning method in this embodiment as an example, which is executed by a computer device (such as terminal device 102 or server 104), Figure 2 This is a flowchart illustrating an optional data table cleaning method according to an embodiment of this application, such as... Figure 2 As shown, the process of this method may include the following steps:

[0022] Step S202: Obtain the multi-source data table for the specified object; the multi-source data table refers to the resource data table about the specified object collected from different data sources; each source data table in the multi-source data table includes resource data of different resource subjects.

[0023] Step S204: In response to the presence of dirty data in the same target resource subject in the multi-source data tables of the specified object, determine the trust level of each source data table based on the data backtracking chain of each source data table; the data backtracking chain of each source data table refers to the flow path of each source data table; the data backtracking chain of each source data table is used to trace the original source of each source data table in reverse.

[0024] Step S206: Based on the trust level of each source data table, select a calibration value from the same target resource subject in the multi-source data table, and use the calibration value to replace the dirty data of the same target resource subject in the multi-source data table.

[0025] The data table cleaning method in this embodiment can be applied to the field of data processing, particularly in scenarios that rely on multi-source heterogeneous data for decision analysis. For example, it can be applied to investment institutions that need to process large amounts of report data in risk assessment, market trend analysis, and compliance checks, ensuring the accuracy of the information used for analysis and improving the accuracy of risk management and investment decisions.

[0026] With the development of the information age, business reports are increasingly used in areas such as corporate operations analysis and market forecasting, placing unprecedented demands on data quality (including accuracy, timeliness, and completeness). Currently, companies regularly publish annual reports, audit reports, or business overviews containing multi-dimensional information such as operating conditions and market performance; these documents constitute a wealth of information resources. Several professional information providers exist in the market, dedicated to compiling and providing digital summaries of these documents for analysis by various business organizations. However, while these services simplify the information acquisition process to some extent, their data often suffers from insufficient accuracy and delayed error correction, making it difficult to meet the internal business needs for timely and high-quality data.

[0027] When dealing with inconsistencies in multi-source data, related technologies typically employ simple rules (such as prioritizing the latest data or prioritizing specific systems), which fail to objectively determine which data source is more reliable. After identifying dirty data, they often lack a scientific calibration value selection mechanism, which may lead to the selection of low-quality data as calibration values, thus affecting the cleaning effect.

[0028] To at least partially solve the above-mentioned technical problems, this embodiment provides a data table cleaning method. This method can intelligently, accurately, and efficiently integrate the "problem discovery capability" of multi-source verification with the "problem-solving capability" of complex document parsing, and can make dynamic decisions based on data credibility to achieve a systematic solution for autonomous closed-loop calibration of resource data.

[0029] The designated object refers to an entity that requires disclosure of resource data tables. For example, a designated object can be a company, an individual, or a third-party organization. A multi-source data table for a designated object refers to a collection of data records about the designated object gathered from multiple independent data providers. Each source data table in the multi-source data table refers to a single data record provided by a single data provider for the designated object. Each source data table contains resource data corresponding to different resource items for the designated object under different reporting periods and report types. A reporting period refers to a specified time period used for collecting, organizing, and displaying data. For example, a reporting period can be daily, weekly, monthly, quarterly, or annual. Report type refers to the data display format organized according to specific purposes and content. For example, report types include, but are not limited to, operational status tables, profit and loss statements, cash flow analysis tables, and annual reports. Resource items refer to the specific measurement items in the data. For example, resource items can be production costs, R&D expenditures, etc. The resource data corresponding to a resource item represents the specific numerical information of each resource item for the designated object within a certain period. For example, source data table A records the data records of resource item Z (corresponding resource subject) in the performance display table (corresponding report type) of entity E in the first quarter (corresponding reporting period).

[0030] In this embodiment, dirty data in the multi-source data tables is identified by comparing resource data under the same reporting period, report type, and resource category in the same multi-source data tables. Dirty data in the multi-source data tables refers to resource data inconsistencies that occur under the same reporting period, report type, and resource category in the same multi-source data tables.

[0031] In some embodiments, because different data providers describe or name the same resource subject differently, it is necessary to convert the multi-source data tables into the same format according to standardized mapping rules before performing multi-source data table comparison. Specifically, after obtaining the multi-source data tables of the specified object, the above method further includes:

[0032] By using a pre-built mapping knowledge base, the different source data tables in the multi-source data table are standardized to obtain a standardized multi-source data table. The mapping knowledge base is used to define the mapping relationship between the standard name and alias of the resource subject of a specified object, the report type to which the resource subject of the specified object belongs, and the logical relationship between the resource subjects of the specified object.

[0033] This involves defining a data model based on the mapping relationship between resource subjects and business logic, and then building a unified mapping knowledge base through this data model. The core of this data model is to create a "resource subject multi-source validation relationship mapping configuration table," which specifies the standard name, report type, and associated resource subjects (such as the logical relationship with "total profit") for each resource subject (e.g., operating revenue), as well as possible aliases or encoding mapping rules in different data source products.

[0034] The standard name of a resource subject is a standardized terminology defined in the mapping knowledge base to uniformly represent and identify resource subjects, aiming to eliminate confusion caused by using different names for the same subject in different source data tables. Aliases for resource subjects refer to non-standard names or codes that may exist for resource subjects in different source data tables. Through the mapping relationships in the mapping knowledge base, the system can identify and convert these aliases, achieving data consistency verification for the same subject.

[0035] Based on the "Resource Subject Multi-Source Validation Relationship Mapping Configuration Table" in the mapping knowledge base, multi-source data tables provided by data providers such as external data source A, external data source B, and external data source C are converted into a standard data stream of a unified internal format in real time. Key processing includes horizontal and vertical table conversion of fields, mapping of synonymous fields, and code value conversion.

[0036] In some embodiments, after converting multi-source data tables into the same format according to standardized mapping rules, the same reporting period, the same report type, and the same resource subject of a specified object are used as the smallest comparison unit. Before comparing the resource data within the smallest comparison unit in the multi-source data tables in parallel, it is necessary to parse the resource data within the same smallest comparison unit in each source data table from the multi-source data tables. Unlike the "end-to-end" full-text parsing model in related technologies (such as Qwen2.5-VL-7B), the embodiments of this application use a parsing engine with a "structure-recognition-relationship" (SRR) ternary architecture to parse each source data table. This parsing engine is not a simple OCR, but rather an SRR ternary architecture that integrates visual, semantic, and rule-based collaboration. Specifically, the parsing engine uses a lightweight object detection network to complete the ultra-fast structure detection of the entire page document in a very short time (e.g., within 20 milliseconds). The input of the object detection network is the source data table, and the output is an accurate bounding box containing all content elements and category labels.

[0037] By tracing and calibrating the source, the system eliminates its dependence on any external data source. In particular, through multimodal recognition and relationship reconstruction in the SRR architecture, namely lightweight structure detection based on meta-learning and dynamic anchor boxes, error-correcting OCR based on visual-semantic interaction, topological reasoning of table relationships based on graph neural networks (GNN), and semantic localization and verification for the financial domain, it achieves accurate and robust extraction of specific financial information from non-standardized and complexly formatted PDF reports. This ensures extremely high accuracy (over 99.99%) in extracting data from complex formats, enabling autonomous control over the quality of core financial data, reducing computational resource consumption by more than 70%, and increasing processing speed several times over.

[0038] After extracting the resource data within the smallest comparison unit from each source data table, the resource data within the smallest comparison unit in the multi-source data tables are compared in parallel. If the logical rules of the resource data within the smallest comparison unit in the multi-source data tables are consistent, it is considered reliable data and is directly entered into the "Resource Subject Verification Pass Information Table". If the logical rules of the resource data within the smallest comparison unit in the multi-source data tables are inconsistent or violate the preset logical rules, the resource data within the smallest comparison unit in the multi-source data tables is marked as dirty data to be calibrated. The key information of the dirty data (such as the specified object, reporting period, resource subject, and data source involved) is written into the "Resource Subject Self-Verification Information Table" and the data cleaning process is triggered.

[0039] In response to the presence of dirty data in the same target resource category across multiple source data tables for a specified object, during data cleaning, the trust level of each source data table is dynamically determined based on its data backtracking chain. The resource data of the target resource category in the source data table with the highest trust level is designated as the calibration value, and this calibration value is used to replace the resource data of the target resource category in the multiple source data tables. The cleaned multi-source data tables are then used for subsequent risk identification, compliance report generation, and other purposes.

[0040] The target resource account refers to the resource account containing dirty data in the multi-source data table. Dirty data refers to inconsistent values ​​or logical errors appearing under the same target resource account in multiple source data tables.

[0041] The data backlink chain for each source data table records the flow path of the source data table, from the final data table back to the most original data table, used to trace the exact source of the data. Each source data table's data backlink chain includes at least one data flow node, which refers to each processing or storage stage the data undergoes from its original generation point to its inclusion in the source data table. Data flow nodes include, but are not limited to, data collection points (such as enterprise resource information reports), data vendors, intermediate tables for data standardization mapping, manual or system review stages, and the final data storage location (such as a database or data warehouse). For example, the data backlink chain for source data table A records that source data table A was obtained according to the following flow path: Enterprise Annual Report (original generation point) → Data Vendor 1 → Data Vendor 2 → Source Data Table A. The data backlink chain for source data table A includes four data flow nodes, such as the enterprise annual report as one data flow node and the data vendor as another.

[0042] In this embodiment, the longer the data backtracking chain of each source data table, the lower the trust level of each source data table; conversely, the more data flow nodes in the data backtracking chain of each source data table, the lower the trust level of each source data table. This is because errors may be introduced each time data passes through a processing or transmission node, whether due to human error, technical mistakes, or differences in format conversion. These factors accumulate as the number of nodes increases, ultimately affecting the accuracy and integrity of the data. Secondly, a longer data backtracking chain often means a longer time interval between data generation and use, and this time delay can cause the data to lose its timeliness value. Finally, each extension of the data backtracking chain may introduce more third-party involvement. The data security measures, processing standards, and motivations of third parties may differ from those of the original data source, increasing the risk of a broken trust chain.

[0043] Optionally, assume that the contribution of each data flow node in the data backtracking chain of each source data table to the trust level of the original data has an exponential decay effect, that is, the trust level decreases exponentially with the increase of data flow nodes; define the initial trust level of each source data table as 100%, and set a decay base λ (0 < λ < 1), representing the proportion of trust level decay after passing through each flow node. If the number of data flow nodes in the data backtracking chain of a certain source data table D is n, the trust level TD of the source data table D can be expressed as: For example, assuming λ=0.9, meaning that the trust level decreases by 10% for each additional flow node. If the number of data flow nodes in the data backtracking chain of source data table D is n=3 (i.e., the data has passed through three different flow nodes, such as from the original PDF to the data provider's database, and then to the internal data processing platform), then the final trust level of its source data table D is 72.9%. This means that after passing through three flow nodes, the original trust level of data table D drops from 100% to 72.9%.

[0044] Optionally, the validity and weight of each data flow node are pre-set, and the validity of each data flow node in the data backtracking chain of each source data table is weighted and summed to obtain the trust level of each source data table.

[0045] For example, suppose the specified object is a company, and source data tables for this company are collected from three different data sources A, B, and C. Each source data table includes resource data for different resource items under the same report type within the same reporting period. During multi-source data validation, it is found that the resource data for "Net Profit" in the source data tables provided by the three data sources is inconsistent. In this case, it is determined that there is dirty data in the resource data for "Net Profit" in the source data tables provided by the three data sources. Specifically, the data backtracking chain of source data table A provided by data source A contains two flow nodes. Assuming the trust decay rate of each node is 0.95, the trust level of source data table A is 90.25%. The data backtracking chain of source data table B provided by data source B contains three flow nodes. Assuming the trust decay rate of each node is 0.95, the trust level of source data table B is 85.74%. The data backtracking chain of source data table C provided by data source C contains one flow node. Assuming the trust decay rate of each node is 0.95, the trust level of source data table C is 0.95%. The comparison shows that source data table C has the highest trust level. Therefore, the specific resource data of the "Net Profit" resource item in source data table C is selected as the calibration value, and the specific values ​​of the "Net Profit" resource item in source data tables A and B are replaced with the calibration values.

[0046] In some embodiments, after obtaining the calibration value, the calibration value, along with its source (original PDF file path, parsing timestamp), is written back to the corresponding dirty data record in the "Financial Statement Subject Self-Verification Information Table" to complete the cleaning and calibration of that data. A "Financial Statement Subject Calibration Result Table" is created as the final output. For each smallest comparison unit, if the smallest comparison unit has the highest "analysis calibration value" in the self-verification information table, generated by parsing, this value is used, and the data source is marked as "Multi-source verification self-calibration passed." If no parsing calibration value exists, but the unit has consistent multi-source data in the verification pass information table, this value is used, and it is marked as "Multi-source verification automatically passed." The entire system runs daily. When new external data is updated or historical data is corrected, the first-stage multi-source verification is automatically triggered. Newly generated dirty data triggers the second-stage precise parsing again, and the calibration result is updated to the final table. Thus, the financial database continuously evolves in a closed loop of "verification-parsing-calibration," maintaining high accuracy and timeliness.

[0047] The embodiments provided in this application standardize resource data from different data sources to form multi-source data tables, solving the problem of difficult horizontal comparison caused by inconsistent fields and chaotic formats between data sources in related technologies. When dirty data is identified in the same target resource item (such as "net profit") in the multi-source data tables of a specified object, its trust level is dynamically assessed based on the data backtracking chain of each source data table, i.e., the flow path and processing history. That is, the original source of the data is traced back through the data backtracking chain to assess the reliability of the source. Considering the processing links in the data flow process, the more flow nodes, the greater the risk of data contamination. The analysis results of the data backtracking chain are transformed into quantifiable trust level indicators, solving the problem of difficulty in reliability assessment. This can more accurately reflect the actual credibility of the source data tables and solve the problem that static rating cannot accurately reflect the timeliness of data and the trust changes in the processing process. The highest level of trust level data is selected from the data of the same target resource item in the multi-source data tables as a calibration value to replace the multi-source data. Based on the dirty data of the same target resource subject in the table, through a closed-loop mechanism of "verification-driven, precise analysis, and autonomous calibration," not only can dirty data be automatically detected, but its authoritative source can also be traced. Through a trigger-based cleaning mechanism, only data identified as dirty data is cleaned, avoiding unnecessary consumption of computing resources and significantly improving processing efficiency. Through a calibration value selection strategy that prioritizes trust, the credibility of calibration values ​​is ensured, avoiding calibration errors caused by improper data source selection. Correct values ​​are used for calibration, thereby forming a highly accurate and timely database. This solves the technical problem in related technologies where the reliability of multi-source data is difficult to assess, leading to the misselection of low-quality data for calibration and resulting in poor cleaning effects. It significantly improves the efficiency and accuracy of data processing, reduces the waste of computing resources, and ensures the high quality and timeliness of data. It achieves three major leaps: from "passive verification" to "active calibration," from "static weighting" to "scenario adaptation," and from "simple replacement" to "precise traceability," providing strong support for resource data analysis and management.

[0048] In an exemplary embodiment, each source data table includes resource data for different resource items under different report types for a specified object in different reporting periods. The method further includes: comparing data in preset comparison units in the multi-source data tables in parallel, and in response to inconsistencies in resource data in at least one target preset comparison unit in the multi-source data tables, marking resource data in at least one target preset comparison unit in the multi-source data tables as dirty data; the preset comparison unit refers to the same resource item of the specified object under the same report type in the same reporting period.

[0049] The pre-defined comparison unit refers to the basic unit used for data comparison, specifically defined as "specified object - reporting period - report type - resource item". For example, "Company A - Q1 2024 - balance sheet - cash and cash equivalents" is a pre-defined comparison unit. The pre-defined comparison unit is the smallest comparison granularity, designed to ensure the granularity and accuracy of data verification. The target pre-defined comparison unit refers to the pre-defined comparison unit where inconsistencies in resource data are found during the multi-source data comparison process. The data in the target pre-defined comparison unit is the focus of the next step of cleaning and calibration.

[0050] Optionally, the server parses the data structure of each source data table, extracting dimensional information, including the specified object identifier, reporting period, report type, and resource subject. Then, the server integrates the dimensional information from all source data tables to generate all possible preset comparison units. The generation of preset comparison units follows the Cartesian product principle, meaning that all combinations of the specified object, reporting period, report type, and resource subject constitute a preset comparison unit. The server employs multi-threading or distributed computing technology to distribute the preset comparison units in the preset comparison unit list to multiple computing nodes or threads for parallel processing. In response to detecting inconsistencies in resource data in at least one target preset comparison unit, the resource data in at least one target preset comparison unit across the multi-source data tables is marked as dirty data.

[0051] This embodiment allows for parallel comparison of resource data of a specified object within a preset comparison unit consisting of three dimensions: the same reporting period, the same report type, and the same resource subject. This enables rapid identification of inconsistencies between data, overcoming the lack of fine-grained comparison in traditional data verification and improving the speed and accuracy of dirty data identification.

[0052] In one exemplary embodiment, the data backtracking chain of each source data table includes at least one data flow node. Determining the trust level of each source data table based on its data backtracking chain includes the following steps:

[0053] 1. Determine the basic trust value of each source data table based on the validity of each data flow node in the data backtracking chain of each source data table.

[0054] 2. Determine the penalty factor for the data backtracking chain of each source data table based on the number of data flow nodes in the data backtracking chain of each source data table.

[0055] Third, the base trust value of each source data table is penalized using the penalty factor of the data backtracking chain of each source data table to obtain the trust value of each source data table.

[0056] The validity of each data flow node refers to its accuracy and compliance. For example, the consistency between the source data tables before and after entering each data flow node can be used to determine the validity of each data flow node. Specifically, the similarity (such as hash value) between the source data tables before and after entering each data flow node can be used to determine the validity of each data flow node. For example, a static validity can be pre-set for each data flow node, and the trust level of each source data table will dynamically change when the data backtracking chain of each source data table contains different data flow nodes.

[0057] The baseline trust value for each source data table is a preliminary value reflecting the overall credit level of the data source, calculated by the system based on the validity scores of the data flow nodes in the data backtracking chain of each source data table. It constitutes the first level of data source trust assessment. For example, the weighted average of the validity of the data flow nodes in the data backtracking chain of each source data table can be used to determine the baseline trust value for each source data table. Alternatively, the baseline trust value can be calculated using a geometric mean method based on the validity of each data flow node in the data backtracking chain of each source data table. Specifically, for each source data table, the validity values ​​of all data flow nodes in its data backtracking chain are first extracted. Then, the computer uses a geometric mean method to calculate the baseline trust value. The formula for the geometric mean method is: Baseline Trust Value = (πNode Validity)^(1 / n), where π represents a multiplication operation, n represents the number of data flow nodes, and node validity represents the validity value of that node. The geometric mean method is characterized by the fact that when the validity of any node is low, the overall trust baseline value will be significantly affected. This aligns with the reality that problems in any link of the data flow process will affect the overall data quality. For example, for a source data table, if the validity of its five data flow nodes are 0.95, 0.90, 0.92, 0.88, and 0.93 respectively, then the trust baseline value = (0.95 × 0.90 × 0.92 × 0.88 × 0.93)^(1 / 5) = (0.648)^(0.2) ≈ 0.916.

[0058] The penalty factor for each source data table's data backtracking chain is a coefficient calculated based on the number of data flow nodes in the backtracking chain. It aims to reflect the negative impact of data flow complexity on data source trust. Too many flow nodes may increase the risk of data mistransmission; therefore, the system penalizes the base trust value of the data source.

[0059] Optionally, the server can calculate the penalty factor using a linear function based on the number of data flow nodes in the data backtracking chain of each source data table. Specifically, for each source data table, the computer device first counts the number of data flow nodes in its data backtracking chain. Then, the computer device calculates the penalty factor using a linear function. The formula for the linear function is: Penalty Factor = Number of Nodes × Single Node Penalty Coefficient, where the single node penalty coefficient represents the penalty strength of each data flow node on trust level, and the value of the single node penalty coefficient is usually between 0 and 0.1. The characteristic of the linear function is that the penalty factor is directly proportional to the number of nodes; the more nodes, the larger the penalty factor. For example, for a source data table whose data backtracking chain contains 5 data flow nodes, if a linear function is used and the single node penalty coefficient is 0.02, then the penalty factor = 5 × 0.02 = 0.1.

[0060] Optionally, the server can calculate the penalty factor using a logarithmic function based on the number of data flow nodes in the data backtrack chain of each source data table. Specifically, for each source data table, the computer device first counts the number of data flow nodes in its data backtrack chain. Then, it calculates the penalty factor using a logarithmic function. The formula for the logarithmic function is: Penalty Factor = log(1 + Number of Nodes) × Node Penalty Coefficient, where the node penalty coefficient represents the overall adjustment coefficient of the penalty factor, and the value of the node penalty coefficient is usually between 0 and 0.5. The characteristic of the logarithmic function is that the penalty factor increases with the increase of the number of nodes, but the growth rate gradually slows down. This reflects the characteristic that the longer the data flow chain, the lower the trust level, but the reduction rate gradually slows down. For example, for a certain source data table, its data backtrack chain contains 5 data flow nodes. If the logarithmic function is used and the node penalty coefficient is 0.3, then the penalty factor = log(1 + 5) × 0.3 = log(6) × 0.3 ≈ 0.778 × 0.3 ≈ 0.233.

[0061] In this embodiment, the trust level of each source data table is calculated by combining the base trust level value and the penalty factor, reflecting the reliability and authority of the data provided by the data source after comprehensive consideration and adjustment.

[0062] Optionally, the computer device uses the penalty factor of the data backtracking chain of each source data table to multiply the base value of the trust level of each source data table to obtain the trust level of each source data table. Figure 3 Here is a flowchart of another optional data table cleaning method provided in the embodiments of this application, such as... Figure 3As shown, specifically, for each source data table, the computer device first obtains its base trust value and penalty factor. Then, the computer device calculates the final trust value using a multiplicative penalty. The formula for the multiplicative penalty is: Trust Value = Base Trust Value × (1 - Penalty Factor). A key characteristic of the multiplicative penalty is that the impact of the penalty factor on the base trust value is relative; the higher the base trust value, the larger the absolute value of the penalty. For example, for a source data table with a base trust value of 0.919 and a penalty factor of 0.1, if a multiplicative penalty is used, then the trust value = 0.919 × (1 - 0.1) = 0.919 × 0.9 = 0.8271.

[0063] Optionally, the computer device uses the penalty factor of the data backtracking chain of each source data table to additively penalize the base trust value of each source data table to obtain the trust value of each source data table. Specifically, for each source data table, the computer device first obtains its base trust value and penalty factor. Then, the computer device calculates the final trust value using additive penalty. The formula for additive penalty is: Trust Value = Base Trust Value - Penalty Factor. The characteristic of additive penalty is that the impact of the penalty factor on the base trust value is absolute; regardless of the base trust value, the absolute value of the penalty is the same. For example, for a source data table with a base trust value of 0.919 and a penalty factor of 0.1, if additive penalty is used, then the trust value = 0.919 - 0.1 = 0.819.

[0064] Optionally, the computer device applies an exponential penalty to the baseline trust value of each source data table using the penalty factor of the data backtracking chain for each source data table, thus obtaining the trust value of each source data table. Specifically, for each source data table, the computer device first obtains its baseline trust value and penalty factor. Then, the computer device calculates the final trust value using exponential penalty. The formula for exponential penalty is: Trust Value = Baseline Trust Value^(1 + Penalty Factor), where ^ represents exponentiation. The characteristic of exponential penalty is that the impact of the penalty factor on the baseline trust value is exponential; the higher the baseline trust value, the larger the absolute value of the penalty, and the stronger the penalty is than multiplicative penalty. For example, for a source data table with a baseline trust value of 0.919 and a penalty factor of 0.1, if exponential penalty is applied, then Trust Value = 0.919^(1+0.1) = 0.919^1.1 ≈ 0.912.

[0065] Optionally, the computer device applies a mixed penalty to the baseline trust value of each source data table using the penalty factor of the data backtracking chain for each source data table, thus obtaining the trust value of each source data table. Specifically, for each source data table, the computer device first obtains its baseline trust value and penalty factor. Then, the computer device calculates the final trust value using a mixed penalty. The formula for the mixed penalty is: Trust Value = α × (Baseline Trust Value - Penalty Factor) + (1 - α) × (Baseline Trust Value × (1 - Penalty Factor)), where α is the mixing coefficient, ranging from 0 to 1, used to adjust the weights of additive and multiplicative penalties. The characteristic of the mixed penalty is that it combines the advantages of additive and multiplicative penalties, allowing for flexible adjustment of the penalty method according to actual needs. For example, for a source data table, its base trust value is 0.919 and the penalty factor is 0.1. If a mixed penalty is used and the mixed coefficient α is 0.5, then the trust value = 0.5 × (0.919 - 0.1) + 0.5 × (0.919 × 0.9) = 0.5 × 0.819 + 0.5 × 0.8271 = 0.4095 + 0.41355 = 0.82305.

[0066] This embodiment achieves quality control over the entire data lifecycle by evaluating the effectiveness of each data flow node, avoiding one-sided evaluations based solely on the final data state, and making trust levels more reflective of the data's authenticity and reliability. The introduction of a penalty factor based on the number of nodes reasonably quantifies the potential risks brought about by an increase in data flow links, meeting the actual needs of data quality management. The more flow nodes there are, the greater the possibility of data contamination, and the penalty mechanism ensures the rationality of the evaluation.

[0067] In one exemplary embodiment, selecting a calibration value from the same target resource subject in multiple source data tables based on the trust level of each source data table includes: determining the reliability of the resource data of the target resource subject in each source data table, and determining the authority of the resource data of the target resource subject in each source data table based on the trust level of each source data table and the reliability of the resource data of the target resource subject in each source data table; and determining the resource data of the target resource subject in the source data table corresponding to the highest authority as the calibration value.

[0068] In this context, the reliability of resource data for each target resource item in the source data table refers to the degree of trustworthiness and stability of the resource data at a specific point in time or within a specific time period. Reliability reflects whether the resource data can accurately reflect the resource distribution of a specified object. Higher reliability indicates more reliable resource data and should be given higher weight in data calibration.

[0069] Optionally, check whether the required fields for the target resource subject in each source data table are complete, and calculate the completeness rate of the required fields. Check whether the data records for the target resource subject in each source data table are complete in different reporting periods, and calculate the data record completeness rate. Calculate a weighted average of the field completeness rate and the record completeness rate to obtain a data integrity score; the weights can be set according to business needs. Establish an authority scoring system for each data source, including indicators such as data source qualification certification, historical data quality performance, and industry recognition. Score each data source according to the scoring system to obtain source authority scores for different indicators. The source authority score refers to the degree of authority of the data source providing the source data table, including the data source qualification, historical performance, and industry recognition. Calculate a weighted average of the source authority scores for different indicators to obtain the final source authority score. Standardize the data integrity score and the source authority score to make them of the same order of magnitude. Calculate the comprehensive reliability score using a weighted average method; for example, the data integrity weight is 0.5, and the source authority weight is 0.5. Determine the reliability level of the resource data for the target resource subject in the source data table based on the comprehensive reliability score.

[0070] The authority level of resource data for the target resource category in each source data table refers to the comprehensive evaluation result derived by considering both the trustworthiness of the source data table and the reliability of the resource data. In this application, authority level is the core basis for selecting calibration values, reflecting the overall credibility of the resource data for the target resource category in the source data table. Authority level comprehensively considers the reliability of the data source and the quality of the data itself, and is the final basis for selecting calibration values ​​in data calibration. The higher the authority level, the more credible the resource data is, and the more likely it should be selected as a calibration value.

[0071] Optionally, the computer device determines the authority of the resource data for the target resource subject in each source data table based on the trust level of each source data table and the reliability of the resource data for the target resource subject in each source data table. Specifically, for each source data table, the computer device first obtains the trust level of the source data table and the reliability of the resource data for the target resource subject in the source data table. Then, the computer device calculates the authority using a weighted average or weighted summation method.

[0072] In this application, the calibration value is selected based on the authority of the resource data of the target resource subject in each source data table, with the highest authority being chosen. The selection of the calibration value follows the principle of authority priority, that is, the resource data in the source data table with the highest authority is selected as the calibration value.

[0073] The maximum authority level refers to the highest authority level among the resource data of the target resource subject in all source data tables. In this application, the resource data in the source data table corresponding to the maximum authority level is the most reliable and should be selected as the calibration value. Determining the maximum authority level requires comparing the authority levels of the resource data of the target resource subject in all source data tables and selecting the highest value.

[0074] This embodiment comprehensively considers the trustworthiness of the source data tables and the reliability of the resource data to accurately select calibration values, ensuring the credibility and accuracy of the calibration values. It proposes an authority calculation scheme that couples the trustworthiness of each source data table with the reliability of the resource data for the target resource subject within each source data table. This organically combines the reliability of the data source with the quality of the data itself, comprehensively evaluating the authority of the resource data. Compared to existing calibration value selection methods that only consider a single factor (such as data update time or data source trustworthiness), this embodiment can more comprehensively and accurately assess the overall credibility of resource data, avoiding misjudgments caused by single-factor evaluation. Furthermore, this embodiment adopts an authority-first calibration value selection strategy, choosing the resource data from the source data table with the highest authority as the calibration value, ensuring the credibility and accuracy of the calibration values ​​and improving the quality of data cleaning.

[0075] In an exemplary embodiment, the reliability of the resource data of the target resource subject in each source data table includes consistency and timeliness; the consistency of the resource data of the target resource subject in each source data table characterizes the matching degree between the resource data of the target resource subject in each source data table; the timeliness of the resource data of the target resource subject in each source data table represents the update speed of the resource data of the target resource subject in each source data table.

[0076] Consistency refers to the degree of matching or similarity between resource data of the same target resource category in multiple source data tables. In this application, consistency is an important component of reliability, used to measure the degree of agreement between the same resource data in different source data tables. Consistency can be quantified by calculating indicators such as similarity, difference, and coefficient of variation between resource data. The higher the consistency, the closer the resource data is in multiple source data tables, and the higher the reliability of the resource data.

[0077] Optionally, for consistency calculation, the computer device first extracts resource data for the target resource item from all source data tables. Then, the computer device calculates the similarity or difference between these resource data. Similarity can be calculated using methods such as cosine similarity, Pearson correlation coefficient, and Euclidean distance. Difference can be calculated using methods such as standard deviation, coefficient of variation, and maximum variance. The computer device converts the similarity or difference into a consistency index, which typically ranges from 0 to 1, with higher values ​​indicating higher consistency. For example, for the target resource item "Company A - Q1 2024 - Balance Sheet - Cash and Cash Equivalents," the computer device calculates the coefficient of variation for the cash and cash equivalents values ​​in the three source data tables to be 0.02. Converting the coefficient of variation into a consistency index, the consistency index = 1 - coefficient of variation = 1 - 0.02 = 0.98.

[0078] Timeliness refers to the speed at which resource data is updated or its freshness. In this application, timeliness is an important component of reliability, used to measure the freshness and timeliness of resource data. Timeliness can be quantified by calculating indicators such as the resource data's publication time, update time, and the time difference with the current time. Higher timeliness indicates newer resource data and higher reliability.

[0079] Optionally, for calculating timeliness, the computer device first extracts the publication time or update time of the resource data for the target resource item in the source data table. Then, the computer device calculates the time difference between that time and the current time. The smaller the time difference, the newer the resource data and the higher the timeliness. The computer device converts the time difference into a timeliness index, which typically ranges from 0 to 1, with a larger value indicating higher timeliness. For example, for a monetary value in a source data table, if its publication date is April 1, 2024, and the current time is April 15, 2024, the time difference is 14 days. The computer device converts the time difference into a timeliness index: Timeliness index = 1 / (1 + Time difference / 30) = 1 / (1 + 14 / 30) = 1 / 1.467 ≈ 0.682.

[0080] In some embodiments, when evaluating "a company's net profit in the first quarter of 2024," the values ​​provided by data source A (official annual report, high credibility but delayed release) and data source B (data provider, high timeliness but moderate credibility) are inconsistent. Conventional solutions only perform static credibility ratings on the data sources; however, static weighting methods cannot handle this complex trade-off between "credibility and timeliness," potentially leading to the use of outdated authoritative data or inaccurate, up-to-date data. Therefore, to address this issue, this embodiment provides a dynamic, adaptive, multi-factor coupled method for determining authority. This method derives authority based on dynamic coupling of multiple factors, resolving the inability of related technologies to handle the complex dependency between credibility and timeliness.

[0081] In some embodiments, the authority of the resource data of the target resource subject in each source data table is determined based on the trust level of each source data table and the reliability of the resource data of the target resource subject in each source data table, including:

[0082] The product of the trust level of each source data table and the timeliness of the resource data of the target resource subject in each source data table is used to determine the trust weight value of each source data table; the sum of the trust weight value of each source data table and the consistency of the resource data of the target resource subject in each source data table is used to determine the authority of the resource data of the target resource subject in each source data table.

[0083] Timeliness is quantified as a decay function based on data release time, while consistency is measured by calculating the coefficient of variation or similarity of values ​​from multiple data sources. The trust-weighted value is the numerical value obtained by multiplying the trust level of the source data table by the timeliness of the resource data in the target resource subject. In this application, the trust-weighted value is an important component of authority calculation, reflecting the combined impact of the reliability and freshness of the data source. A higher trust-weighted value indicates a more reliable and fresher data source for the source data table, thus increasing the credibility of the resource data. The formula for calculating the trust-weighted value is: Trust-weighted value = Trust level × Timeliness.

[0084] It should be noted that this embodiment uses multiplication to calculate the weighted trust value, i.e., "Weighted Trust Value = Trust Value × Timeliness," rather than addition or a weighted average. This is primarily due to the following reasons: there is a strong coupling between trust value and timeliness; both are indispensable. Only when both are high is the reliability of the resource data truly high. Multiplication accurately expresses this coupling, ensuring that when either trust value or timeliness is low, the weighted trust value will significantly decrease, thus avoiding misjudgments caused by a single factor being high. For example, if the trust value is 0.9 and the timeliness is 0.1, the weighted trust value obtained using multiplication is 0.09, while the weighted trust value obtained using addition is 1.0. Obviously, multiplication better reflects the actual situation. Furthermore, multiplication inherently has a penalty mechanism, effectively avoiding the "weakest link effect," where a deficiency in any factor leads to a decline in the overall evaluation result.

[0085] Authority refers to a comprehensive evaluation result that takes into account the trustworthiness of the source data table, the timeliness and consistency of the resource data. For example... Figure 3As shown, in this application, the formula for calculating authority is: Authority = Trust Weighted Value + Consistency = Trust × Timeliness + Consistency. A higher authority indicates more reliable resource data and should be selected as the calibration value. In this embodiment, the formula for calculating authority balances the inherent reliability of the data source, the freshness of the data, and the collaborative verification effect among multiple data sources.

[0086] It should be noted that this embodiment uses a summation operation to calculate the authority score, i.e., "Authority Score = Trust Weighted Value + Consistency," rather than a multiplication operation or a weighted average. This is primarily because: the trust weighted value and consistency are two relatively independent factors, evaluating the credibility of resource data from different dimensions, and there is no strong coupling relationship. Therefore, using a summation operation can reasonably combine these two factors. Furthermore, the trust weighted value and consistency are complementary, compensating for each other's shortcomings. The summation operation can balance the contributions of the two evaluation dimensions, ensuring the rationality and accuracy of the authority score calculation. For example, if the trust weighted value is 0.5 and the consistency is 0.5, the authority score obtained by the summation operation is 1.0, reflecting a moderate overall credibility of the resource data. However, the authority score obtained by the multiplication operation, 0.25, might be too conservative.

[0087] Optionally, for a target resource subject, the computer device first extracts the resource data for that target resource subject from all source data tables. Then, the computer device calculates the similarity or difference between the resource data for that target resource subject in all source data tables, converting the similarity or difference into a consistency index. For each source data table, the computer device first extracts the publication time or update time of the resource data for the target resource subject in that source data table, and calculates the time difference between that time and the current time. The computer device converts the time difference into a timeliness index. The computer device determines the trust weight value for each source data table by multiplying the trust value of each source data table by the timeliness of the resource data for the target resource subject in each source data table; and determines the authority of the resource data for the target resource subject in each source data table by summing the trust weight value for each source data table and the consistency of the resource data for the target resource subject in each source data table.

[0088] Compared to existing technologies that use addition or weighted averaging, which fail to accurately reflect the coupling relationship between trust and timeliness and are prone to misjudgment, this embodiment uses the formula "Authority = Trust × Timeliness + Consistency" to solve the problem of improper handling of the relationship between trust and timeliness in related technologies. Multiplication accurately reflects the coupling relationship between trust and timeliness, while summation reasonably combines the relatively independent factors of trust weighting and consistency, improving the accuracy of authority calculation. Secondly, it enhances the comprehensiveness of authority calculation by comprehensively considering the three dimensions of trust, timeliness, and consistency, enabling a more comprehensive assessment of the overall credibility of resource data. It also improves the quality of data cleaning, ensuring the accuracy and reliability of calibration values, ultimately outputting a high-quality, self-developed resource database with a clear source and cross-validated or authoritative source calibration.

[0089] In one exemplary embodiment, trust level and timeliness are not independent. For example, the timeliness of an official annual report with the highest trust level (published annually) may be lower than that of a data provider with medium trust level (published quarterly or monthly). Directly multiplying them would systematically underestimate the value of high-frequency, high-trust data. Therefore, to address this issue, in this embodiment, such as Figure 3 As shown, by introducing a learnable timeliness sensitivity coefficient R (R>0), the modified authority calculation formula can be expressed as: R is not a fixed value, but rather learned through regression analysis of historical data. For a specific subject or data source type, R > 1 indicates that the impact of timeliness is amplified (such as liquidity indicators applicable to high-frequency trading scenarios), while R < 1 indicates that the impact of timeliness is suppressed (such as capital structure indicators that place more emphasis on historical stability). For example, for liquidity indicators (such as the "quick ratio"), where the market changes rapidly, the system will learn an R > 1 value, amplifying the weight of timeliness and giving a greater advantage to data source B, which has higher timeliness, in the calculation. Conversely, for capital structure indicators (such as the "debt-to-equity ratio"), the system will learn an R < 1 value, suppressing the impact of timeliness and favoring data source A, which has higher trustworthiness. This resolves the non-linear adjustment relationship that cannot be expressed by simply multiplying trustworthiness and timeliness.

[0090] In some embodiments, the above method further includes:

[0091] 1. Based on the importance of the timeliness of the resource data of the target resource subject in each source data table, determine the timeliness sensitivity coefficient of the resource data of the target resource subject in each source data table; the timeliness sensitivity coefficient of the resource data of the target resource subject in each source data table is used to quantify the degree of influence of the timeliness of the resource data of the target resource subject in each source data table on the authority.

[0092] Second, the timeliness of the modified resource data of the target resource subject in each source data table will be determined by using the timeliness of the resource data of the target resource subject in each source data table as the base and the timeliness sensitivity coefficient of the resource data of the target resource subject in each source data table as the exponent.

[0093] The importance of timeliness refers to its relative significance or weight in the calculation of authority. In this application, the importance of timeliness reflects the sensitivity of different target resource items to the freshness of data. For some resource items, such as liquidity indicators and market data, timeliness is more important because these data change rapidly over time and need to be updated promptly to reflect the true situation; for some resource items, such as historical costs and fixed assets, timeliness is less important because these data are relatively stable and change slowly over time. The importance of timeliness can be determined through various methods such as historical data analysis, expert evaluation, and business needs.

[0094] Optionally, the computer equipment collects historical data for the resource subject at multiple points in time and calculates the degree of change of the historical data over time. The degree of change can be measured by calculating indicators such as volatility, trend, and autocorrelation of the time series data. The greater the volatility, the stronger the trend, and the weaker the autocorrelation, the faster the resource subject changes over time, and the higher the importance of timeliness. The computer equipment converts the degree of change into an importance index of timeliness. The value of the importance index of timeliness is usually between 0 and 1, with a larger value indicating higher importance of timeliness.

[0095] Optionally, the computer equipment determines the importance of timeliness based on business needs. Business needs include multiple aspects such as data usage frequency, data update frequency, and the degree of impact of data on decisions. The higher the data usage frequency, the higher the data update frequency, and the greater the degree of impact of data on decisions, the higher the importance of timeliness. The computer equipment converts business needs into timeliness importance indicators, which typically range from 0 to 1, with higher values ​​indicating higher timeliness importance.

[0096] The timeliness sensitivity coefficient is an exponential coefficient used to quantify the impact of timeliness on authority. In this application, the timeliness sensitivity coefficient is a real number greater than 0, used to adjust the degree of influence of timeliness in the authority calculation. The larger the timeliness sensitivity coefficient, the greater the impact of timeliness on authority; the smaller the timeliness sensitivity coefficient, the smaller the impact of timeliness on authority. The value of the timeliness sensitivity coefficient is usually between 0.5 and 2, and the specific value is determined according to the importance of timeliness. The introduction of the timeliness sensitivity coefficient enables the authority calculation to dynamically adjust the degree of influence of timeliness according to the characteristics of different resource subjects, improving the accuracy and flexibility of authority calculation.

[0097] It should be noted that different resource categories exhibit significantly different sensitivities to data freshness. Applying a uniform timeliness weight would fail to accurately reflect the characteristics of each category. For example, liquidity indicators such as cash and cash equivalents experience rapid data changes over time, making timeliness crucial and thus requiring a higher weight. Conversely, relatively stable categories like fixed assets have lower timeliness importance and should be assigned a lower weight. By determining a timeliness sensitivity coefficient, the impact of timeliness on authority calculation can be dynamically adjusted based on the characteristics of different resource categories, improving the accuracy and flexibility of authority calculation. This mechanism better adapts to the characteristics of different resource categories, ensuring a more scientific and reasonable approach to authority calculation.

[0098] Optionally, the computer equipment uses a linear mapping method to convert the timeliness importance index into a timeliness sensitivity coefficient. The timeliness sensitivity coefficient = 0.5 + timeliness importance index × 1.5. When the timeliness importance index is 0, the timeliness sensitivity coefficient is 0.5; when the timeliness importance index is 1, the timeliness sensitivity coefficient is 2.0. This mapping method ensures that the timeliness sensitivity coefficient ranges from 0.5 to 2.0.

[0099] Optionally, the computer equipment uses a segmented mapping method to convert the timeliness importance index into a timeliness sensitivity coefficient. When the timeliness importance index is less than 0.3, the timeliness sensitivity coefficient = 0.5; when the timeliness importance index is between 0.3 and 0.7, the timeliness sensitivity coefficient = 0.5 + (timeliness importance index - 0.3) / 0.4 × 1.0; when the timeliness importance index is greater than 0.7, the timeliness sensitivity coefficient = 1.5 + (timeliness importance index - 0.7) / 0.3 × 0.5. This mapping method allows the timeliness sensitivity coefficient to grow at different rates within different importance ranges, providing greater flexibility.

[0100] The modified timeliness refers to the timeliness value obtained by exponentially calculating the original timeliness. In this application, the modified timeliness is obtained by using the original timeliness as the base and the timeliness sensitivity coefficient as the exponent. The formula for calculating the modified timeliness is: Modified Timeliness = Timeliness^Timeliness Sensitivity Coefficient. When the timeliness sensitivity coefficient is greater than 1, the modified timeliness will be less than the original timeliness, and the impact of timeliness is amplified; when the timeliness sensitivity coefficient is less than 1, the modified timeliness will be greater than the original timeliness, and the impact of timeliness is suppressed; when the timeliness sensitivity coefficient is equal to 1, the modified timeliness is equal to the original timeliness, and the impact of timeliness remains unchanged. The modified timeliness is used to replace the original timeliness in the calculation of authority.

[0101] This embodiment introduces a timeliness sensitivity coefficient. By analyzing multiple dimensions such as the characteristics of resource subjects, business needs, and historical data quality, the importance of timeliness is determined, and thus the timeliness sensitivity coefficient is determined. Compared with the existing method that uses a fixed timeliness weight, the timeliness sensitivity coefficient determination mechanism in this embodiment can dynamically adjust the influence of timeliness in authority calculation according to the characteristics of different resource subjects, improving the accuracy and flexibility of authority calculation. Simultaneously, this embodiment uses an exponential calculation method, with timeliness as the base and the timeliness sensitivity coefficient as the exponent, which can more accurately quantify the influence of timeliness on authority, avoiding the inflexibility of the fixed-weight method. This mechanism can better adapt to the characteristics of different resource subjects, ensuring that authority calculation is more scientific and reasonable, ultimately outputting a high-quality, self-developed resource database with clear sources and cross-validation or authoritative source calibration.

[0102] In one exemplary embodiment, during practical application, it was found that when the timeliness of a certain data source is abnormally high (e.g., sudden high frequency of releases, early releases), it excessively affects the overall authority. To solve this single data source "overshooting" problem, such as... Figure 3 As shown, this embodiment introduces a smoothing factor S, and the modified authority calculation formula can be expressed as: The smoothing factor S acts as a "buffer" for timeliness calculations. When the timeliness of a data source is abnormally high, The growth curve will be greater than The system achieves a smoother, more gradual transition, preventing the distortion of overall authority assessment due to the outstanding performance of a single dimension. This makes the model more robust to abnormal fluctuations and more comprehensively reflects the authenticity and credibility of the data. The system can intelligently adapt to the different data characteristic requirements of various business scenarios, achieving a leap from "static weighting" to "scenario-adaptive" computing. By introducing a timeliness sensitivity coefficient R and a smoothing factor S, the system can adapt to the data characteristic requirements of different business scenarios. A daily closed-loop mechanism ensures continuous absorption of new data, self-verification, and calibration, making the database a dynamically evolving, highly reliable database. This significantly reduces operational costs and forms a complete closed loop from automatic problem discovery and authoritative source tracing to intelligent result integration and continuous iteration, making the resource database a dynamically highly reliable database capable of continuous self-verification, self-calibration, and self-evolution.

[0103] In some embodiments, the above method further includes:

[0104] Obtain a pre-determined smoothing factor; the smoothing factor is used to balance the impact of the timeliness fluctuation of the resource data of the target resource subject in each source data table on the authority; the sum of the timeliness of the resource data of the target resource subject in each source data table and the smoothing factor is used to determine the modified timeliness of the resource data of the target resource subject in each source data table.

[0105] The smoothing factor is an adjustment coefficient used to balance the impact of timeliness fluctuations on authority. Timeliness fluctuations refer to the degree or magnitude of timeliness changes over time. In this application, the smoothing factor is a real number greater than or equal to 0, used to smooth timeliness and reduce the impact of timeliness fluctuations on authority. The introduction of the smoothing factor allows the authority calculation to remain relatively stable when timeliness fluctuations are large, avoiding large fluctuations in authority due to drastic changes in timeliness. The value of the smoothing factor is usually between 0 and 0.5, and the specific value is determined by various methods such as business needs, historical data analysis, and expert evaluation. The larger the smoothing factor, the greater the suppression of timeliness fluctuations; the smaller the smoothing factor, the less the suppression of timeliness fluctuations.

[0106] Optionally, the computer device pre-collects historical time-sensitive data of the target resource subject from the source data table with the highest trust level, calculates the mean and standard deviation of the historical time-sensitive data of the target resource subject, and determines the ratio between the standard deviation and the mean as the smoothing factor.

[0107] In this embodiment, the modified timeliness is calculated using addition, i.e., "Modified Timeliness = Timeliness + Smoothing Factor," instead of other calculation methods, for the following reasons: Timeliness fluctuations can significantly impact authority calculation. When timeliness fluctuations are large, directly using the original timeliness will lead to significant fluctuations in authority, affecting the accuracy of the calibration value selection. By introducing a smoothing factor and adjusting the timeliness additively, the impact of timeliness fluctuations on authority can be effectively suppressed, maintaining relative stability of authority. Addition can directly and simply adjust timeliness, avoiding the computational overhead of complex calculations and improving computational efficiency. Simultaneously, the introduction of the smoothing factor provides the system with a flexible adjustment mechanism, allowing dynamic adjustment of timeliness based on the characteristics of different resource subjects, improving the adaptability and accuracy of authority calculation. For example, if the timeliness is 0.682 and the smoothing factor is 0.3, then the modified timeliness = 0.682 + 0.3 = 0.982, resulting in a moderate improvement in timeliness and reducing the impact of fluctuations.

[0108] This embodiment proposes a timeliness fluctuation balancing mechanism based on a smoothing factor. By adding the timeliness and the smoothing factor, a modified timeliness is obtained, reducing the impact of timeliness fluctuations on authority and maintaining the relative stability of authority. This solves the problem of authority instability caused by timeliness fluctuations in related technologies, improves the accuracy and reliability of authority calculation, and provides a more reliable basis for the selection of calibration values.

[0109] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0110] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as read-only memory (ROM) / random access memory (RAM), magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0111] According to another aspect of the embodiments of this application, a data table cleaning apparatus is also provided, which can be used to implement the data table cleaning method provided in the above embodiments, and will not be repeated hereafter. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0112] Figure 4 This is a structural block diagram of an optional data table cleaning device according to an embodiment of this application, such as... Figure 4 As shown, the data table cleaning device includes:

[0113] The acquisition module 402 is used to acquire a multi-source data table for a specified object; the multi-source data table refers to a table of resource data about the specified object collected from different data sources; each source data table in the multi-source data table includes resource data of different resource subjects;

[0114] The trust confirmation module 404 is used to respond to the presence of dirty data in the same target resource subject in the multi-source data tables of a specified object, and to determine the trust level of each source data table based on the data backtracking chain of each source data table; the data backtracking chain of each source data table refers to the flow path of each source data table; the data backtracking chain of each source data table is used to trace the original source of each source data table in reverse.

[0115] The data cleaning module 406 is used to select calibration values ​​from the same target resource subject in multiple source data tables based on the trust level of each source data table, and use the calibration values ​​to replace the dirty data of the same target resource subject in the multiple source data tables.

[0116] In an exemplary embodiment, each source data table includes resource data of different resource subjects under different report types for a specified object in different reporting periods; the acquisition module 402 is further configured to compare the data in preset comparison units in the multi-source data tables in parallel, and in response to inconsistencies in resource data in at least one target preset comparison unit in the multi-source data tables, mark the resource data in at least one target preset comparison unit in the multi-source data tables as dirty data; the preset comparison unit refers to the same resource subject under the same report type for the specified object in the same reporting period.

[0117] In an exemplary embodiment, the data backtracking chain of each source data table includes at least one data flow node. The trust confirmation module 404 is further configured to determine the basic trust value of each source data table based on the validity of each data flow node in the data backtracking chain of each source data table; determine the penalty factor of the data backtracking chain of each source data table based on the number of data flow nodes in the data backtracking chain of each source data table; and use the penalty factor of the data backtracking chain of each source data table to penalize the basic trust value of each source data table to obtain the trust value of each source data table.

[0118] In an exemplary embodiment, the data cleaning module 406 is further configured to determine the reliability of the resource data of the target resource subject in each source data table, and to determine the authority of the resource data of the target resource subject in each source data table based on the trust level of each source data table and the reliability of the resource data of the target resource subject in each source data table;

[0119] The resource data of the target resource subject in the source data table corresponding to the highest authority is determined as the calibration value.

[0120] In an exemplary embodiment, the reliability of the resource data of the target resource subject in each source data table includes consistency and timeliness; the consistency of the resource data of the target resource subject in each source data table characterizes the matching degree between the resource data of the target resource subject in each source data table; the timeliness of the resource data of the target resource subject in each source data table represents the update speed of the resource data of the target resource subject in each source data table; the data cleaning module 406 is further configured to determine the trust weight value corresponding to each source data table by multiplying the trust value of each source data table and the timeliness of the resource data of the target resource subject in each source data table; and to determine the authority of the resource data of the target resource subject in each source data table by summing the trust weight value corresponding to each source data table and the consistency of the resource data of the target resource subject in each source data table.

[0121] In an exemplary embodiment, the data cleaning module 406 is further configured to determine the timeliness sensitivity coefficient of the resource data of the target resource subject in each source data table based on the importance of the timeliness of the resource data of the target resource subject in each source data table; the timeliness sensitivity coefficient of the resource data of the target resource subject in each source data table is used to quantify the degree of influence of the timeliness of the resource data of the target resource subject in each source data table on the authority; and the modified timeliness of the resource data of the target resource subject in each source data table is determined by using the timeliness of the resource data of the target resource subject in each source data table as the base and the timeliness sensitivity coefficient of the resource data of the target resource subject in each source data table as the exponent.

[0122] In an exemplary embodiment, the data cleaning module 406 is further configured to obtain a predetermined smoothing factor; the smoothing factor is used to balance the impact of the timeliness fluctuation of resource data of target resource subjects in each source data table on the authority.

[0123] The timeliness of the resource data of the target resource subject in each source data table is determined by summing the timeliness and smoothing factor of the resource data.

[0124] It should be noted that the above modules can be implemented by software or hardware. For the latter, they can be implemented in the following ways, but are not limited to: all the above modules are located in the same processor; or, the above modules are located in different processors in any combination.

[0125] According to another aspect of the embodiments of this application, a computer-readable storage medium is provided, the computer-readable storage medium including a stored program, wherein the program executes the steps in any of the above method embodiments when it is run.

[0126] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as USB flash drives, ROMs, RAMs, portable hard drives, magnetic disks, or optical disks.

[0127] According to another aspect of the embodiments of this application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor is configured to perform the steps of any of the method embodiments described above via the computer program. In an exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.

[0128] Specific examples in this embodiment can be found in the examples described in the above embodiments and exemplary implementations, and will not be repeated here.

[0129] According to another aspect of the embodiments of this application, a computer program product is also provided, comprising a computer program / instructions containing program code for performing the methods shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via communication section 509, and / or installed from removable medium 511. When the computer program is executed by central processing unit 501, it performs various functions provided in the embodiments of this application. The sequence numbers of the embodiments of this application above are merely descriptive and do not represent the superiority or inferiority of the embodiments.

[0130] Figure 5 A schematic block diagram of a computer system architecture for implementing embodiments of the present application is shown. Figure 5 As shown, the computer system 500 includes a Central Processing Unit (CPU) 501, which can perform various appropriate actions and processes based on programs stored in ROM 502 or programs loaded into RAM 503 from storage section 508. Random access memory 503 also stores various programs and data required for system operation. The CPU 501, ROM 502, and RAM 503 are interconnected via bus 504. Input / output (I / O) interface 505 is also connected to bus 504.

[0131] The following components are connected to I / O interface 505: input section 506 including keyboard, mouse, etc.; output section 507 including cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; storage section 508 including hard disk, etc.; and communication section 509 including network interface card, modem, etc. Communication section 509 performs communication processing via a network such as the Internet. Drive 510 is also connected to I / O interface 505 as needed. Removable media 511, such as disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 510 as needed so that computer programs read from them can be installed into storage section 508 as needed.

[0132] Specifically, according to embodiments of this application, the processes described in the various method flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 509, and / or installed from removable medium 511. When the computer program is executed by central processing unit 501, it performs various functions defined in the system of this application.

[0133] It should be noted that, Figure 5 The computer system 500 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0134] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. They can be implemented using computer-executable program code, and thus can be stored in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those described herein, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.

[0135] The above are merely preferred embodiments of this application and are not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this application should be included within the protection scope of this application.

Claims

1. A data table cleaning method, characterized in that, include: Retrieve multi-source data tables for a specified object; The multi-source data table refers to a resource data table about the specified object collected from different data sources; Each source data table in the multi-source data table includes resource data for different resource categories; In response to the presence of dirty data in the same target resource category across multiple source data tables of a specified object, the trust level of each source data table is determined based on its data backtracking chain. The data backtracking chain of each source data table refers to the flow path of each source data table. The data backtracking chain of each source data table is used to trace the original source of each source data table in reverse. Based on the trust level of each source data table, a calibration value is selected from the same target resource subject in the multi-source data tables, and the calibration value is used to replace the dirty data of the same target resource subject in the multi-source data tables.

2. The method according to claim 1, characterized in that, Each source data table includes resource data for different resource categories under different report types for the specified object in different reporting periods; the method further includes: The data in the preset comparison units of the multi-source data table are compared in parallel. In response to the inconsistency of resource data in at least one target preset comparison unit in the multi-source data table, the resource data in the at least one target preset comparison unit in the multi-source data table is marked as dirty data. The preset comparison unit refers to the same resource subject under the same report type for the specified object in the same reporting period.

3. The method according to claim 1, characterized in that, The data backtracking chain of each source data table includes at least one data flow node. Determining the trust level of each source data table based on its data backtracking chain includes: Based on the validity of each data flow node in the data backtracking chain of each source data table, determine the basic trust value of each source data table; The penalty factor for the data backtracking chain of each source data table is determined based on the number of data flow nodes in the data backtracking chain of each source data table. The trust level of each source data table is obtained by penalizing the base trust value of each source data table using the penalty factor of the data backtracking chain of each source data table.

4. The method according to claim 1, characterized in that, The step of selecting a calibration value from the same target resource item in the multi-source data tables based on the trust level of each source data table includes: Determine the reliability of the resource data of the target resource subject in each source data table, and determine the authority of the resource data of the target resource subject in each source data table based on the trust level of each source data table and the reliability of the resource data of the target resource subject in each source data table; The resource data of the target resource subject in the source data table corresponding to the highest authority is determined as the calibration value.

5. The method according to claim 4, characterized in that, The reliability of the resource data of the target resource items in each source data table includes consistency and timeliness; the consistency of the resource data of the target resource items in each source data table represents the degree of matching between the resource data of the target resource items in each source data table; The timeliness of the resource data for the target resource category in each source data table represents the update speed of the resource data for the target resource category in each source data table; The step of determining the authority of the resource data of the target resource subject in each source data table based on the trust level of each source data table and the reliability of the resource data of the target resource subject in each source data table includes: The product of the trust level of each source data table and the timeliness of the resource data of the target resource subject in each source data table is determined as the trust level weighted value corresponding to each source data table. The authority of the resource data for the target resource subject in each source data table is determined by summing the weighted trust value corresponding to each source data table and the consistency of the resource data for the target resource subject in each source data table.

6. The method according to claim 5, characterized in that, The method further includes: Based on the importance of the timeliness of the resource data of the target resource subject in each source data table, a timeliness sensitivity coefficient for the resource data of the target resource subject in each source data table is determined; the timeliness sensitivity coefficient for the resource data of the target resource subject in each source data table is used to quantify the degree of influence of the timeliness of the resource data of the target resource subject in each source data table on the authority. The modified timeliness of the resource data of the target resource subject in each source data table is determined by using the timeliness of the resource data of the target resource subject in each source data table as the base and the timeliness sensitivity coefficient of the resource data of the target resource subject in each source data table as the exponent.

7. The method according to claim 5 or 6, characterized in that, The method further includes: Obtain a predetermined smoothing factor; the smoothing factor is used to balance the impact of the timeliness fluctuation of resource data of target resource subjects in each source data table on the authority. The timeliness of the resource data of the target resource item in each source data table and the sum of the smoothing factor are determined as the modified timeliness of the resource data of the target resource item in each source data table.

8. A data table cleaning device, characterized in that, include: The acquisition module is used to acquire a multi-source data table of a specified object; the multi-source data table refers to a resource data table about the specified object collected from different data sources; Each source data table in the multi-source data table includes resource data for different resource categories; The trust confirmation module is used to determine the trust level of each source data table based on the data backtracking chain of each source data table in response to the presence of dirty data in the same target resource subject in the multi-source data tables of a specified object. The data backtracking chain of each source data table refers to the flow path of each source data table. The data backtracking chain of each source data table is used to trace the original source of each source data table in reverse. The data cleaning module is used to select calibration values ​​from the same target resource subject in the multi-source data tables according to the trust level of each source data table, and use the calibration values ​​to replace the dirty data of the same target resource subject in the multi-source data tables.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the method according to any one of claims 1 to 7.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Data processing method, device and system and storage medium

    CN116910050A

  • Multi-data-source data verification method and system, computer equipment and storage medium

    CN118484447A

  • System and method for auditing and supervising real-time traceability of multi-source data

    CN120632744A

  • Multi-source data processing and form generation method and device, equipment and medium

    CN120745576A

  • Evaluating a trust value of a data report from a data processing tool

    US20130080197A1