Accounting data intelligent correction method and system based on data difference identification
By using a composite feature extraction method based on data profiling and business rule injection, and a multi-objective optimization correction decision mechanism, the problems of insufficient adaptability and imbalance in correction decisions in accounting data processing are solved, achieving efficient and accurate correction of accounting data.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-11
- Publication Date
- 2026-03-27
AI Technical Summary
In existing accounting data processing, data standardization and difference identification are not well adapted, the extraction of difference features is too simplistic and cannot fully capture complex differences, and the correction decision lacks multi-dimensional dynamic balancing, resulting in low correction efficiency and insufficient accuracy.
By using a differential composite feature extraction method based on data profiling and business rule injection, combined with a multi-objective optimization and correction decision mechanism, a high degree of matching between data standardization and business scenarios is achieved. It dynamically extracts features of numerical deviation, time series anomalies, and business logic conflicts, and makes cost-risk-efficiency optimization decisions to generate the optimal correction strategy.
It significantly improves the comprehensiveness and accuracy of accounting data discrepancy identification, realizes the scientific nature of accounting data correction and the rationality of resource utilization, and meets the high requirements of modern accounting business for data quality and processing efficiency.
Smart Images

Figure CN121743683A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent data correction technology, and in particular to an intelligent accounting data correction method and system based on data difference identification. Background Technology
[0002] In the field of accounting data processing, data standardization and discrepancy identification are core aspects of ensuring data quality, but existing technologies generally have significant shortcomings. Traditional methods often use a single standardization rule to process accounting data uniformly, without combining the inherent attributes of the data to build data profiles. This results in insufficient adaptability of standardization rules to data types, making it difficult to accurately match the needs of different business scenarios. At the same time, discrepancy feature extraction is often limited to a single dimension, lacking deep integration of business rules, and failing to comprehensively capture complex discrepancies such as numerical deviations, time series anomalies, and business logic conflicts. This significantly reduces the comprehensiveness and accuracy of discrepancy identification, thereby affecting the targeted nature of subsequent correction operations.
[0003] Furthermore, existing accounting data correction decision-making processes often focus only on optimizing a single objective, failing to establish a multi-dimensional dynamic trade-off mechanism for cost, risk, and efficiency. Traditional correction strategies often rely on fixed rules or manual experience, ignoring key factors such as differences in data importance and historical correction success rates. This leads to unreasonable resource allocation during the correction process, either excessively pursuing efficiency while ignoring correction risks, or increasing unnecessary correction costs to reduce risks. It is impossible to achieve the optimal balance between correction effectiveness and resource consumption. These shortcomings result in low efficiency and insufficient accuracy in accounting data correction, making it difficult to meet the high requirements of modern accounting operations for data quality and processing efficiency. Therefore, how to comprehensively and accurately identify differences and improve the accuracy of data correction has become an urgent problem to be solved. Summary of the Invention
[0004] This invention provides an intelligent correction method and system for accounting data based on data difference identification, in order to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, the present invention provides an intelligent accounting data correction method based on data difference identification, comprising: S1, based on accounting data and preset business rules, simultaneously performs data standardization and data profile construction, and outputs standardized accounting data with data profile tags; S2, perform business rule-injected feature extraction on the standardized accounting data to obtain a composite feature vector of differences including numerical deviations, time series anomalies and business logic conflicts; S3, Based on the difference composite feature vector, perform multi-level confidence classification on the accounting data and generate difference type labels with decision path descriptions; S4. Based on the difference type label, perform cost-risk-efficiency optimization and correction decisions on the standardized accounting data to obtain the optimal correction strategy instruction and initial draft of the traceability record for the accounting data. S5. Based on the optimal correction strategy instruction and the initial draft of the traceability record, the standardized accounting data is subject to targeted correction process control, and the corrected accounting data and complete operation traceability log are generated. S6. Based on the complete operation traceability log, search for similar cases from the historical traceability log, perform case-assisted verification on the corrected accounting data, and generate the correction verification results and rule optimization suggestions for the accounting data. S7. Based on the correction verification results, output the corrected accounting data and the corresponding complete correction traceability report.
[0006] In a preferred embodiment, the step of simultaneously performing data standardization and data profiling based on original accounting data and preset business rules, and outputting standardized accounting data with data profiling tags, includes: Extract the inherent attributes of accounting data to obtain an attribute set; The various indicators in the attribute set are weighted and combined to output labels that represent data behavior patterns, thereby constructing a data profile and obtaining the data profile labels for the accounting data. Based on the data profile tags, a subset of rules applicable to the current data type is selected from the preset business rule base to dynamically select the appropriate standardized rule set, thus obtaining the selected rule set; Based on the selected rule set, the accounting data is subjected to data standardization processing to obtain standardized data; The data profile tags are attached to the standardized data to output standardized data with data profile tags.
[0007] In a preferred embodiment, the step of performing business rule-injected feature extraction on the standardized accounting data to obtain a composite feature vector containing numerical deviations, time series anomalies, and business logic conflicts includes: Based on the data profile tags in the standardized accounting data, the corresponding feature calculation models are called and integrated into a feature calculation model set; Based on the feature calculation model set, the numerical deviation features of the standardized accounting data are extracted by calculating the absolute deviation value; Based on the feature calculation model set and the standardized accounting data, a time series anomaly feature extraction operation is performed to obtain time series anomaly features; Based on the feature calculation model set and the standardized accounting data, a business logic conflict feature extraction operation is performed to obtain business logic conflict features; Based on the numerical deviation characteristics, time series anomaly characteristics, and business logic conflict characteristics, feature vector combination operations are performed to obtain differential composite feature vectors.
[0008] In a preferred embodiment, the step of performing multi-level confidence classification on the accounting data based on the difference composite feature vector and generating difference type labels with decision path descriptions includes: The differential composite feature vector is scaled and mapped to a uniform numerical range to obtain a standardized feature vector; A multi-level classification initial confidence analysis is performed on the standardized feature vectors to obtain preliminary classification results and confidence scores; Based on the preliminary classification results and the confidence scores, a decision path backtracking operation is performed on the accounting data to generate difference type labels with accompanying decision path descriptions.
[0009] In a preferred embodiment, the step of making cost-risk-efficiency optimization and correction decisions on the standardized accounting data based on the difference type labels to obtain the optimal correction strategy instructions and a draft of the traceability record for the accounting data includes: Data importance indicators are extracted from the difference type labels and the data profile labels in the standardized accounting data to obtain a quantitative value of data importance; Query the historical correction operation records corresponding to the difference type labels, filter out past correction records with similar difference types, and extract their success rates and related parameters to obtain a historical success rate dataset; Based on the data importance quantification value, the historical success rate dataset, and the difference type label, a cost-risk-efficiency optimization decision algorithm is used to perform dynamic trade-off calculations to obtain the optimal correction strategy instruction for the accounting data. Based on the optimal correction strategy instruction and the difference type label, the decision parameters, strategy selection and difference type information are integrated to generate a draft of the source tracing record.
[0010] In a preferred embodiment, the mathematical expression of the cost-risk-efficiency optimization decision algorithm is as follows: In the formula, S is the overall score, C is the cost factor, R is the risk factor, and E is the efficiency factor. These are the weighting coefficients for cost, risk, and efficiency, respectively.
[0011] In a preferred embodiment, the process of controlling the targeted correction of the standardized accounting data based on the optimal correction strategy instruction and the initial draft of the traceability record, and generating corrected accounting data and a complete operation traceability log, includes: Based on the optimal correction strategy instruction, a targeted correction operation is performed on the standardized accounting data to obtain intermediate correction data; The intermediate calibration data and calibration operations are bound to operation records. All operation steps, parameters and results in the calibration process are associated with the intermediate calibration data and logged to obtain a complete operation traceability log. Based on the complete operation traceability log, corrected accounting data is generated.
[0012] In a preferred embodiment, the traceability correction and verification module is used to search for similar cases from historical traceability logs based on the complete operation traceability log, perform case-assisted verification on the corrected accounting data, and generate correction and verification results and rule optimization suggestions for the accounting data, including: Based on the complete operation traceability log and the historical traceability log, the log features are compared, and historical cases with high similarity to the current correction process are selected for similar case search to obtain a similar case set. The similar case set and the corrected accounting data are subjected to case-assisted verification operations to evaluate the degree of consistency between the corrected accounting data and historical successful cases in terms of data consistency and business compliance, so as to obtain the correction verification results; Based on the correction and verification results, a rule optimization suggestion generation operation is performed to generate rule optimization suggestions for the accounting data.
[0013] In a preferred embodiment, the step of outputting the corrected accounting data and the corresponding complete correction traceability report based on the correction verification results includes: Based on the correction verification results, the final state of the corrected accounting data is determined, and the verified corrected data is obtained. A complete correction and tracing report is constructed based on the verified correction data and the complete operation tracing log.
[0014] To address the aforementioned problems, the present invention also provides an intelligent accounting data correction system based on data difference identification, the system comprising: The multi-dimensional standardization processing module is used to simultaneously perform data standardization and data profile construction based on accounting data and preset business rules, and output standardized accounting data with data profile tags. The difference feature extraction module is used to perform business rule-injected feature extraction on the standardized accounting data to obtain a composite feature vector of differences that includes numerical deviations, time series anomalies, and business logic conflicts. The difference type classification module is used to perform multi-level confidence classification on the accounting data based on the difference composite feature vector and generate difference type labels with decision path descriptions; The correction strategy decision module is used to make cost-risk-efficiency optimization correction decisions on the standardized accounting data based on the difference type labels, and to obtain the optimal correction strategy instruction and the initial draft of the traceability record for the accounting data. The traceable targeted correction module is used to manage the targeted correction process of the standardized accounting data based on the optimal correction strategy instruction and the initial draft of the traceability record, and to generate corrected accounting data and a complete operation traceability log of the accounting data. The traceability correction and verification module is used to search for similar cases from historical traceability logs based on the complete operation traceability logs, perform case-assisted verification on the corrected accounting data, and generate the correction and verification results and rule optimization suggestions for the accounting data. Output feedback module: Based on the correction verification results, output the corrected accounting data and the corresponding complete correction traceability report.
[0015] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention significantly improves the comprehensiveness and accuracy of accounting data difference identification through a "difference composite feature extraction method based on data profiling and business rule injection". This method first constructs a data profile based on the inherent attributes of accounting data, dynamically selects a suitable standardized rule set from a preset business rule base to ensure that data standardization is highly matched with business scenarios, and then calls the corresponding feature calculation model set with the data profile tags as a guide to simultaneously extract three core features: numerical deviation, time series anomalies, and business logic conflicts, and combine them into a difference composite feature vector. Compared with traditional single-dimensional feature extraction, this method achieves deep integration of business rules and feature extraction. It can accurately capture multi-dimensional differences in data at the numerical, time series, and logical levels, and avoid the disconnect between feature extraction and data type and business scenario through dynamic matching of data profile and model. This provides a more comprehensive and targeted basis for subsequent correction operations and greatly reduces correction deviations caused by omissions or misjudgments of differences.
[0016] 2. This invention achieves both the scientific nature of accounting data correction and the rationality of resource utilization through a "multi-objective optimization correction decision-making mechanism that integrates cost, risk, and efficiency." This mechanism first extracts quantitative values of data importance from difference type labels and data profile labels. Combined with a historical correction success rate dataset of similar difference types, it performs dynamic trade-off calculations through a cost-risk-efficiency optimization decision-making algorithm. Compared to traditional single-objective correction decisions, this mechanism not only considers the three core elements of correction operation cost consumption, failure risk, and processing efficiency, but also introduces key influencing factors such as data importance and historical experience. Through weight allocation, it achieves a dynamic balance of multiple objectives. It can generate optimal correction strategy instructions that balance correction effectiveness and resource consumption for accounting data of different importance and different types of differences. This mechanism effectively solves the problems of unreasonable resource allocation and objective imbalance in traditional correction. While ensuring correction accuracy, it reduces unnecessary cost input, improves overall correction efficiency, and meets the dual demands of modern accounting operations for data correction quality and processing efficiency. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating an intelligent accounting data correction method based on data difference identification, provided in an embodiment of the present invention. Figure 2 A functional block diagram of an intelligent accounting data correction system based on data difference identification, provided in an embodiment of the present invention; The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0018] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0019] This application provides an intelligent accounting data correction method based on data difference identification. The executing entity of this intelligent accounting data correction method based on data difference identification includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this application: a server, a terminal, etc. In other words, the intelligent accounting data correction method based on data difference identification can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0020] Reference Figure 1 The diagram shown is a flowchart illustrating an intelligent accounting data correction method based on data difference identification, according to an embodiment of the present invention. In this embodiment, the intelligent accounting data correction method based on data difference identification includes: S1, based on accounting data and preset business rules, simultaneously performs data standardization and data profile construction, and outputs standardized accounting data with data profile tags; In this embodiment of the invention, the step of simultaneously performing data standardization and data profiling based on original accounting data and preset business rules, and outputting standardized accounting data with data profiling tags, includes: Extract the inherent attributes of accounting data to obtain an attribute set; The various indicators in the attribute set are weighted and combined to output labels that represent data behavior patterns, thereby constructing a data profile and obtaining the data profile labels for the accounting data. Based on the data profile tags, a subset of rules applicable to the current data type is selected from the preset business rule base to dynamically select the appropriate standardized rule set, thus obtaining the selected rule set; Based on the selected rule set, the accounting data is subjected to data standardization processing to obtain standardized data; The data profile tags are attached to the standardized data to output standardized data with data profile tags.
[0021] It should be noted that the inherent attributes of the extracted data are based on the structured fields of the original accounting data. Key features such as data category, frequency of historical changes, and strength of related accounts are identified through attribute parsing operations to form an attribute set.
[0022] Furthermore, the essence of attribute parsing is to use statistical analysis and relational mapping techniques to quantitatively extract descriptive indicators from accounting data. For example, data category represents the classification code of accounting subjects, historical change frequency reflects the update rate of data in the time dimension, and related subject strength measures the business correlation between different accounting subjects.
[0023] It should be noted that the attribute set is a structured summary of the inherent attributes of the original accounting data, used for subsequent data profiling and serving to provide a basis for dynamic rule selection.
[0024] It should be noted that the generation of data profiles is based on a set of attributes. Data features are comprehensively calculated through a profile building algorithm to generate data profile labels.
[0025] Furthermore, the essence of the profile building algorithm is to use multi-dimensional feature fusion technology to weight and combine various indicators in the attribute set to output labels that represent data behavior patterns. For example, high-frequency transaction flow corresponds to streaming processing labels, and fixed asset data corresponds to batch verification labels. In the process of weighting and combining, the weight coefficients of each indicator are evenly distributed and add up to 1.
[0026] It should be noted that data profile tags are abstract representations of inherent data attributes, used to guide the adaptive selection of standardized rules.
[0027] It should be noted that the dynamic selection of the standardized rule set is based on data profile tags. The selected rule set is obtained by filtering a subset of rules applicable to the current data type from the preset business rule library through rule matching operations.
[0028] Furthermore, the essence of rule matching operation is to use a tag-driven query mechanism to map data profile tags to corresponding standardized rules. For example, streaming standardized rules are used for high-frequency transaction data, and batch verification rules are used for fixed asset data.
[0029] It should be noted that the selected rule set is a combination of rules from the preset business rules that match the data profile tags, used to ensure data standardization and adaptability to business scenarios.
[0030] It should be noted that data standardization processing is based on a selected set of rules, which involves performing format conversion, value range verification, and structure alignment operations on the original accounting data to eliminate data inconsistencies and obtain standardized data.
[0031] Furthermore, the format conversion operation includes unified data encoding and unit standardization, the value range verification operation verifies whether the data value range conforms to business specifications, and the structure alignment operation adjusts the data fields to match the target pattern.
[0032] It should be noted that standardized data is accounting data that has undergone normalization processing. It is clean data after eliminating ambiguity in the original data, and is used to provide consistent input for subsequent discrepancy identification.
[0033] It should be noted that attaching data profile tags to standardized data is done by using a tag binding operation to associate the data profile tags with the standardized data, so as to output standardized data with data profile tags attached.
[0034] Furthermore, the tag binding operation is implemented by writing data profile tags as additional attributes into extended fields of standardized data.
[0035] It should be noted that standardized data with data profile tags is normalized data that integrates data feature identifiers. It is a standardized output carrying business semantics and is used to support dynamic routing in subsequent feature extraction and classification processes.
[0036] S2, perform business rule-injected feature extraction on the standardized accounting data to obtain a composite feature vector of differences including numerical deviations, time series anomalies and business logic conflicts; In this embodiment of the invention, the step of performing business rule-injected feature extraction on the standardized accounting data to obtain a composite feature vector containing numerical deviations, time series anomalies, and business logic conflicts includes: Based on the data profile tags in the standardized accounting data, the corresponding feature calculation models are called and integrated into a feature calculation model set; Based on the feature calculation model set, the numerical deviation features of the standardized accounting data are extracted by calculating the absolute deviation value; Based on the feature calculation model set and the standardized accounting data, a time series anomaly feature extraction operation is performed to obtain time series anomaly features; Based on the feature calculation model set and the standardized accounting data, a business logic conflict feature extraction operation is performed to obtain business logic conflict features; Based on the numerical deviation characteristics, time series anomaly characteristics, and business logic conflict characteristics, feature vector combination operations are performed to obtain differential composite feature vectors.
[0037] It should be noted that the corresponding feature calculation model is invoked based on the data profile label. The feature calculation model corresponding to the data profile label is matched from the preset feature model library through the model selection operation to form a feature calculation model set.
[0038] Furthermore, during the matching feature calculation model process, based on data profile tags, a tag-driven routing mechanism dynamically selects the model corresponding to the data type from the preset feature model library. Specific feature models include streaming numerical deviation models for high-frequency transaction flow data and batch time series models for fixed asset data. These models achieve adaptive selection by mapping data profile tags, such as "streaming processing tags" or "batch verification tags", to ensure that feature extraction is adapted to the business scenario.
[0039] Furthermore, the essence of the model selection operation is to use a label-driven routing mechanism to dynamically map data profile labels to applicable feature calculation models. For example, for high-frequency transaction flow data, a streaming numerical deviation model is called, and for fixed asset data, a batch time series model is called.
[0040] It should be noted that the feature calculation model set is a collection of feature extraction algorithms that match data profile tags. It is a combination of configuration parameters for the feature extraction process, used to ensure the adaptability of feature extraction to data types and business scenarios, and to improve the accuracy and efficiency of feature extraction.
[0041] It should be noted that the numerical deviation feature extraction operation is based on the streaming numerical deviation model in the feature calculation model set. It calculates the numerical deviation of standardized accounting data to quantify the degree of deviation of data values from the expected range and obtains numerical deviation features.
[0042] Furthermore, the essence of numerical deviation calculation is the operation of calculating the absolute deviation between the actual value and the standard value, that is, the absolute value of the difference between the actual value and the standard value, which is used to identify data entry or calculation errors. In the numerical deviation feature extraction operation, the actual value refers to the specific value in standardized accounting data, such as the amount, and the standard value refers to the expected value defined based on business rules or historical data, such as the standard cost. When calculating the absolute deviation, the absolute difference between the actual value and the standard value is used to quantify the degree of data deviation.
[0043] It should be noted that numerical deviation characteristics are quantitative indicators of data values deviating from the normal range. They are measures of the degree of numerical anomaly and are used to provide numerical evidence of anomalies for subsequent difference classification.
[0044] It should be noted that the time series anomaly feature extraction operation is a processing method based on batch time series models to detect anomalies in standardized accounting data in order to identify abnormal patterns in data points over time.
[0045] Furthermore, the anomaly detection operation first calculates the mean and standard deviation of the time series data, then calculates the absolute deviation of each data point value from the mean, and finally divides the absolute deviation by the standard deviation to obtain a standardized anomaly score, which is used to identify anomalies that exceed a preset threshold.
[0046] It should be noted that time series anomaly features are identifiers of abnormal fluctuations in time series data. They are pattern features of data anomalies in the time dimension and are used to detect periodic or sudden anomalies, enhancing the time-series sensitivity of difference identification.
[0047] It should be noted that the business logic conflict feature extraction operation is a process of verifying standardized accounting data against business rules to identify whether the data violates the preset business logic.
[0048] Furthermore, business rule verification operations include checking the balance of accounting subjects, the continuity of transaction flows, or the consistency of related data, such as verifying whether the debit and credit amounts are balanced, or detecting whether there are gaps or duplicates in the transaction flow.
[0049] It should be noted that business logic conflict characteristics are indicators of the severity of data violations of business rules. They are a measure of the violation of business logic consistency and are used to ensure that data conforms to business specifications and prevent the propagation of logical errors.
[0050] It should be noted that the feature vector combination operation is based on numerical deviation features, time series anomaly features, and business logic conflict features. It integrates multiple features into a single vector through vector concatenation to obtain a differential composite feature vector.
[0051] Furthermore, the vector concatenation operation connects the feature values in a preset order to form a high-dimensional vector. The preset order is: first, numerical deviation features, then time series anomaly features, and finally business logic conflict features. This order is designed based on feature priority to ensure that numerical anomalies are processed first, thereby enhancing the interpretability of the vector in the classification model.
[0052] S3, Based on the difference composite feature vector, perform multi-level confidence classification on the accounting data and generate difference type labels with decision path descriptions; In this embodiment of the invention, the step of performing multi-level confidence classification on the accounting data based on the difference composite feature vector and generating difference type labels with decision path descriptions includes: The differential composite feature vector is scaled and mapped to a uniform numerical range to obtain a standardized feature vector; A multi-level classification initial confidence analysis is performed on the standardized feature vectors to obtain preliminary classification results and confidence scores; Based on the preliminary classification results and the confidence scores, a decision path backtracking operation is performed on the accounting data to generate difference type labels with accompanying decision path descriptions.
[0053] It should be noted that the differential composite feature vector is scaled and mapped to a uniform numerical range by using the minimum-maximum scaling method to map the feature values to the range of [0,1], so as to eliminate the difference in feature dimensions and obtain a standardized feature vector.
[0054] Furthermore, the mathematical expression for the min-max scaling method is as follows: In the formula, Here, X represents the standardized feature values, and X represents the original feature values. The minimum value of the characteristic. This represents the maximum value of the characteristic.
[0055] Furthermore, the essence of the min-max scaling method is to adjust the original feature values to the [0,1] interval through linear transformation, in order to ensure that different features are comparable in the classification model and improve the stability and accuracy of classification.
[0056] It should be noted that the standardized feature vector is a feature vector that has undergone normalization processing. It is a feature representation after eliminating differences in dimensions and is used to provide consistent input for multi-level classification operations.
[0057] It should be noted that the multi-level classification initial confidence analysis is a process of first performing multi-level classification operations on the standardized feature vectors, and then calculating the confidence score based on the classification results. The multi-level classification operation first makes conditional judgments at the decision nodes of the decision tree model based on the feature values. For example, it compares whether the numerical deviation features exceed the preset 95th percentile threshold. Then, it gradually guides the data to the leaf nodes, with each leaf node corresponding to a difference type, to obtain the initial classification result.
[0058] Furthermore, the confidence score is calculated by dividing the number of correctly classified samples in the leaf nodes by the total number of samples in the leaf nodes.
[0059] It should be noted that the preliminary classification results are used to identify the possible categories of discrepancies in accounting data. The confidence score is a quantitative value of the confidence level of the classification results, representing a probabilistic measure of the reliability of the classification. It is used to assess the reliability of the classification results and to provide a basis for subsequent decision-making.
[0060] It should be noted that the decision path backtracking operation is based on the preliminary classification results and confidence scores. By backtracking the decision-making process of the classification model, a textual description of the classification logic is generated to obtain the decision path description.
[0061] Furthermore, the essence of decision path backtracking is to extract the importance of decision rules or features from the classification model and form an interpretable path description. For example, in the generated decision tree model, the path from the root node to the leaf node is the decision path.
[0062] It should be noted that the difference type label is a classification result with accompanying decision path description. It represents the identifier of the type of accounting data difference and the textual description of the classification reason, which is used to provide interpretable classification output and support subsequent correction decisions.
[0063] S4. Based on the difference type label, perform cost-risk-efficiency optimization and correction decisions on the standardized accounting data to obtain the optimal correction strategy instruction and initial draft of the traceability record for the accounting data. In this embodiment of the invention, the step of performing cost-risk-efficiency optimization and correction decisions on the standardized accounting data based on the difference type labels to obtain the optimal correction strategy instruction and initial draft of the traceability record for the accounting data includes: Data importance indicators are extracted from the difference type labels and the data profile labels in the standardized accounting data to obtain a quantitative value of data importance; Query the historical correction operation records corresponding to the difference type labels, filter out past correction records with similar difference types, and extract their success rates and related parameters to obtain a historical success rate dataset; Based on the data importance quantification value, the historical success rate dataset, and the difference type label, a cost-risk-efficiency optimization decision algorithm is used to perform dynamic trade-off calculations to obtain the optimal correction strategy instruction for the accounting data. Based on the optimal correction strategy instruction and the difference type label, the decision parameters, strategy selection and difference type information are integrated to generate a draft of the source tracing record.
[0064] It should be noted that the extracted data importance index is based on the difference type label and the data profile label. The importance assessment operation is used to quantify the criticality of the data in the business scenario to obtain the data importance quantification value.
[0065] Furthermore, the essence of importance assessment is to use a weighted scoring method, combining the degree of impact of the difference type on the business and the inherent attributes in the data profile, to calculate the data importance score.
[0066] Furthermore, the calculation of the data importance score first extracts the degree of impact of the difference type on the business from the difference type label and assigns an impact weight. Then, it extracts the inherent attributes from the data profile label and assigns attribute weights. Finally, it performs a weighted sum of the impact weights and attribute weights to obtain the data importance score, which is used to quantify the criticality of data in the business. In this score, both the impact weight and the attribute weight are equal. Among the degree of impact of the difference type on the business: numerical deviation is assigned 0.8 points, time series anomalies are assigned 0.5 points, and business logic conflicts are assigned 0.9 points.
[0067] It should be noted that the data importance quantification value is a numerical representation of the criticality of data in accounting operations. It is a metric for data correction priority and is used to provide a weighting basis for subsequent optimization decisions, ensuring that highly important data is processed first.
[0068] It should be noted that the query history is based on a preset traceability database. Historical correction cases that match the difference type labels are obtained through database retrieval operations to obtain the historical success rate dataset.
[0069] Furthermore, the essence of database retrieval operations is to perform structured queries, filter out past correction records with similar difference types, and extract their success rates and related parameters.
[0070] It should be noted that the historical success rate dataset is a statistical collection of the success rates of past correction operations. It serves as historical evidence of the effectiveness of correction strategies and is used to provide empirical data for dynamic trade-offs, thereby reducing decision-making risks.
[0071] It should be noted that the optimal correction strategy instruction is the correction method instruction selected after optimization decision-making. It represents the best operating guide after balancing cost, risk, and efficiency, and is used to guide the subsequent correction process to improve correction accuracy and resource utilization.
[0072] It should be noted that the initial draft of the traceability record is generated based on the optimal correction strategy instruction and the difference type label. A preliminary correction process document is created through the record construction operation to obtain the initial draft of the traceability record.
[0073] Furthermore, the essence of recording the construction operation is to integrate decision parameters, strategy selection, and difference type information to form a structured log draft.
[0074] It should be noted that the initial draft of the traceability record is a preliminary record of the corrective decision-making process and a basic document for the traceability of corrective operations. It is used to provide an initial framework for subsequent complete traceability, ensuring operational transparency and audit compliance.
[0075] In this embodiment of the invention, the mathematical expression of the cost-risk-efficiency optimization decision algorithm is as follows: In the formula, S is the overall score, C is the cost factor, R is the risk factor, and E is the efficiency factor. These are the weighting coefficients for cost, risk, and efficiency, respectively.
[0076] It should be noted that the cost factor represents the quantitative value of the resources required for the correction operation, the risk factor represents the quantitative value of the probability of correction failure, and the efficiency factor represents the quantitative value of the correction speed and processing capacity. The weight coefficients for cost, risk, and efficiency are all one-third.
[0077] Furthermore, resource consumption, failure cases, and processing speed data are extracted from the historical correction database. Subsequently, the human resource cost is normalized by minimum-maximum scaling to map the original value to the range [0,1] to obtain the cost factor. The failure case ratio is normalized to obtain the risk factor. The processing speed data is similarly normalized to obtain the efficiency factor.
[0078] It should be noted that the cost-risk-efficiency optimization decision algorithm is based on the data importance quantification value, the historical success rate dataset, and the difference type label. It calculates the comprehensive score of the correction strategy, and then selects the strategy with the highest comprehensive score as the optimal correction strategy instruction. This transformation uses sorting and selection logic to ensure that the instruction achieves a balance between cost, risk, and efficiency, and directly maps it to the correction operation parameters.
[0079] S5. Based on the optimal correction strategy instruction and the initial draft of the traceability record, the standardized accounting data is subject to targeted correction process control, and the corrected accounting data and complete operation traceability log are generated. In this embodiment of the invention, the process of controlling the targeted correction of the standardized accounting data based on the optimal correction strategy instruction and the initial draft of the traceability record, and generating corrected accounting data and a complete operation traceability log, includes: Based on the optimal correction strategy instruction, a targeted correction operation is performed on the standardized accounting data to obtain intermediate correction data; The intermediate calibration data and calibration operations are bound to operation records. All operation steps, parameters and results in the calibration process are associated with the intermediate calibration data and logged to obtain a complete operation traceability log. Based on the complete operation traceability log, corrected accounting data is generated.
[0080] It should be noted that the targeted correction operation is based on the optimal correction strategy instruction. It performs numerical adjustment, journal entry reconstruction and supplementary notes through the correction process to eliminate data differences and obtain intermediate corrected data. Among them, the numerical adjustment is based on the optimal correction strategy instruction. It identifies numerical errors in standardized accounting data, performs correction processing, and adjusts the erroneous values to the compliance range to eliminate data differences. Journal entry restructuring is based on the optimal correction strategy instruction to reorganize the logical structure of accounting entries, including analyzing debit and credit imbalances, optimizing account correspondences, and ensuring that the total debit and credit amounts of the restructured entries are equal to eliminate logical conflicts. The supplementary notes, based on the optimal correction strategy directive, add explanatory notes and metadata to the corrected accounting data, such as the reasons for the operation and the version of the rule, to provide complete operational tracing and business context, enhancing data transparency and audit compliance.
[0081] Furthermore, the essence of targeted correction operations is to systematically correct accounting data according to the parameters and rules in the strategy instructions, including adjusting numerical errors, reconstructing journal entry logic, and supplementing missing information.
[0082] It should be noted that intermediate correction data is accounting data that has undergone preliminary correction but has not been linked to traceability. Its physical meaning is the intermediate state of data after correction, which is used to provide clean and structurally complete input for binding operation records.
[0083] It should be noted that the operation record binding is based on the initial draft of the traceability record. It associates and logs all operation steps, parameters and results in the calibration process with intermediate calibration data to generate a complete operation traceability log.
[0084] Furthermore, the essence of operation record binding is to integrate operation time, operation type, execution parameters, and intermediate results to form a structured, traceable document.
[0085] It should be noted that the complete operation traceability log is a complete collection of records of the correction process. It is a multi-dimensional traceable document of the data correction history, used to support auditing, verification and subsequent analysis, and to ensure operational transparency and compliance.
[0086] It should be noted that the corrected accounting data is based on the complete operation traceability log. The intermediate corrected data is version marked and status updated to output the final corrected version, thus obtaining the corrected accounting data.
[0087] Furthermore, the generation of corrected accounting data includes data version control, status identification, and output formatting to ensure that the data is available for subsequent business processing.
[0088] It should be noted that the corrected accounting data is the final data that has been corrected and comes with complete traceability. It represents the output of the correction process and is used to provide clean, reliable and auditable accounting data for report generation or decision support.
[0089] S6. Based on the complete operation traceability log, search for similar cases from the historical traceability log, perform case-assisted verification on the corrected accounting data, and generate the correction verification results and rule optimization suggestions for the accounting data. In this embodiment of the invention, the source tracing correction and verification module is used to search for similar cases from historical source tracing logs based on the complete operation source tracing logs, perform case-assisted verification on the corrected accounting data, and generate correction and verification results and rule optimization suggestions for the accounting data, including: Based on the complete operation traceability log and the historical traceability log, the log features are compared, and historical cases with high similarity to the current correction process are selected for similar case search to obtain a similar case set. The similar case set and the corrected accounting data are subjected to case-assisted verification operations to evaluate the degree of consistency between the corrected accounting data and historical successful cases in terms of data consistency and business compliance, so as to obtain the correction verification results; Based on the correction and verification results, a rule optimization suggestion generation operation is performed to generate rule optimization suggestions for the accounting data.
[0090] It should be noted that the similar case search operation is based on the complete operation traceability log and historical traceability log. By comparing log features through similarity calculation, historical cases with high similarity to the current correction process are selected to obtain a similar case set.
[0091] Furthermore, the essence of similarity calculation is to use feature vector distance metric to calculate the degree of matching between the current source log and historical source log in terms of operation type, data features, and business context. The mathematical expression involved in the similarity calculation is as follows: In the formula, K is the similarity score, and D is the Euclidean distance between the feature vectors of the current source log and the historical source log; This expression maps the distance to the range [0,1] by comparing the degree of matching between operation type, data features, and business context, and is used to quantify similarity.
[0092] Furthermore, the Euclidean distance between the feature vectors of the current source log and the historical source log is calculated using the following formula: In the formula, The current source log number 1 eigenvalue, The first in the historical origin log 1 eigenvalue, is the dimension of the feature vector.
[0093] It should be noted that the similar case set is a collection of historical cases that are highly relevant to the current correction process. It is an aggregation of past successful correction experiences. Cases with a similarity score greater than 0.7 are identified as highly relevant and used to provide a reference benchmark for case-assisted verification, thereby improving the reliability of the verification.
[0094] It should be noted that the case-assisted verification operation is based on a set of similar cases and corrected accounting data. A consistency check algorithm is used to evaluate the degree of consistency between the corrected data and historical successful cases in terms of data consistency and business compliance in order to obtain the correction verification results.
[0095] Furthermore, the essence of the consistency check algorithm is to compare the numerical range, logical relationship, and business rule compliance of the corrected data with the successful patterns in similar cases, and to compare the degree of matching between the corrected data and historical successful cases in terms of numerical range, logical relationship, and business rule compliance. Specifically, this includes calculating the absolute value difference of numerical deviations, verifying the balance of borrowing and lending, such as the difference between the total amount of debits and the total amount of credits, and checking the continuity of transaction flow.
[0096] Furthermore, the compliance with business rules is checked by the business rules engine to see if the corrected data conforms to the preset business specifications, calculate the proportion of data that conforms to the rules, and generate a quantitative score, which is obtained by dividing the number of data points that conform to the rules by the total number of data points.
[0097] It should be noted that the calibration verification result is a quantitative assessment of the consistency between the calibrated data and historical cases. It is a verification indicator of the effectiveness of the calibration operation, used to confirm the correctness of the calibration result and provide a basis for rule optimization.
[0098] It should be noted that the rule optimization suggestion generation operation is based on the correction and verification results. By identifying potential defects and improvement points in the correction process, it generates optimization suggestions for business rules. Essentially, it extracts deviations from the verification results, maps them to the corresponding business rules, and proposes adjustment plans. Specifically, based on the deviations in the correction and verification results, it identifies potential defects in the correction process, such as frequent numerical errors or logical conflicts, maps them to the corresponding business rules, and proposes adjustment plans.
[0099] It should be noted that the rule optimization suggestions are guidance for improving business rules and are specific measures for optimizing the correction process. They are used to improve the accuracy and efficiency of future correction operations and prevent similar discrepancies from recurring.
[0100] S7. Based on the correction verification results, output the corrected accounting data and the corresponding complete correction traceability report.
[0101] In this embodiment of the invention, the step of outputting the corrected accounting data and the corresponding complete correction traceability report based on the correction verification result includes: Based on the correction verification results, the final state of the corrected accounting data is determined, and the verified corrected data is obtained. A complete correction and tracing report is constructed based on the verified correction data and the complete operation tracing log.
[0102] It should be noted that the final status confirmation operation is based on the correction verification results. The correctness and business compliance of the corrected accounting data are confirmed through data verification and status marking operations to obtain the verified correction data.
[0103] Furthermore, the essence of data verification and status marking operations is to compare the compliance of the passing indicators and business rules in the verification results, perform a final audit on the data, and attach a verification status label. The specific content of data verification and status marking operations includes verifying the integrity and correctness of the corrected data, such as verifying whether the numerical range is within the business specifications and whether the logical relationship is consistent, and marking the data status through status labels. The specific steps are as follows: First, compare the compliance of the passing indicators and business rules in the verification results; then, perform a final audit on the data and attach a verification status label to the data field.
[0104] It should be noted that verified and corrected data refers to corrected accounting data that has been finally confirmed and has a verification status. It represents the final reliable output of the correction process and is used to ensure that the data can be used for subsequent business processing or report generation, thus avoiding the circulation of unverified data.
[0105] It should be noted that the structured report generation process is based on verified correction data and complete operation traceability logs. Through document integration and format standardization, data, operation history, verification results, and rule optimization suggestions are combined to generate a complete correction traceability report.
[0106] Furthermore, the essence of document integration and format standardization is to extract the key attributes of the verified correction data, the time-series records of the complete operation traceability log, and the summary of the correction verification results, and then arrange them in a structured manner according to a preset template.
[0107] It should be noted that a complete calibration traceability report is a comprehensive document of the entire calibration process. It represents a readable summary of the data calibration history and verification results, and is used to provide a basis for audit trails, decision support, and process optimization, thereby enhancing data transparency and credibility.
[0108] like Figure 2 The diagram shown is a functional block diagram of an accounting data intelligent correction system based on data difference identification, provided by an embodiment of the present invention.
[0109] The intelligent accounting data correction system 100 based on data difference identification described in this invention can be installed in an electronic device. Depending on the functions implemented, the intelligent accounting data correction system 100 may include a multi-dimensional standardization processing module 101, a difference feature extraction module 102, a difference type classification module 103, a correction strategy decision module 104, a traceable directional correction module 105, a traceable correction verification module 106, and an output feedback module 107. The module described in this invention can also be called a unit, which refers to a series of computer program segments that can be executed by the processor of an electronic device and can perform a fixed function, stored in the memory of the electronic device.
[0110] In this embodiment, the functions of each module / unit are as follows: The multi-dimensional standardization processing module is used to simultaneously perform data standardization and data profile construction based on accounting data and preset business rules, and output standardized accounting data with data profile tags. The difference feature extraction module is used to perform business rule-injected feature extraction on the standardized accounting data to obtain a composite feature vector of differences that includes numerical deviations, time series anomalies, and business logic conflicts. The difference type classification module is used to perform multi-level confidence classification on the accounting data based on the difference composite feature vector, and generate difference type labels with decision path descriptions; The correction strategy decision module is used to make cost-risk-efficiency optimization correction decisions on the standardized accounting data based on the difference type label, and to obtain the optimal correction strategy instruction and the initial draft of the traceability record for the accounting data. The traceable directional correction module is used to manage the directional correction process of the standardized accounting data based on the optimal correction strategy instruction and the initial draft of the traceability record, and to generate corrected accounting data and a complete operation traceability log of the accounting data. The source tracing correction and verification module is used to search for similar cases from historical source tracing logs based on the complete operation source tracing logs, perform case-assisted verification on the corrected accounting data, and generate the correction and verification results and rule optimization suggestions for the accounting data. The output feedback module is used to output the corrected accounting data and the corresponding complete correction traceability report based on the correction verification results.
[0111] In the several embodiments provided by this invention, it should be understood that the disclosed methods and systems can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.
[0112] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0113] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0114] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0115] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0116] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. An accounting data intelligent correction method based on data difference recognition, characterized in that, The method comprises: S1, based on the accounting data and the preset business rules, synchronously performing data standardization and data portrait construction, and outputting standardized accounting data attached with data portrait labels; S2, performing business rule injection type feature extraction on the standardized accounting data to obtain a difference composite feature vector containing numerical deviation, time series anomaly and business logic conflict; S3, based on the difference composite feature vector, performing multi-level confidence classification on the accounting data to generate a difference type label with decision path description; S4, based on the difference type label, performing cost-risk-efficiency optimization correction decision on the standardized accounting data to obtain an optimal correction strategy instruction and a draft of the traceability record of the accounting data; S5, based on the optimal correction strategy instruction and the draft of the traceability record, performing directional correction process control on the standardized accounting data to generate corrected accounting data and a complete operation traceability log of the accounting data; S6, based on the complete operation traceability log, searching similar cases from historical traceability logs, performing case-assisted verification on the corrected accounting data to generate a correction verification result and a rule optimization suggestion of the accounting data; S7, based on the correction verification result, outputting the corrected accounting data and the corresponding complete correction traceability report.
2. The accounting data intelligent correction method based on data difference identification according to claim 1, characterized in that, The method comprises: extracting inherent attributes of data from the accounting data to obtain an attribute set; weighting and combining each index in the attribute set to output a label representing data behavior mode to construct a data portrait, so as to obtain a data portrait label of the accounting data; based on the data portrait label, filtering a rule subset suitable for the current data type from a preset business rule library to dynamically select an adaptive standardized rule set, and obtaining a selected rule set; based on the selected rule set, performing data standardization processing on the accounting data to obtain standardized data; attaching the data portrait label to the standardized data to output standardized data attached with the data portrait label.
3. The accounting data intelligent correction method based on data difference recognition of claim 1, wherein, The method comprises: based on the data portrait label in the standardized accounting data, calling a corresponding feature calculation model to integrate into a feature calculation model set; based on the feature calculation model set, extracting numerical deviation features of the standardized accounting data by calculating absolute deviation values; based on the feature calculation model set and the standardized accounting data, performing time series anomaly feature extraction operations to obtain time series anomaly features; based on the feature calculation model set and the standardized accounting data, performing business logic conflict feature extraction operations to obtain business logic conflict features; based on the numerical deviation features, time series anomaly features and business logic conflict features, performing feature vector combination operations to obtain a difference composite feature vector.
4. The accounting data intelligent correction method based on data difference recognition of claim 1, wherein, The multi-level confidence classification of the accounting data based on the difference composite feature vector generates a difference type label with decision path description, including: scaling and mapping the difference composite feature vector to a unified numerical range to obtain a standardized feature vector; performing multi-level classification initial confidence analysis on the standardized feature vector to obtain a preliminary classification result and a confidence score; based on the preliminary classification result and the confidence score, performing a decision path backtracking operation on the accounting data to generate a difference type label with decision path description.
5. The accounting data intelligent correction method based on data difference recognition of claim 1, wherein, The cost-risk-efficiency optimization correction decision of the standardized accounting data based on the difference type label obtains the optimal correction strategy instruction and the preliminary draft of the traceability record of the accounting data, including: extracting a data importance index from the difference type label and the data portrait label in the standardized accounting data to obtain a data importance quantitative value; querying the correction operation historical record corresponding to the difference type label, screening out the past correction records of similar difference types, and extracting their success rates and related parameters to obtain a historical success rate data set; based on the data importance quantitative value, the historical success rate data set, and the difference type label, performing dynamic trade-off calculation through a cost-risk-efficiency optimization decision algorithm to obtain the optimal correction strategy instruction of the accounting data; based on the optimal correction strategy instruction and the difference type label, integrating decision parameters, strategy selection, and difference type information to generate a preliminary draft of the traceability record.
6. The accounting data intelligent correction method based on data difference recognition of claim 5, wherein, The mathematical expression of the cost-risk-efficiency optimization decision algorithm is as follows: where S is the composite score, C is the cost factor, R is the risk factor, and E is the efficiency factor, are the weight coefficients for cost, risk, and efficiency, respectively.
7. The accounting data intelligent correction method based on data difference recognition of claim 1, wherein, Based on the optimal correction strategy instruction and the preliminary draft of the traceability record, the standardized accounting data is subjected to directional correction process control to generate corrected accounting data and complete operation traceability log of the accounting data, including: based on the optimal correction strategy instruction, performing directional correction operation on the standardized accounting data to obtain intermediate correction data; binding operation records of the intermediate correction data and correction operation, associating and logging all operation steps, parameters, and results in the correction process with the intermediate correction data to obtain a complete operation traceability log; based on the complete operation traceability log, generating corrected accounting data.
8. The accounting data intelligent correction method based on data difference recognition of claim 1, wherein, The traceability correction verification module is used to search for similar cases from historical traceability logs based on the complete operation traceability log, perform case-assisted verification on the corrected accounting data, and generate correction verification results and rule optimization suggestions of the accounting data, including: based on the complete operation traceability log and the historical traceability log, comparing log features to screen out historical cases with high similarity to the current correction process for similar case searching operation to obtain a similar case set; performing case-assisted verification operation on the similar case set and the corrected accounting data to evaluate the consistency and business compliance of the corrected accounting data with historical successful cases to obtain correction verification results; based on the correction verification results, performing rule optimization suggestion generation operation to generate rule optimization suggestions of the accounting data.
9. The accounting data intelligent correction method based on data difference recognition of claim 1, wherein, The outputting the corrected accounting data and the complete correction traceability report corresponding to the corrected accounting data based on the correction verification result comprises: determining a final state of the corrected accounting data based on the correction verification result to obtain verified corrected data; constructing a complete correction traceability report based on the verified corrected data and the complete operation traceability log.
10. An accounting data intelligent correction system based on data difference identification, used to implement the accounting data intelligent correction method based on data difference identification in claims 1-9, characterized in that, The system comprises: a multi-dimensional standardization processing module configured to perform data standardization and data portrait construction based on accounting data and preset business rules, and output standardized accounting data with data portrait labels; a difference feature extraction module configured to perform business rule injection feature extraction on the standardized accounting data to obtain a difference composite feature vector containing numerical deviation, time series anomaly and business logic conflict; a difference type division module configured to perform multi-level confidence classification on the accounting data based on the difference composite feature vector to generate difference type labels with decision path descriptions; a correction strategy decision module configured to perform cost-risk-efficiency optimization correction decision on the standardized accounting data based on the difference type labels to obtain optimal correction strategy instructions and draft traceability records of the accounting data; a traceable directional correction module configured to perform directional correction process control on the standardized accounting data based on the optimal correction strategy instructions and the draft traceability records to generate corrected accounting data and a complete operation traceability log of the accounting data; a traceability correction verification module configured to search for similar cases from historical traceability logs based on the complete operation traceability log, perform case-assisted verification on the corrected accounting data, and generate correction verification results and rule optimization suggestions of the accounting data; an output feedback module configured to output the corrected accounting data and the complete correction traceability report corresponding to the corrected accounting data based on the correction verification result.
Citation Information
Cited By
Knowledge distillation-based tax meeting difference automatic adjustment method
CN121998784A
A tax difference automatic adjustment method based on knowledge distillation
CN121998784B