Heterogeneous account subject adaptive data mapping method in audit supervision
Patent Information
- Application Number
- CN202611299997.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-26
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]本发明旨在解决现有技术过度依赖孤立文本字面特征导致非标准账目科目映射失真以及单向级联处理放大匹配偏差的问题
[0029](1)在审计监督中的异构账目科目自适应数据映射中,将科目名称与凭证分录流数据中的资金借贷流向共同作为科目对照依据,并根据科目层级调整两类信息的权重,对于表层科目,能够保留科目名称的判断作用;对于深层明细科目,则更多利用实际借贷关系进行判断,由此可减少企业自定义名称混乱、同一名称含义不一或者科目名称与实际用途不符造成的干扰,提高异构账套中科目对照的准确性,并使数据清洗结果更加稳定。
Smart Images

Figure CN122817847A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of financial data processing technology, and in particular relates to an adaptive data mapping method for heterogeneous account items in audit supervision. Background Technology
[0002] In the standardization and audit supervision system for financial accounting systems, the financial management software architecture used by different enterprises is not uniform. Each enterprise also sets its own account coding rules and detailed account names, resulting in significant differences in the coding and names of the same type of accounting account in different accounting systems. Traditional data standardization processing usually relies on a preset static rule base or basic text similarity algorithm to extract account numbers and account names from the source accounting data. Then, through character comparison and word frequency statistics, the enterprise's self-defined accounts are mapped to standard accounting accounts. In the existing financial audit data cleaning system, not only does the underlying system hardware storage environment or data carrier have inherent limitations in heterogeneous compatibility, but its upper-level software mapping control methods also face fundamental shortcomings. For example, the Chinese invention patent with publication number CN121579453A... The application discloses a multi-link financial account mapping method and system that integrates semantic models and rule matching. It adopts a serial strategy to align accounts through multiple links such as complete matching, financial rules, and semantic similarity. However, this technique relies on the literal text features of account names or isolated semantic model predictions. In its serial transmission mechanism, once a preceding link outputs a high-confidence misjudgment result when faced with a high-noise non-standard name, the subsequent links will immediately terminate, resulting in the inability to correct the cascade matching deviation. This multi-link design fails to deeply explore the inherent topological flow constraints of the underlying debit and credit balance relationship of financial vouchers and lacks a two-way dynamic verification closed loop between account comparison and auxiliary accounting identification. When faced with enterprise-customized names, its data cleaning convergence accuracy and anti-interference stability are still insufficient.
[0003] When processing non-standard accounting data, the names of self-defined accounts and auxiliary accounting items are often not standardized, and the same name may express different meanings. Existing technologies usually only establish mapping relationships based on the literal characteristics of account names or separately set rules, without making full use of the account associations formed by the flow of funds and the debit-credit balance in the voucher entries. Simply increasing the literal matching threshold will increase the number of account mismatches. Continuously expanding the static dictionary is also difficult to eliminate the ambiguity caused by non-standard naming, and may also bring rule conflicts and additional computing overhead. Therefore, although existing solutions add auxiliary rules to supplement the judgment, their processing still mainly relies on the matching of isolated texts, and it is difficult to fundamentally solve the mapping conflicts of heterogeneous account accounts.
[0004] Therefore, how to achieve adaptive comparison and consistency updates of non-standard accounting subjects by utilizing the fund flow relationship between subjects and through mutual verification of subject comparison and auxiliary accounting identification without over-reliance on the literal matching of subject names remains a problem that existing technologies need to solve. Summary of the Invention
[0005] The present invention aims to solve the problems of existing technologies that rely too much on isolated text literal features, resulting in distortion of non-standard account subject mapping and amplification of matching deviations by one-way cascading processing.
[0006] Text dispersion is characterized by the entropy of character information in the subject name to represent the degree of disorder in the naming of heterogeneous subjects. The higher the entropy value, the more chaotic the enterprise's custom naming.
[0007] The topology baseline parameter is the basic offset parameter for the normalization of directed edge weights and the calculation of topological correlation in the subject transaction network. The static baseline range is 0.45 to 0.55.
[0008] The time-decrease factor is a correction coefficient that decreases with the increase of the accounting cycle, used to reduce the impact of historical voucher data on the current topology correction;
[0009] Accounting feature tags are standardized identifiers extracted from auxiliary accounting items, including five categories: customers, suppliers, projects, departments, and fixed assets.
[0010] The multidimensional weighted fusion weight is the proportion of text similarity and topological similarity in the comprehensive similarity score, and the sum of the two is always 1;
[0011] Mapping consistency time series is a time series array consisting of the number of changes in the mapping relationship between accounts within a continuous accounting period;
[0012] Volatility variance is the variance of the number of changes over time, used to quantify the standardization stability of the accounting system;
[0013] The standard type space is a set of optional types for auxiliary accounting defined by standard accounting standards;
[0014] Topological correlation is the degree to which the upstream and downstream capital flows of two subject nodes overlap, i.e., the network flow topological similarity.
[0015] In this technical solution, the adaptive data mapping method for heterogeneous account items in audit supervision includes the following steps:
[0016] Step S101: Receive discrete account detail data from multiple heterogeneous accounting sets through the data acquisition interface module, and extract the corresponding voucher entry stream data.
[0017] Step S102: Extract the text features of heterogeneous account names from the discrete account details data and calculate the text feature similarity. At the same time, construct a multi-level account transaction network structure based on the directed fund lending flow between heterogeneous accounts in the voucher entry flow data, calculate the network flow topology similarity between different heterogeneous account nodes, and perform multi-dimensional weighted fusion of text feature similarity and network flow topology similarity to determine the standard account candidate comparison relationship.
[0018] Step S103: Analyze the detailed auxiliary accounting items in the voucher entry flow data, extract accounting feature labels, use the standard account candidate comparison relationship as a pre-fixed rigid boundary, and adaptively trim the candidate standard type space of the detailed auxiliary accounting items; at the same time, in response to the real-time detection frequency of accounting feature labels in the voucher entry flow data exceeding the set counting threshold, dynamically correct the topology benchmark parameters when calculating the network flow topology similarity in situ, and update the standard account candidate comparison relationship, generate the adaptive mapping status result from heterogeneous account accounts to standard accounts, and write it into the relational database storage medium.
[0019] Preferably, step S102 includes the following sub-steps: Step S1021, calculate the cosine distance between the text feature vectors of heterogeneous subject names in the discrete subject detail data as the text feature similarity; Step S1022, construct a multi-level subject transaction network structure with heterogeneous subjects as nodes and voucher lending relationships as directed edges, and calculate the topological correlation degree based on the flow distribution between nodes as the network flow topological similarity; Step S1023, weightedly fuse the text feature similarity and the network flow topological similarity to determine the standard subject candidate comparison relationship.
[0020] Preferably, after the adaptive mapping state result is written to the relational database storage medium, the following steps are also included: Step S301, read the standard enterprise business registration database from the external data source and extract the detailed accounting entity name and control relationship chain data from it; Step S302, compare the detailed accounting entity name with the detailed accounting entity name under the account that has been mapped in the adaptive mapping state result, and complete the normalized renaming; Step S303, establish a relationship topology structure of the same control group based on the control relationship chain data, convert the independent external enterprise customer and supplier names into structured network nodes containing equity control levels and related attributes, and use the relationship topology structure of the same control group to mark and verify the transaction flow of related parties during cross-account set flow verification and internal reconciliation.
[0021] Preferably, after the adaptive mapping state result is written to the relational database storage medium, the following steps are also included: Step S401, continuously record and store the iteration frequency data of the adaptive mapping state result in multiple consecutive financial accounting cycles to establish a mapping consistency time series; Step S402, analyze the mapping consistency time series and calculate the volatility variance value of the mapping consistency time series as a quantitative indicator of the trend of accounting standardization; Step S403, when the volatility variance value exceeds the set standardization threshold, generate a high-risk accounting set anomaly marking instruction and transmit the high-risk accounting set anomaly marking instruction to the background audit supervision interaction interface to switch the financial data of the corresponding accounting set to the penetration verification mode.
[0022] Preferably, in step S102, before weighted fusion of text feature similarity and network flow topology similarity, the text dispersion of the name character distribution in the discrete subject detail data is calculated first, and the fusion weight of network flow topology similarity is increased in response to the increase of text dispersion; wherein, the multidimensional weighted fusion weights of text feature similarity and network flow topology similarity are dynamically adjusted to the numerical ranges of 0.4 to 0.6 and 0.6 to 0.4 respectively, and the sum of the two weights is always equal to 1; and by comparing the convergence state, the comprehensive similarity judgment threshold of the determined standard subject candidate comparison relationship is set in the closed interval of 0.75 to 0.85.
[0023] Preferably, the method further includes the following steps: Step S601, traversing the voucher entry stream data, matching business keywords in the voucher summary stream, and extracting the pairing structure of debit account categories and credit account categories; Step S602, based on the preset transfer account correspondence rules as the pairing constraint rules for the pairing structure, classifying and identifying period-end transfer vouchers, cost transfer vouchers, and profit and loss transfer vouchers, and removing the corresponding voucher transaction amounts from the standardized accounting data; Step S603, capturing the coverage status change signal from the data source end, calling the incremental data synchronization algorithm, and writing the changed mapping status into the enterprise-level account relational database for synchronous flushing.
[0024] Preferably, the step S101 of obtaining discrete account detail data from multiple heterogeneous external accounting sets includes the following sub-steps: Step S1011, removing leading and trailing spaces, tabs, and non-printable implicit control characters from the original account data of multiple heterogeneous external accounting sets; Step S1012, identifying English letters embedded in the original account data, uniformly converting lowercase English letters to uppercase English letters, converting full-width characters to half-width characters, and outputting the cleaned standard text stream as discrete account detail data.
[0025] Preferably, the calculation of topological correlation degree based on the flow distribution between nodes in step S1022 includes the following sub-steps: Step S10221, normalize the weights of the directed edges in the constructed multi-level subject transaction network structure; Step S10222, for any two heterogeneous subject nodes, respectively count the set of downstream account nodes they both point to and the set of upstream account nodes they both originate from; Step S10223, calculate the topological correlation degree based on the overlap of fund flow between the set of downstream account nodes and the set of upstream account nodes.
[0026] Preferably, after the adaptive mapping status result is written to the relational database storage medium, the following steps are also included: Step S901, outputting the generated adaptive mapping status result and the corresponding discrete candidate standard subject comparison list to the interactive interface; Step S902, capturing manual confirmation instructions or batch overwrite instructions for large-scale comparison electronic forms from the input port, and updating and solidifying the adaptive mapping status result in the relational database storage medium according to the instructions.
[0027] Preferably, in the process of dynamically correcting the topology reference parameters when calculating the network flow topology similarity in situ, a step coefficient is used for correction, and the step coefficient gradually decreases with a negative exponential law as the calculation cycle progresses. The step coefficient used in the current calculation cycle is calculated in the following way: obtain the ordinal number of the current calculation cycle, calculate the difference between the ordinal number and the constant 1, multiply the obtained difference by a preset attenuation coefficient to obtain the attenuation exponent; calculate the constant power of the base of the natural logarithm when the attenuation exponent is negative, and use the obtained value as the time-effect attenuation factor; multiply the preset initial step coefficient by the time-effect attenuation factor to obtain the step coefficient of the current calculation cycle.
[0028] The present invention has the following beneficial effects:
[0029] (1) In the adaptive data mapping of heterogeneous accounts in audit supervision, the account name and the fund lending and borrowing flow in the voucher entry flow data are used as the basis for account comparison. The weight of the two types of information is adjusted according to the account level. For surface accounts, the judgment role of account name can be retained; for deep detailed accounts, the actual lending and borrowing relationship is used more to make judgment. This can reduce the interference caused by the confusion of enterprise custom names, different meanings of the same name, or the account name not matching the actual use, improve the accuracy of account comparison in heterogeneous accounts, and make the data cleaning results more stable.
[0030] (2) Enables mutual verification between subject comparison and auxiliary accounting identification. The standard subject candidate comparison relationship can limit the selection range of detailed auxiliary accounting items. The accounting feature label can also reverse the parameters on which the subject comparison is based, thereby eliminating irrelevant auxiliary accounting types, reducing the transmission of deviations from the previous processing stage to subsequent stages, reducing continuous mismatches caused by non-standard voucher summaries and ambiguous texts, and improving the anti-interference ability and processing stability of the heterogeneous account cleaning process.
[0031] (3) By unifying the names of detailed accounting entities and establishing the same control group relationship topology based on the control relationship chain, the identification ambiguity caused by enterprise abbreviations, custom names and different writing methods can be reduced, and the relationships of customers, suppliers and their related parties can be recorded in a unified manner. When checking cross-account flow and reconciling internal transactions, the transactions between related parties can be identified more accurately, reducing omissions, duplicate entries and verification errors caused by inconsistent names, and improving the traceability integrity of financial data under complex equity control relationships. Attached Figure Description
[0032] Figure 1 This is a flowchart of the heterogeneous accounting set subject candidate comparison and adaptive mapping process of the present invention;
[0033] Figure 2 This is a time decay curve of the step coefficient corresponding to the accounting cycle of this invention. Detailed Implementation
[0034] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0035] The adaptive data mapping method for heterogeneous account items in audit supervision includes the following steps:
[0036] Step S101: Receive discrete account detail data from multiple heterogeneous accounting sets through the data acquisition interface module, and extract the corresponding voucher entry stream data.
[0037] The boundaries for text dispersion are defined as follows: low value boundary = 0.3, high value boundary = 0.7. If the dispersion < 0.3, it is normalized to 0; if it > 0.7, it is normalized to 1. The normalization function for text dispersion is: The auxiliary accounting counting threshold is based on 50, and supports custom ranges of 40-80 for different industry accounting systems. The sliding time window is fixed at 12 accounting cycles and can only be adjusted within the range of 6-24. The risk of the accounting system is determined by the variance of fluctuations > 0.035 for three consecutive cycles, which triggers a penetration check. The keyword library for carry-over vouchers includes current year profit and carry-over profit and loss, cost carry-over includes production cost and completed goods received, and profit and loss carry-over includes main business cost and sales expenses. The subject text segmentation uses a two-character sliding window, and the benchmark dictionary includes all first-level and second-level standard subject names of the enterprise accounting standards. The industrial and commercial entity matching uses the unified social credit code as the unique matching identifier. When there is no code, the full name and administrative division are matched with double verification. Specifically, the enterprise full name, administrative division and industry suffix are split into three fields. If the similarity of the three fields is ≥0.8, it is determined to be the same entity. If any one of them is lower than 0.6, the match fails and the original name is retained without automatic renaming. The enterprise full name, administrative division and industry suffix are calculated using Stext cosine similarity, and the three scores are judged independently.
[0038] Step S102: Extract the text features of heterogeneous account names from the discrete account details data and calculate the text feature similarity. At the same time, construct a multi-level account transaction network structure based on the directed fund lending flow between heterogeneous accounts in the voucher entry flow data, calculate the network flow topology similarity between different heterogeneous account nodes, and perform multi-dimensional weighted fusion of text feature similarity and network flow topology similarity to determine the standard account candidate comparison relationship.
[0039] Step S103: Analyze the detailed auxiliary accounting items in the voucher entry flow data, extract accounting feature labels, use the standard account candidate comparison relationship as a pre-fixed rigid boundary, and adaptively trim the candidate standard type space of the detailed auxiliary accounting items; at the same time, in response to the real-time detection frequency of accounting feature labels in the voucher entry flow data exceeding the set counting threshold, dynamically correct the topology benchmark parameters when calculating the network flow topology similarity in situ, and update the standard account candidate comparison relationship, generate the adaptive mapping status result from heterogeneous account accounts to standard accounts, and write it into the relational database storage medium.
[0040] Preferably, step S102 includes the following sub-steps: Step S1021, calculate the cosine distance between the text feature vectors of heterogeneous subject names in the discrete subject detail data as the text feature similarity; Step S1022, construct a multi-level subject transaction network structure with heterogeneous subjects as nodes and voucher lending relationships as directed edges, and calculate the topological correlation degree based on the flow distribution between nodes as the network flow topological similarity; the voucher entry splitting rule is as follows: when a single voucher has multiple debits and credits, each group of credit subjects and debit subjects independently generates a directed edge, and the weight of each edge is taken as the absolute value of the entry amount corresponding to the lending pair, and repeated edges are accumulated: when the same credit or debit subject appears multiple times, the weight of the corresponding directed edge is accumulated by the absolute value of all the entry amounts; the raw weights of all outgoing edges of node X are calculated. w Total outflow from nodes Normalized border rights
[0041] Total inflow to node X Normalization into border rights The multi-level network layering rule is as follows: each subject independently constructs a single-layer subnet, and cross-level directed edges are added between subjects through subject subordinate codes. The weight of the cross-level edge is the sum of all occurrences of the lower-level subject. Step S1023: Weighted fusion of text feature similarity and network flow topology similarity determines the candidate comparison relationship of standard subjects.
[0042] Preferably, after the adaptive mapping state result is written to the relational database storage medium, the following steps are also included: Step S301, read the standard enterprise business registration database from the external data source and extract the detailed accounting entity names and control relationship chain data from it; Step S302, compare the detailed accounting entity names with the detailed accounting entity names under the accounts receivable and payable that have been mapped in the adaptive mapping state result, and complete the normalized renaming; Step S303, establish a relationship topology structure of the same control group based on the control relationship chain data, and convert the independent external enterprise customer and supplier names into a structured network containing equity control hierarchy and related attributes. The system utilizes the topology of the same controlling group relationship to mark and verify the flow of related party transactions during cross-account flow verification and internal reconciliation. It traces upwards along the directed equity edge to the ultimate parent company; companies under the same actual controller are identified as related parties. When the difference between the debit and credit amounts of related parties exceeds 0.5% of the total assets of the account set, it marks an abnormal transaction. Cross-subsidiary accounts automatically merge and offset internal transactions, outputting offset details for audit penetration verification. The absolute difference between debit and credit amounts of related accounts is divided by the corresponding account set's ending total assets; a ratio > 0.005 marks an abnormal related party transaction. The abnormal related party transaction ratio R = |Total debit transactions - Total credit transactions| ÷ Total assets of the account set at the end of the period; when R > 0.005, it marks an abnormal related party transaction. For transactions between two parties within the same controlling group, the smaller of the total debit and credit amounts is taken as the offset amount, and the difference is retained as the unoffset balance. The offset details field includes the subsidiary code, standard inter-company account, offset amount, remaining outstanding amount, and equity control level. Entities that fail to match are separately stored in a list awaiting manual review and output to the audit interaction interface, supporting batch manual binding of business entities.
[0043] Preferably, after the adaptive mapping state results are written to the relational database storage medium, the following steps are further included: Step S401, continuously record and store the iteration frequency data of the adaptive mapping state results in multiple consecutive financial accounting cycles to establish a mapping consistency time series; within a single accounting cycle, the total number of newly added and modified entries in the mapping relationship between heterogeneous accounts and standard accounts is the time series value for that cycle. The sliding window consists of 12 consecutive periods forming a sequence. Step S402: Analyze the mapping consistency time series and calculate the variance of the mapping consistency time series as a quantitative indicator of the trend of accounting standardization; specifically, first calculate the window mean. Fluctuation variance Step S403: When the variance of the fluctuation exceeds the set normalization threshold, a high-risk account set anomaly marking instruction is generated, and the high-risk account set anomaly marking instruction is transmitted to the background audit supervision interaction interface to switch the financial data of the corresponding account set to the penetration verification mode.
[0044] Preferably, in step S102, before weighted fusion of text feature similarity and network flow topology similarity, the text dispersion of the name character distribution in the discrete subject detail data is calculated first, and the fusion weight of network flow topology similarity is increased in response to the increase of text dispersion; wherein, the multidimensional weighted fusion weights of text feature similarity and network flow topology similarity are dynamically adjusted to the numerical ranges of 0.4 to 0.6 and 0.6 to 0.4 respectively, and the sum of the two weights is always equal to 1; and by comparing the convergence state, the comprehensive similarity judgment threshold of the determined standard subject candidate comparison relationship is set in the closed interval of 0.75 to 0.85.
[0045] Preferably, the method further includes the following steps: Step S601, traversing the voucher entry flow data, matching business keywords in the voucher summary flow, and extracting the pairing structure of debit account categories and credit account categories; Step S602, based on the preset carry-over account correspondence rules as the pairing constraint rules for the pairing structure, classifying and identifying period-end carry-over vouchers, cost carry-over vouchers, and profit and loss carry-over vouchers, and removing the corresponding voucher transaction flow from the standardized accounting data; vouchers identified as carry-over are not included in the directed edge weight accumulation only when constructing the account transaction network, and the original vouchers are completely retained in the underlying audit database. A voucher simultaneously contains carry-over entries. When recording and managing entries, only the debit / credit flows corresponding to the carry-forward entries are removed; the managing entries participate normally in the topology network calculation. The summary precisely matches the carry-forward keywords. The entry must belong to both profit / loss and cost carry-forward categories to be considered a carry-forward entry. When constructing the network, the directed edge weights corresponding to this entry are removed. Entries that only match keywords but whose accounts do not belong to the carry-forward category are not removed. Entries that only carry-forward accounts but whose summaries lack keywords are subject to manual review as a fallback and are not automatically removed. Voucher summaries use precise word segmentation matching; containing complete carry-forward keywords indicates a matching voucher type. Fuzzy word segmentation matching is only used as a fallback and is not directly removed. Entries are categorized by account type, including profit / loss and cost carry-forward entries. The entry line for this type of carry-forward account is determined to be a carry-forward entry, and the rest are operating entries; when constructing the account transaction network, only the directed edge weights corresponding to the carry-forward entry lines are ignored, and the weights of operating entry lines are accumulated normally; all original vouchers are stored in the audit database without deletion; in step S603, the overlay status change signal from the data source is captured, the incremental data synchronization algorithm is called, and the changed mapping status is written to the enterprise-level account relational database for synchronous flushing; the historical mapping relationship table in the database is read, and the generated mapping is compared with the heterogeneous account code as the unique key; newly added or changed entries are marked as incremental datasets, and database transaction batch writing is started; if it fails, it is rolled back. Change logs are recorded for audit traceability; the maximum number of mapping records that can be written in a single transaction is 1000; in the event of primary key conflicts or database IO exceptions, the entire batch of data will be rolled back without partial commits; fixed fields in the change log include log ID, accounting period, heterogeneous account code, original standard account code, updated standard account code, change type, operation source, and timestamp. The change type includes addition, modification, and manual locking, and the operation source includes automatic algorithm, manual batch, and single confirmation; when multiple update instructions exist for the same heterogeneous account code, they are overwritten according to instruction priority, with higher priority instructions overwriting lower priority old records.
[0046] Preferably, the step S101 of obtaining discrete account detail data from multiple heterogeneous external accounting sets includes the following sub-steps: Step S1011, removing leading and trailing spaces, tabs, and non-printable implicit control characters from the original account data of multiple heterogeneous external accounting sets; Step S1012, identifying English letters embedded in the original account data, uniformly converting lowercase English letters to uppercase English letters, converting full-width characters to half-width characters, and outputting the cleaned standard text stream as discrete account detail data.
[0047] Preferably, the calculation of topological correlation degree based on the flow distribution between nodes in step S1022 includes the following sub-steps: Step S10221, normalize the weights of the directed edges in the constructed multi-level subject transaction network structure; Step S10222, for any two heterogeneous subject nodes, respectively count the set of downstream account nodes they both point to and the set of upstream account nodes they both originate from; Step S10223, calculate the topological correlation degree based on the overlap of fund flow between the set of downstream account nodes and the set of upstream account nodes.
[0048] Preferably, after the adaptive mapping status result is written to the relational database storage medium, the following steps are also included: Step S901, outputting the generated adaptive mapping status result and the corresponding discrete candidate standard subject comparison list to the interactive interface; Step S902, capturing manual confirmation instructions or batch overwrite instructions of large-scale comparison electronic forms from the input port, and updating and solidifying the adaptive mapping status result in the relational database storage medium according to the instructions; the manual confirmation instructions have higher priority than the automatic mapping results, and the batch forms use heterogeneous subject codes and standard subject codes as primary keys to overwrite the original automatic mapping; the instruction priority order is: manual single confirmation instruction > batch form overwrite instruction > algorithm automatically generated mapping result, and when multiple instructions are triggered at the same time, they are executed in order of priority; the batch forms only overwrite the heterogeneous subject codes that exist in the form, and subjects that do not appear in the form retain the original mapping relationship; the manual solidification locking mechanism is that the mapping entries that have been manually confirmed or batch solidified are marked with a lock, and the subsequent periodic topology parameters are automatically corrected, and the algorithm recalculation must not overwrite the locked entries, and only the unlocked entries are automatically updated.
[0049] Preferably, in the process of dynamically correcting the topology reference parameters when calculating the network flow topology similarity in situ, a step coefficient is used for correction, and the step coefficient gradually decreases with a negative exponential law as the calculation cycle progresses. The step coefficient used in the current calculation cycle is calculated in the following way: obtain the ordinal number of the current calculation cycle, calculate the difference between the ordinal number and the constant 1, multiply the obtained difference by a preset attenuation coefficient to obtain the attenuation exponent; calculate the constant power of the base of the natural logarithm when the attenuation exponent is negative, and use the obtained value as the time-effect attenuation factor; multiply the preset initial step coefficient by the time-effect attenuation factor to obtain the step coefficient of the current calculation cycle.
[0050] Example 1: The data acquisition interface module receives discrete account detail data from multiple authorized heterogeneous accounting sets and extracts the corresponding voucher entry stream data. The data cleaning unit removes leading and trailing spaces, tabs, and non-printable implicit control characters from the discrete account detail data and voucher entry stream data, identifies embedded English letters, converts lowercase English letters to uppercase English letters, converts full-width characters to half-width characters, and outputs the cleaned standard text stream.
[0051] The logic operation unit extracts heterogeneous subject name text features from the standard text stream. It constructs heterogeneous subject name text feature vectors using character-level word frequency statistics and a bag-of-words model. The system pre-establishes a benchmark dictionary containing standard accounting terminology and common financial business vocabulary. The cleaned heterogeneous subject names are segmented into discrete character fragments using a two-character sliding window. A fixed-dimensional multidimensional sparse vector is then constructed based on the position index of each character fragment in the benchmark dictionary. The value at the corresponding index position in the vector is the number of times the character fragment appears in the current subject name; non-appearing character fragments have a value of 0. The logic operation unit calculates the cosine distance between the text feature vectors of different heterogeneous subject names and uses the result as the text feature similarity. The two-character segmentation boundary handling method is as follows: when the subject name character length is less than 2, a unified placeholder # is added at the end to complete the two characters, sliding position by position with a step size of 1, without truncation. The benchmark dictionary mapping method is as follows: standard first-level and second-level subject vocabularies are pre-stored, each word is assigned a unique fixed index, and the vector dimension is equal to the total vocabulary size of the dictionary. Unmatched words are uniformly mapped to the last common noise index. Let the subject A vector be... Subject B vector ,but The calculation result range is [0,1], and negative numbers are uniformly set to 0; the index corresponding to the unmatched character segment is 0, and the word frequency accumulation only counts the number of times the segment appears in the current subject.
[0052] When constructing a multi-level account transaction network structure, heterogeneous accounts are used as nodes. Directed edges are set according to the credit account pointing to the debit account in the same voucher. The absolute value of the transaction amount of the corresponding voucher entries is accumulated as the weight of the directed edge. For each node, the weight of the corresponding directed edge is normalized based on the sum of the weights of all outgoing edges and the sum of the weights of all incoming edges of that node. For any two heterogeneous account nodes, the set of downstream account nodes they point to and the set of upstream account nodes they originate from are counted respectively. Within each set, the smaller value of the normalized edge weights corresponding to the two heterogeneous account nodes is accumulated. And accumulate the larger value in the corresponding normalized edge weights. The results serve as the comparison base, where nodes A and B are two heterogeneous subjects whose topological similarity is to be calculated; node i traverses all downstream subject nodes that A and B both point to; node j traverses all upstream subject nodes that A and B both trace back to; WiA is the normalized edge weight of subject A pointing to downstream node i, and WiB is the normalized edge weight of subject B pointing to downstream node i; WjA is the normalized edge weight of upstream node j pointing to subject A, and WjB is the normalized edge weight of upstream node j pointing to subject B; the overlap of downstream and upstream fund flows is obtained respectively, and the logical operation unit takes the average of the two fund flow overlaps as the topological correlation degree. And this topological correlation is used as the network flow topological similarity;
[0053] The logic unit counts the unique characters contained in the names of all heterogeneous accounts within the current heterogeneous accounting set. It calculates the probability distribution of each character based on its frequency of occurrence in all heterogeneous account names. Then, it sums the product of each character's probability and its base-2 logarithm, taking the negative of the sum. This result is used as the text dispersion. Subsequently, it calculates the initial values of the hierarchical weights for text feature similarity and network flow topology similarity based on the absolute hierarchical value of the currently processed account. , ,in, The fusion weights for text feature similarity. The fusion weights are the network flow topology similarity. The maximum account level value is preset for heterogeneous accounting sets. The automatic identification rule for the maximum level Lmax is to traverse all account level fields in the current accounting set and take the maximum value of the absolute level value of all accounts as Lmax. If there is no level field, the default level is 5 according to accounting standards. In this embodiment, it is 5. L is the absolute level value of the account being processed, and its value is an integer from 1 to 5.
[0054] The logic unit uses the initial hierarchical weights calculated when L is 1 and 5 as mapping endpoints, and performs a linear mapping according to the relative position of the current initial hierarchical weights between the two endpoints, so that... Located in the range of 0.4 to 0.6, making Within the range of 0.6 to 0.4, the current text dispersion is normalized to 0 to 1 according to the low-value boundaries and high-value boundaries of the known mapped samples; and If the calculated value exceeds the 0.4–0.6 range, it is forcibly pruned to the range boundary to ensure that the sum of the two values is always equal to 1. When the text dispersion is below the low-value boundary, the normalization result is set to 0; when it is above the high-value boundary, the normalization result is set to 1. When the text dispersion increases, the normalization result is multiplied by the current value. The difference between 0 and 0.6 is used as the adjustment amount to increase and will Synchronous adjustment to The sum of the two is always equal to 1; the fusion weight adjustment method for text dispersion is as follows: , , ,in This represents the normalized value of the dispersion. This represents the basic topological weight of the hierarchy.
[0055] After weighting adjustments, the overall similarity score satisfies: Where S is the overall similarity score, For text feature similarity, The network flow topology similarity is represented by dimensionless values. For surface-level subjects, text feature similarity maintains a high weight in the overall similarity score. For deep-level detailed subjects or heterogeneous accounts with high text dispersion, the fusion weight of network flow topology similarity is increased accordingly. The selection rule for candidate standard subjects is that if the overall similarity is greater than or equal to the threshold, it is included in the candidate set. The candidate set retains a maximum of Top 3, and standard subjects with the same similarity are matched first with those of the same subject level.
[0056] The logic unit sequentially tests the comprehensive similarity judgment threshold within a closed interval of 0.75 to 0.85 according to a preset step size, and compares the standard subject candidate comparison relationship generated under each threshold with the known mapping sample. When the standard subject candidate comparison relationship obtained from two adjacent tests is consistent and the sum of the number of mismatches and the number of missed matches no longer decreases, the comparison is considered to have reached a convergence state, and the current threshold is used as the comprehensive similarity judgment threshold. Standard subjects with a comprehensive similarity score that reaches the threshold are written into the candidate set of the corresponding heterogeneous subjects and arranged in descending order of comprehensive similarity score to form the standard subject candidate comparison relationship.
[0057] The logic unit parses the detailed auxiliary accounting items in the voucher entry flow data, extracts accounting feature tags based on the original type and name fields of the detailed auxiliary accounting items, and after the standard account candidate comparison relationship is determined, the logic unit reads the subordinate account attributes and accounting specification constraints of the candidate standard account under the standard accounting standards, and uses the standard account candidate comparison relationship as the pre-constraint boundary to adaptively trim the candidate standard type space of the detailed auxiliary accounting items. When the candidate standard account belongs to the receivables and payables category, the candidate standard type space of its subordinate detailed auxiliary accounting items is limited to the preset external unit or individual entity directory. All candidate standard types of the detailed auxiliary accounting item are traversed, and the similarity score of accounting types such as fixed asset category and internal R&D project category that do not belong to the external unit or individual entity directory is set to 0.
[0058] Within each financial accounting cycle, the logic operation unit accumulates the number of occurrences of the same accounting feature tags extracted from the detailed auxiliary accounting items, and uses the accumulated result as the real-time detection frequency of the accounting feature tag. When the real-time detection frequency exceeds a set counting threshold, the frequency exceeding the counting threshold is multiplied by a preset step coefficient, and the result is used as the directed edge normalization weight correction amount corresponding to the accounting feature tag. The topology benchmark parameters used when calculating the network flow topology similarity are dynamically corrected in situ. After the parameter correction is completed, the directed edge weight normalization and topology correlation calculation are re-executed, and the standard subject candidate comparison relationship is updated. Thus, the standard subject candidate comparison relationship limits the space of candidate standard types for detailed auxiliary accounting items. The detection result of the accounting feature tag simultaneously corrects the topology benchmark parameters used for subject comparison. Based on this, the logic operation unit generates an adaptive mapping state result from heterogeneous account subjects to standard subjects and writes it into the relational database storage medium.
[0059] The logic operation unit traverses the voucher entry stream data, matches business keywords in the voucher summary stream, and extracts the pairing structure of debit and credit account categories. The preset transfer account correspondence rules record the corresponding set of business keywords, allowed debit and credit account categories according to the voucher type. Only when the business keywords match the corresponding set and the debit and credit account categories simultaneously satisfy the corresponding pairing constraint rules, is the voucher classified as a period-end transfer voucher, cost transfer voucher, or profit and loss transfer voucher, and the corresponding transaction amount flow of the voucher is removed from the standardized accounting data.
[0060] The data acquisition interface module reads the standard enterprise business registration database from an external data source, extracting detailed accounting entity names and control relationship chain data. The logic operation unit first performs the same character cleansing on the detailed accounting entity names in the standard enterprise business registration database and the detailed accounting entity names under the mapped accounts receivable / payable, just like the discrete account detailed data. Then, it establishes a corresponding relationship based on the identifier field in the standard enterprise business registration database used to uniquely identify the enterprise entity. When a unique enterprise entity can be identified, the original name is replaced with the detailed accounting entity name in the standard enterprise business registration database, completing the standardized renaming. When a unique enterprise entity cannot be identified, the original name is retained, and automatic renaming is not performed. The omission rate of non-standardized enterprise names is 18.7%, which is reduced to 2.1% after adopting the industrial and commercial group topology. The test sample contains 12,000 customer or supplier entity names, covering four major industries: manufacturing, commerce, construction, and services. The original automatic matching omission rate is 18.7%, which is reduced to 2.1% after introducing the group equity topology network.
[0061] When establishing the topology of the same control group relationship, the enterprise entities in the standard enterprise business registration database are used as nodes, and the control direction recorded in the control relationship chain data is used as the directed edges between nodes. The equity control level is recorded by the path depth in the control relationship chain, and the associated attributes are written into the corresponding nodes or directed edges. The independently recorded external enterprise customers and suppliers are thus transformed into structured network nodes containing equity control levels and associated attributes. In the process of cross-account set flow verification and internal reconciliation, the logic operation unit retrieves the enterprise entities in the same control relationship chain along the directed edges, marks the corresponding related party transaction flow, and verifies the account mapping results and transaction amount records of the two parties.
[0062] Within multiple consecutive financial accounting cycles, the logic operation unit records the number of times the adaptive mapping state results change and arranges them according to the financial accounting cycle to form a mapping consistency time series. Then, it calculates the volatility variance value of the mapping consistency time series and uses it as a quantitative indicator of the accounting standardization trend. When the volatility variance value exceeds the set standardization threshold, a high-risk accounting abnormality marking instruction is generated and transmitted to the back-end audit supervision interaction interface. After receiving the instruction, the back-end audit supervision interaction interface switches the financial data of the corresponding heterogeneous accounting set to the look-through verification mode and outputs the original account, the mapped standard account, the mapping status change record, and the related party transaction flow according to the voucher entries for auditors to verify.
[0063] The data cleaning unit continuously captures overlay status change signals from the data source. When it detects that the account set data has been re-imported or the mapping status has changed, it calls the incremental data synchronization algorithm to compare the existing mapping status in the relational database storage medium with the current adaptive mapping status result. Only the changed mapping status is written to the enterprise-level account relational database and a synchronous disk flush is performed.
[0064] Example 2: The current data cleaning unit is configured with an experimental environment in a processing core using a double-precision floating-point arithmetic architecture. The original experimental data is taken from an authorized cross-industry audit working paper entry database, which contains 50,000 rows of discrete accounting voucher entry data with custom naming interference. The logic operation unit sets the sliding time window length for calculating the variance of fluctuation to 12 consecutive accounting cycles to balance the timeliness of data sampling and the storage load of the processing core. When the window length is less than 6 accounting cycles, the sample included in the mapping consistency time series is insufficient, and the time domain statistical results are prone to deviation. When the window length exceeds 24 accounting cycles, the account data of earlier accounting cycles will increase the delay of the statistical results. Therefore, this example uses 12 consecutive accounting cycles as the data span for observing the changing trend of non-standard accounting sets.
[0065] To address the text dispersion of character distribution in the names of discrete account details, the logic operation unit counts the total number of unique characters in all heterogeneous account names within the current heterogeneous accounting set, calculates the probability distribution of each character in all names, multiplies each character's probability by its base-2 logarithm, sums the products, and takes the opposite value. The resulting information entropy value is used as the text dispersion. When the number of character types in heterogeneous account names increases and the character distribution becomes more dispersed, the text dispersion increases accordingly. Based on this, the logic operation unit increases the fusion weight of network flow topology similarity.
[0066] The experimental sequence was set with three naming pollution feature ratios of 10%, 20%, and 30%. Each group first calculated the initial value of the hierarchical weight based on the absolute hierarchical value of the current subject, and then adjusted the weights actually used for multidimensional weighted fusion according to the established weight mapping rules and text dispersion. The 0.83 and 0.17, 0.50 and 0.50 in the following records are the initial values of the hierarchical weight before the interval mapping was performed. In the actual fusion process, the fusion weight of text feature similarity was limited to the range of 0.4 to 0.6, and the fusion weight of network flow topology similarity was adjusted in the opposite direction from 0.6 to 0.4. The sum of the two was kept at 1. The comprehensive similarity judgment threshold was determined according to the comparison convergence state within the closed interval of 0.75 to 0.85.
[0067] When the proportion of naming pollution features is 10%, in sample group one of this invention, the initial values of the hierarchical weights of text feature similarity and network flow topology similarity are 0.83 and 0.17, respectively, and the comprehensive similarity judgment threshold is 0.75. After completing multi-dimensional weighted fusion and standard subject candidate comparison relationship judgment on the input data, the final subject mapping accuracy rate is 98.4%, and the quantitative index of accounting standardization trend is 0.0124. In sample group two of this invention, the initial values of the two hierarchical weights are both 0.50, and the comprehensive similarity judgment threshold is 0.80. The final subject mapping accuracy rate is 96.2%, and the quantitative index of accounting standardization trend is 0.0135.
[0068] After increasing the proportion of named contamination features to 20%, in sample group three of this invention, the absolute level value of the subjects with a value of 1 was processed. The initial values of the two level weights were still 0.83 and 0.17, respectively. The comprehensive similarity judgment threshold was adjusted to 0.78, and the final subject mapping accuracy rate was 97.1%, with a trend quantitative index of accounting standardization of 0.0142. In sample group four of this invention, the absolute level value of the subjects with a value of 5 was processed. The initial values of the two level weights were both 0.50, and the comprehensive similarity judgment threshold was adjusted to 0.82. The final subject mapping accuracy rate was 94.7%, with a trend quantitative index of accounting standardization of 0.0158.
[0069] With a naming pollution feature ratio of 30%, in sample group five of this invention, the initial values of the hierarchical weights for text feature similarity and network flow topology similarity were 0.83 and 0.17, respectively, and the comprehensive similarity judgment threshold was set to 0.85. The final subject mapping accuracy was 95.3%, and the trend quantification index of accounting standardization was 0.0189. In sample group six of this invention, the initial values of the two hierarchical weights were both 0.50, and the comprehensive similarity judgment threshold was also set to 0.85. The final subject mapping accuracy was 93.8%, and the trend quantification index of accounting standardization was 0.0214. As can be seen from the processing results of the three groups of naming pollution feature ratios, as the character interference in heterogeneous subject names increases, the final subject mapping accuracy gradually decreases, and the trend quantification index of accounting standardization increases accordingly. The multi-level subject transaction network structure still provides a reference for deep detailed subjects through the directed fund lending flow.
[0070] In control sample group one, subjects with an absolute level value of 5 were processed under a naming pollution feature ratio of 20%, and network flow topology similarity was isolated. The weights of text feature similarity and network flow topology similarity were set to 1.00 and 0.00, respectively, and the comprehensive similarity judgment threshold was set to 0.80. After determining the candidate comparison relationship of standard subjects based solely on text feature similarity, the final subject mapping accuracy rate was 64.1%, and the quantitative index of accounting standardization trend was 0.0845. Compared with sample group four of this invention, which had the same naming pollution feature ratio and absolute level value, this group lacked the constraint of directed fund lending flow, and the ambiguity interference in heterogeneous subject names directly entered the mapping results.
[0071] Control group two uses the initial values of the two hierarchical weights from control group four of this invention, and increases the comprehensive similarity judgment threshold to 0.92. The final subject mapping accuracy rate is 51.3%, and the quantitative index of the accounting standardization trend is 0.1246. After this threshold exceeds the preset closed interval, some heterogeneous subjects that actually correspond to the standard subjects fail to enter the standard subject candidate comparison relationship, and the number of missed matches increases accordingly. Control group three also uses a 20% naming pollution feature ratio and subjects with an absolute hierarchical value of 5, and sets the weights of text feature similarity and network flow topology similarity to 0.90 and 0.10, respectively. The overall similarity threshold was set to 0.80, and the final subject mapping accuracy was 68.4%, with a trend quantification index of accounting standardization of 0.0762. Although this group retained the network flow topology similarity, its weight was insufficient and still could not offset the interference of custom naming of deep detailed subjects. The control sample group four had the auxiliary accounting reverse correction topology parameter function disabled, 20% naming pollution and level 5, and the recorded mapping accuracy was 73.2%, with a fluctuation variance of 0.0716. Compared with the sample group four of this invention (94.7% and 0.0158), it is proved that bidirectional verification can reduce matching deviation.
[0072] All experimental and control groups used a uniform sample of 50,000 cross-industry voucher entries, the same standard subject dictionary, the same 12-cycle sliding window, and the same counting threshold of 50. Only the corresponding single variable was modified to ensure comparison of a single variable.
[0073] The weight boundaries of the multidimensional weighted fusion are used to ensure that both name text and fund flow information participate in the determination of the candidate comparison relationship of standard subjects. When the fusion weight of text feature similarity is less than 0.4, the basic literal information in heterogeneous subject names is insufficient to contribute to the comprehensive similarity score, and standard subjects with similar literal meanings are prone to splitting. When the fusion weight of network flow topology similarity is less than 0.4, the directed fund lending flow of deep detailed subjects is difficult to offset the interference caused by custom naming. When the comprehensive similarity judgment threshold is less than 0.75, unrelated heterogeneous subjects are prone to enter the candidate set. When the threshold is higher than 0.85, the accounting differences in non-standard accounting sets will cause the actual corresponding standard subjects to be excluded. Based on this, the experimental data keeps the two fusion weights within a mutually compensating adjustment range and limits the comprehensive similarity judgment threshold to an interval that can take into account both mismatches and missed matches.
[0074] After completing the experimental verification of the standard subject candidate comparison relationship, the logic operation unit parses the detailed auxiliary accounting items in the voucher entry flow data and extracts the accounting feature labels. Using the standard subject candidate comparison relationship as a rigid boundary, the space of candidate standard types for detailed auxiliary accounting items is trimmed. The system counts the number of times the same accounting feature label appears in the voucher entry flow data and uses it as the real-time detection frequency. When the real-time detection frequency exceeds the set counting threshold, the topology benchmark parameters used when calculating the network flow topology similarity are dynamically corrected in situ and the standard subject candidate comparison relationship is updated. The data cleaning unit records the iteration frequency data of the adaptive mapping state results in each financial accounting cycle, establishes the mapping consistency time series and calculates the fluctuation variance value. The obtained fluctuation variance value is used as the quantitative indicator of the accounting standardization trend of the accounting set. When this indicator exceeds the set standardization threshold, a high-risk accounting set anomaly marking instruction is generated and transmitted to the background audit supervision interaction interface, and the financial data of the corresponding heterogeneous accounting set is switched to the penetration verification mode. The changed mapping state is then written into the enterprise-level subject relational database and synchronous flushing is performed.
[0075] Example 3: The logic operation unit reads the standard text stream processed by the data cleaning unit from the allocated memory buffer. This standard text stream contains discrete account detail data and voucher entry stream data output by the multi-tenant enterprise resource planning server under financial audit supervision. When importing heterogeneous accounting sets across systems, the auxiliary accounting text of the detail accounts has literal expression variations and semantic ambiguity. The logic operation unit allocates a cache register, extracts the entry summary field according to the preset string delimiter, and converts the summary characters into a data displacement stream; at the same time, it parses the detailed auxiliary accounting items in the voucher entry stream data and extracts accounting feature tags based on their original type field and name field.
[0076] The logic unit loads the accounting feature tags into the cache register and writes them into the corresponding memory storage array address space according to the recognition order. The status detection loop in the processor dynamically polls the number of times the same accounting feature tag appears in the voucher entry stream data within a sliding time window of 12 accounting cycles. The cumulative number is used as the real-time detection frequency. When the real-time detection frequency exceeds the set counting threshold of 50 times, the status detection loop triggers an in-situ reverse correction interrupt. In this process, the real-time detection frequency of the accounting feature tag within the current sliding time window is 58.
[0077] After the in-situ reverse correction interrupt is triggered, the logic unit calls the initially set topology reference parameters stored in the non-volatile storage medium, and dynamically adjusts the topology reference parameters used when calculating the network flow topology similarity according to the following formula: ,in, These are the corrected topology baseline parameters. These are the initial topology baseline parameters. This is the preset step size, with a value of 0.05. To calculate the real-time detection frequency of feature labels within the current sliding time window; a corrected method is used. The network edge weights are renormalized, the topological similarity is recalculated, and the candidate comparison relationships are updated. Each calculation cycle can iterate a maximum of 3 times. If the candidate comparison relationships of standard subjects generated in two consecutive iterations are completely consistent, the iteration stops, and the current topological baseline parameters are locked. ; If the calculated value is less than 0.45, then take 0.45; if it is greater than 0.55, then take 0.55.
[0078] During parameter adjustment, the logic unit applies the correction increment formed by the real-time detection frequency and the step coefficient to the directed edge transfer probability of the network corresponding to the same accounting feature label, and renormalizes the flow distribution matrix between network nodes. The real-time detection frequency is determined by the actual number of times the detailed auxiliary accounting item appears in the voucher entry flow data, which is used to reflect the frequency of the corresponding accounting entity in the account flow. After normalization, the correction increment is written into the topology benchmark parameter, so that the frequency change of the local accounting feature label affects the directed edge weight in the multi-level account transaction network structure. Based on this, the logic unit recalculates the network flow topology similarity and updates the standard account candidate comparison relationship.
[0079] After the topology baseline parameters are adjusted, the logic operation unit uses the candidate comparison relationship of standard subjects as a rigid boundary. Through the similarity tensor matrix, it restricts the space of candidate standard types for detailed auxiliary accounting items under the last-level detailed subject with an absolute level value of L to the range of the trading units, excluding accounting types that do not belong to this range. The candidate comparison relationship of standard subjects thus limits the space of candidate standard types for detailed auxiliary accounting items. The real-time detection frequency of accounting feature labels simultaneously corrects the topology baseline parameters used for network flow topology similarity.
[0080] The logic operation unit continuously performs topology baseline parameter correction, standard account candidate comparison relationship update, and alternative standard type space pruning to generate adaptive mapping status results from heterogeneous account accounts to standard accounts. The database control port writes the mapped relational data into the relational database storage medium, which adopts the relational key-value table in the enterprise-level account relational database. The background audit supervision interaction interface reads the relational key-value table during the flow check to obtain a consistent structured data carrying graph. At this time, the data mapping cleaning stability of the monitoring indicator output is maintained at 95.6%.
[0081] A first-in-first-out (FIFO) rolling window can be used. For each new accounting period, the earliest frequency data in the window is removed. The number of occurrences of the same type of accounting label is accumulated in real time within the window. The counting threshold can be customized from 40 to 80, with a default of 50. The frequency of five types of accounting labels—customer, supplier, project, department, and fixed asset—is counted independently. Only the topological baseline parameters of the directed edges associated with the corresponding subject nodes are corrected, without interference between them. The topological parameters are recalculated a maximum of 3 times within a single accounting period. If there is no change in the candidate comparison relationship for two consecutive times, the iteration stops to avoid cyclic calculation. The updated topological baseline parameters are written to non-volatile storage and read during the initialization of the next accounting period.
[0082] Example 4: This example combines Figures 1 to 2 This section explains the adaptive data mapping method for heterogeneous account items in audit supervision, such as... Figure 1As shown, the system receives discrete account detail data from multiple heterogeneous accounting sets through the data acquisition interface module and extracts the corresponding voucher entry flow data. It extracts the text features of heterogeneous account names from the discrete account detail data and calculates text feature similarity. Simultaneously, based on the directed fund lending and borrowing flows between heterogeneous accounts in the voucher entry flow data, it constructs a multi-level account transaction network structure, calculates the network flow topology similarity between different heterogeneous account nodes, and performs multi-dimensional weighted fusion of text feature similarity and network flow topology similarity to determine the standard account candidate comparison relationship. It parses the detailed auxiliary accounting items in the voucher entry flow data, extracts accounting feature labels, and uses the standard account candidate comparison relationship as a pre-defined rigid boundary to adaptively trim the candidate standard type space of the detailed auxiliary accounting items, generating an adaptive mapping state result from heterogeneous account accounts to standard accounts and writing it into the relational database storage medium. In response to the real-time detection frequency of accounting feature labels in the voucher entry flow data exceeding a set counting threshold, it dynamically corrects the topology benchmark parameters when calculating the network flow topology similarity in situ and updates the standard account candidate comparison relationship accordingly.
[0083] like Figure 2 As shown, the step coefficient decay curve reflects the numerical evolution of the step coefficient over time. The horizontal axis represents the ordinal period of the accounting cycle, and the vertical axis represents the step coefficient. The coordinate scale on the horizontal axis is marked from left to right as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, and the coordinate scale on the vertical axis is marked from bottom to top as 0.00, 0.01, 0.02, 0.03, 0.04, 0.05, 0.06. The initial step coefficient scale corresponding to the accounting cycle ordinal number is 0.05, and as the accounting cycle ordinal number increases, the corresponding step coefficient value gradually decreases along the curve with a negative exponential trend.
[0084] Example 5: Before receiving discrete subject detail data, the current data acquisition interface module performs pre-calibration to determine the initial benchmark point of the topology benchmark parameters. The logic operation unit loads the benchmark entry stream with known mapping relationships into the hardware temporary storage, extracts the text features of heterogeneous subject names from it, and solves the initially set topology benchmark parameters in reverse through the multi-level subject transaction network structure. The obtained static value serves as the zero-point offset for the heterogeneity differences between the underlying database architectures of different financial software. The status detection loop reads the test bit in the address space of the memory storage array. After confirming that the static value is continuously within the range of 0.45 to 0.55, the logic operation unit opens the data transmission channel for the heterogeneous account set adaptive mapping.
[0085] The benchmark journal entry stream uses a standard historical financial voucher dataset that has been manually verified and whose account mapping relationships have been determined, as a benchmark reference for pre-calibration. During reverse engineering, the logic unit inputs the benchmark journal entry stream into the already constructed multi-level account transaction network structure, keeping the directed edge weights and known standard account candidate correspondences unchanged. The initially set topology benchmark parameters are treated as unknown variables, and a gradient descent iterative algorithm is used for reverse derivation to ensure that the network output network flow topology similarity is consistent with the known mapping relationship corresponding to the benchmark journal entry stream. When the iteration reaches convergence, the obtained parameter algebraic solution is used as the initially set topology benchmark parameters. The static values are: loss function = squared difference between predicted topological similarity and benchmark sample similarity; iteration convergence threshold = difference less than 0.001; fixed learning rate is 0.02; maximum number of iterations is 1000.
[0086] When the logic unit processes the voucher entry stream data, it continuously reads the variance value of the mapping consistency time series. When the variance value monotonically increases over three consecutive sampling points and reaches a set normalization threshold, the state detection loop corrects the step coefficient according to the time decay factor within the sliding time window. This causes the step coefficient to decrease exponentially with the calculation cycle, and the logic unit calculates the topology reference parameters based on the corrected step coefficient. The current control value is then overwritten in situ to the non-volatile storage medium. Subsequently, the network flow topology similarity is recalculated and the candidate comparison relationships for standard subjects are updated, causing the fluctuation variance value to converge towards the preset safe interval. The database control port writes the updated adaptive mapping state result to the relational database storage medium. After correction... The constraint is set within the static reference range of 0.45 to 0.55; if it exceeds this range, the boundary value of the range is used.
[0087] The aging attenuation factor is determined based on the time span between the current accounting cycle and the initial calibration cycle. It is used to reduce the impact of data from earlier accounting cycles on the current correction magnitude. The state detection loop first obtains the ordinal number of the currently running accounting cycle, calculates the difference between this ordinal number and the constant 1, and then multiplies the difference by a preset attenuation coefficient of 0.1 to obtain the attenuation exponent. Subsequently, it calculates the base of the natural logarithm to the power of the constant when the attenuation exponent is negative, and uses the resulting value as the aging attenuation factor. Finally, it multiplies the preset initial step coefficient by the aging attenuation factor to obtain the step coefficient used in the current accounting cycle, which decreases exponentially over time. Let t be the ordinal number of the current accounting cycle, attenuation coefficient k = 0.1, attenuation exponent I = (t-1) × k, and aging attenuation factor γ = e -I The current step coefficient η t =η0×γ, where η0 is the initial step coefficient, which defaults to 0.05, η tThe calculation result has a lower limit of 0 and negative numbers are not allowed.
[0088] Example 6: The data acquisition interface module receives historical voucher entry stream data from non-standard accounting sets in the test environment. The logic operation unit performs online self-healing adaptive compensation for sudden changes in the variance value of the mapped consistency time series within the continuous financial accounting period to verify the adaptability of data mapping in a noisy environment. To determine the switching sensitivity of the penetration verification mode after the high-risk accounting set anomaly marking instruction is triggered, the state detection loop establishes multiple sets of anomaly simulation sequences based on the cleaned standard text stream. The normalization threshold of the variance value is set to 0.035, and non-real business data entry noise is injected to raise the variance value from the normal median of 0.015 to the extreme boundary upper limit of 0.048. When the state detection loop detects that the variance value of three consecutive sampling points is greater than the normalization threshold, it triggers the in-situ reverse dynamic correction interruption. If the Var calculated by three consecutive adjacent sliding windows is greater than 0.035, a high-risk marking is triggered. Only when a single window exceeds the threshold, it is not triggered.
[0089] The normalization threshold of 0.035 is used to distinguish between routine accounting adjustments and frequent, large-scale changes in account mapping. Within a normal financial accounting cycle, cross-period carry-over or occasional error corrections cause slight fluctuations in the mapping consistency time series, with a normal median of 0.015 for the variance. When the variance is below 0.035, the adaptive mapping state remains relatively stable across financial accounting cycles. When the variance exceeds 0.035, it indicates frequent, large-scale account mapping tampering or account restructuring in the corresponding heterogeneous accounting set. Based on this, the state detection loop generates a high-risk accounting set anomaly flag instruction and switches the financial data of the corresponding accounting set to the penetration verification mode. The output fields for each voucher include the original heterogeneous account code, the original account name, the mapped standard account, the mapping change record, the auxiliary accounting entity, the related party flag, the debit and credit amount, and the internal offset amount.
[0090] After the in-situ reverse dynamic correction interruption is triggered, the logic operation unit recalculates the topology baseline parameters in real time according to the predetermined update relationship. The step coefficient η performs accelerated step correction under the current out-of-tolerance operating condition, and the real-time detection frequency f of the calculated feature label is in a high-speed accumulation state, which makes the fusion weight of network flow topology similarity more effective. As the boundary value approaches the upper limit of 0.60, the weights of the directed edges in the multi-level subject transaction network structure are adjusted accordingly, the fusion weights of text feature similarity are reduced accordingly, and the similarity tensor matrix is reconstructed accordingly. After the correction is completed, the final subject mapping accuracy rate recovers from the low value of 71.3% and stabilizes at 94.2%.
[0091] During the dynamic adjustment process, the absolute level value is still used to determine the basic fusion weight of the network flow topology similarity. When the real-time detection frequency of the accounting feature label increases rapidly, the logic operation unit adjusts the scalar amplitude of the network flow topology similarity matrix. This amplitude forms a non-linear gain in the matrix reconstruction and normalization process, so that the actual contribution of the network flow topology similarity in the comprehensive similarity score is close to 0.60, without changing the basic weight determined by the absolute level value. Thus, the system reduces the impact of text feature similarity on the candidate comparison relationship of standard subjects by adjusting the internal parameters of the matrix.
[0092] The high-risk account set anomaly marking instruction synchronously locks the corresponding account set's financial data in the penetration verification mode. The database control port calls the incremental data synchronization algorithm to write the changed adaptive mapping status results into the enterprise-level account relational database and execute synchronous disk flushing, so that the constraint transmission between the standard account candidate comparison relationship, detailed auxiliary accounting item identification and mapping status storage reaches a convergence state.
[0093] The embodiments of this application have been described above with reference to the accompanying drawings. Unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other. This application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit of this application and the scope of protection of this invention, and all of these forms are within the protection scope of this application.
Claims
1. An adaptive data mapping method for heterogeneous account items in audit supervision, characterized by: Includes the following steps: Step S101: Receive discrete account detail data from multiple heterogeneous accounting sets through the data acquisition interface module, and extract the corresponding voucher entry stream data. Step S102: Extract the text features of heterogeneous account names from the discrete account details data and calculate the text feature similarity. At the same time, construct a multi-level account transaction network structure based on the directed fund lending flow between heterogeneous accounts in the voucher entry flow data, calculate the network flow topology similarity between different heterogeneous account nodes, and perform multi-dimensional weighted fusion of text feature similarity and network flow topology similarity to determine the standard account candidate comparison relationship. Step S103: Analyze the detailed auxiliary accounting items in the voucher entry flow data, extract accounting feature labels, use the standard account candidate comparison relationship as a pre-fixed rigid boundary, and adaptively trim the candidate standard type space of the detailed auxiliary accounting items; at the same time, in response to the real-time detection frequency of accounting feature labels in the voucher entry flow data exceeding the set counting threshold, dynamically correct the topology benchmark parameters when calculating the network flow topology similarity in situ, and update the standard account candidate comparison relationship, generate the adaptive mapping status result from heterogeneous account accounts to standard accounts, and write it into the relational database storage medium.
2. The adaptive data mapping method for heterogeneous account items in audit supervision according to claim 1, characterized in that: Step S102 includes the following sub-steps: Step S1021, calculate the cosine distance between the text feature vectors of heterogeneous subject names in the discrete subject detail data as the text feature similarity; Step S1022, construct a multi-level subject transaction network structure with heterogeneous subjects as nodes and voucher lending relationships as directed edges, and calculate the topological correlation degree based on the flow distribution between nodes as the network flow topological similarity; Step S1023, weightedly fuse the text feature similarity and the network flow topological similarity to determine the standard subject candidate comparison relationship.
3. The adaptive data mapping method for heterogeneous account items in audit supervision according to claim 1, characterized in that: in After the adaptive mapping status result is written to the relational database storage medium, the following steps are also included: Step S301, read the standard enterprise business registration database from the external data source and extract the detailed accounting entity name and control relationship chain data from it; Step S302, compare the detailed accounting entity name with the detailed accounting entity name under the account that has been mapped in the adaptive mapping status result, and complete the normalized renaming; Step S303, establish a relationship topology structure of the same control group based on the control relationship chain data, convert the independent external enterprise customer and supplier names into structured network nodes containing equity control levels and related attributes, and use the relationship topology structure of the same control group to mark and verify the transaction flow of related parties during cross-account set flow verification and internal reconciliation.
4. The adaptive data mapping method for heterogeneous account items in audit supervision according to claim 1, characterized in that: After the adaptive mapping state results are written to the relational database storage medium, the following steps are also included: Step S401, continuously record and store the iteration frequency data of the adaptive mapping state results in multiple consecutive financial accounting cycles to establish a mapping consistency time series; Step S402, analyze the mapping consistency time series and calculate the volatility variance value of the mapping consistency time series as a quantitative indicator of the trend of accounting standardization; Step S403, when the volatility variance value exceeds the set standardization threshold, generate a high-risk accounting set anomaly marking instruction and transmit the high-risk accounting set anomaly marking instruction to the background audit supervision interaction interface to switch the financial data of the corresponding accounting set to the penetration verification mode.
5. The adaptive data mapping method for heterogeneous account items in audit supervision according to claim 1, characterized in that: In step S102, before weighted fusion of text feature similarity and network flow topology similarity, the text dispersion of the name character distribution in the discrete subject detail data is calculated first, and the fusion weight of network flow topology similarity is increased in response to the increase of text dispersion; wherein, the multidimensional weighted fusion weights of text feature similarity and network flow topology similarity are dynamically adjusted to the numerical ranges of 0.4 to 0.6 and 0.6 to 0.4 respectively, and the sum of the two weights is always equal to 1; and by comparing the convergence state, the comprehensive similarity judgment threshold of the determined standard subject candidate comparison relationship is set in the closed interval of 0.75 to 0.
85.
6. The adaptive data mapping method for heterogeneous account items in audit supervision according to claim 1, characterized in that: It also includes the following steps: Step S601: Traverse the voucher entry stream data, match business keywords in the voucher summary stream, and extract the pairing structure of debit account categories and credit account categories; Step S602: Based on the preset transfer account correspondence rules as the pairing constraint rules for the pairing structure, classify and identify period-end transfer vouchers, cost transfer vouchers, and profit and loss transfer vouchers, and remove the corresponding voucher transaction amounts from the standardized accounting data; Step S603: Capture the overlay status change signal from the data source, call the incremental data synchronization algorithm, and write the changed mapping status into the enterprise-level account relational database for synchronous flushing.
7. The adaptive data mapping method for heterogeneous account items in audit supervision according to claim 1, characterized in that: Step S101, obtaining discrete account detail data from multiple heterogeneous external accounting sets, includes the following sub-steps: Step S1011, removing leading and trailing spaces, tabs, and non-printable implicit control characters from the original account data of multiple heterogeneous external accounting sets; Step S1012, identifying English letters embedded in the original account data, uniformly converting lowercase English letters to uppercase English letters, converting full-width characters to half-width characters, and outputting the cleaned standard text stream as discrete account detail data.
8. The adaptive data mapping method for heterogeneous account items in audit supervision according to claim 2, characterized in that: Step S1022, which calculates the topological correlation degree based on the flow distribution between nodes, includes the following sub-steps: Step S10221, normalize the weights of the directed edges in the constructed multi-level subject transaction network structure; Step S10222, for any two heterogeneous subject nodes, respectively count the set of downstream account nodes they both point to and the set of upstream account nodes they both originate from; Step S10223, calculate the topological correlation degree based on the overlap of fund flows between the downstream account node set and the upstream account node set.
9. The adaptive data mapping method for heterogeneous account items in audit supervision according to claim 1, characterized in that: in After the adaptive mapping status result is written to the relational database storage medium, the following steps are also included: Step S901, output the generated adaptive mapping status result and the corresponding discrete candidate standard subject comparison list to the interactive interface; Step S902, capture the manual confirmation instruction or the batch overwrite instruction of the large batch comparison electronic form from the input port, and update and solidify the adaptive mapping status result in the relational database storage medium according to the instruction.
10. The adaptive data mapping method for heterogeneous account items in audit supervision according to claim 1, characterized in that: in In the process of in-situ reverse dynamic correction of the topological reference parameters when calculating the network flow topology similarity, a step coefficient is used for correction, and the step coefficient gradually decreases with a negative exponential law as the calculation cycle progresses. The step coefficient used in the current calculation cycle is calculated in the following way: obtain the ordinal number of the current calculation cycle, calculate the difference between the ordinal number and the constant 1, multiply the obtained difference by the preset attenuation coefficient to obtain the attenuation exponent; calculate the constant power of the base of the natural logarithm when the attenuation exponent is negative, and use the obtained value as the time-effect attenuation factor; multiply the preset initial step coefficient by the time-effect attenuation factor to obtain the step coefficient of the current calculation cycle.
Citation Information
Patent Citations
Multilink financial subject mapping method and system fusing semantic model and rule matching
CN121579453A