Financial data processing method and device, electronic equipment and storage medium

CN122616488APending Publication Date: 2026-08-21PING AN INT FINANCIAL LEASING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610748180.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-27
Publication Date
2026-08-21

AI Technical Summary

Technical Problem

[0004]本发明提供一种财务数据处理方法、装置、电子设备及存储介质,以解决现有的财务数据处理方案因缺乏语义理解和容错机制,导致异构数据匹配准确率低且人工干预成本高昂的技术问题

Benefits of technology

[0009]上述财务数据处理方法、装置、电子设备及存储介质所实现的方案中,可以获取待处理的原始财务数据,并提取原始财务数据中的待识别字段;对待识别字段与多维财务术语知识库中的标准字段进行初步匹配分析,得到候选标准字段及对应的初步置信度;在初步置信度未满足预设的高置信度条件时,提取待识别字段的语义嵌入向量,并结合上下文句法角色计算语义嵌入向量与标准字段间的语义相似度,确定最终匹配置信度;根据最终匹配置信度以及预设的科目风控规则,触发对应的分级响应策略将待识别字段映射至对应的目标标准字段,得到目标财务填报数据。在本申请中,针对多源异构财务数据合并与智能填报等复杂业务场景,先是通过初步模糊粗筛与深层语义特征分析的联动机制,精准解析同义词、缩写及行业用语等复杂表述的深层含义与句法特征以输出高精度的匹配置信度,再将该置信度结合核心业务属性等风控规则实施包含自动填充、备选推荐、专家复核及异常阻断的动态干预机制,能有效地提升非标准陌生字段的匹配映射精度,并在精准拦截逆向错误或违背财务逻辑异常的同时大幅减少人工全量复核的负担,能够解决现有的财务数据处理方案因缺乏语义理解和容错机制,导致异构数据匹配准确率低且人工干预成本高昂的技术问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122616488A_ABST
    Figure CN122616488A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of data processing and artificial intelligence, and discloses a financial data processing method and device, an electronic device and a storage medium, which comprise the following steps: obtaining original financial data to be processed, and extracting a field to be recognized in the original financial data; performing preliminary matching analysis on the field to be recognized and standard fields in a multi-dimensional financial term knowledge base to obtain candidate standard fields and a preliminary confidence degree; when the preliminary confidence degree does not satisfy a high confidence degree condition, extracting a semantic embedding vector of the field to be recognized, combining a context syntax role to calculate semantic similarity between the semantic embedding vector and the standard fields, and determining a final matching confidence degree; according to the final matching confidence degree and a subject risk control rule, triggering a hierarchical response strategy to map the field to be recognized to a corresponding target standard field, and obtaining target financial reporting data. The application can be applied to business scenarios such as intelligent reporting of enterprise finance, can reduce the cost of manual checking, and can improve the accuracy and security of financial data processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of data processing and artificial intelligence technology, and can be applied to the financial business field, particularly to a financial data processing method, apparatus, electronic device and storage medium. Background Technology

[0002] In financial and accounting operations, such as the merging and reporting of multi-source heterogeneous financial data, the most commonly used techniques for cleaning and mapping financial data and fields are based on fixed rules, preset field mapping tables, or simple keyword matching techniques. This involves manually configuring field correspondences or requiring users to provide highly structured input data for import and parsing using templates in financial software.

[0003] The inventors realized that existing financial data processing solutions rely heavily on static rule-based matching systems. When faced with heterogeneous financial data with diverse expressions in actual business, the semantic understanding is weak, and data matching is prone to failure due to differences in terminology or implicit formula-derived logic. At the same time, existing systems lack flexible fault tolerance and error correction mechanisms when faced with matching anomalies. When errors occur, it is often necessary to manually rerun the entire process, which in turn leads to low data matching accuracy and high manual intervention costs in existing financial data processing solutions. Summary of the Invention

[0004] This invention provides a financial data processing method, apparatus, electronic device, and storage medium to solve the technical problems of low accuracy in heterogeneous data matching and high cost of manual intervention caused by the lack of semantic understanding and fault tolerance mechanisms in existing financial data processing solutions.

[0005] To achieve the above objectives, a first aspect of this application provides a financial data processing method, the method comprising: Obtain the initial financial data to be processed, and extract the fields to be identified from the initial financial data; A preliminary matching analysis is performed on the field to be identified and the standard fields in the multidimensional financial terminology knowledge base to obtain candidate standard fields and their corresponding preliminary confidence levels. When the initial confidence level does not meet the preset high confidence level condition, the semantic embedding vector of the field to be identified is extracted, and the semantic similarity between the semantic embedding vector and the standard field is calculated in combination with the context syntactic roles to determine the final matching confidence level. Based on the final matching confidence level and the preset subject risk control rules, the corresponding hierarchical response strategy is triggered to map the field to be identified to the corresponding target standard field and generate target financial reporting data; wherein, the hierarchical response strategy includes an automatic filling strategy, an alternative recommendation strategy, an expert review strategy, and an anomaly blocking strategy.

[0006] To achieve the above objectives, a second aspect of this application provides a financial data processing apparatus, the apparatus comprising: The acquisition module is used to acquire the initial financial data to be processed and extract the fields to be identified from the initial financial data; The preliminary matching module is used to perform preliminary matching analysis between the field to be identified and the standard fields in the multidimensional financial terminology knowledge base to obtain candidate standard fields and their corresponding preliminary confidence levels. The determination module is used to extract the semantic embedding vector of the field to be identified when the initial confidence does not meet the preset high confidence condition, and calculate the semantic similarity between the semantic embedding vector and the standard field in combination with the context syntactic roles to determine the final matching confidence. The generation module is used to trigger a corresponding tiered response strategy to map the field to be identified to the corresponding target standard field based on the final matching confidence level and the preset subject risk control rules, thereby generating target financial reporting data; wherein, the tiered response strategy includes an automatic filling strategy, an alternative recommendation strategy, an expert review strategy, and an anomaly blocking strategy.

[0007] To achieve the above objectives, a third aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect.

[0008] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method of the first aspect described above.

[0009] The above-mentioned financial data processing method, device, electronic equipment, and storage medium can acquire the original financial data to be processed and extract the fields to be identified from the original financial data; perform preliminary matching analysis between the fields to be identified and the standard fields in the multidimensional financial terminology knowledge base to obtain candidate standard fields and their corresponding preliminary confidence levels; when the preliminary confidence level does not meet the preset high confidence level condition, extract the semantic embedding vector of the fields to be identified, and calculate the semantic similarity between the semantic embedding vector and the standard fields in combination with the context syntactic roles to determine the final matching confidence level; based on the final matching confidence level and the preset subject risk control rules, trigger the corresponding hierarchical response strategy to map the fields to be identified to the corresponding target standard fields to obtain the target financial reporting data. In this application, for complex business scenarios such as merging multi-source heterogeneous financial data and intelligent data entry, a linkage mechanism of preliminary fuzzy screening and deep semantic feature analysis is first used to accurately analyze the deep meaning and syntactic features of complex expressions such as synonyms, abbreviations and industry terms to output high-precision matching confidence. Then, this confidence is combined with risk control rules such as core business attributes to implement a dynamic intervention mechanism including automatic filling, alternative recommendation, expert review and anomaly blocking. This can effectively improve the matching accuracy of non-standard unfamiliar fields, and while accurately blocking reverse errors or anomalies that violate financial logic, it can significantly reduce the burden of manual full-scale review. This can solve the technical problem that existing financial data processing solutions lack semantic understanding and fault tolerance mechanisms, resulting in low accuracy of heterogeneous data matching and high cost of manual intervention. Attached Figure Description

[0010] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a schematic diagram of an application environment for a financial data processing method according to an embodiment of the present invention; Figure 2 This is a flowchart of the financial data processing method provided in the embodiments of this application; Figure 3 This is another flowchart of the financial data processing method provided in the embodiments of this application; Figure 4 This is another flowchart of the financial data processing method provided in the embodiments of this application; Figure 5 This is another flowchart of the financial data processing method provided in the embodiments of this application; Figure 6 This is another flowchart of the financial data processing method provided in the embodiments of this application; Figure 7 This is another flowchart of the financial data processing method provided in the embodiments of this application; Figure 8 This is another flowchart of the financial data processing method provided in the embodiments of this application; Figure 9 This is a schematic diagram of the structure of the financial data processing device provided in the embodiments of this application; Figure 10 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Figure 11 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0012] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0013] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0014] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0015] In financial and accounting operations, such as the merging and reporting of multi-source heterogeneous financial data, the most commonly used techniques for cleaning and mapping financial data and fields are based on fixed rules, preset field mapping tables, or simple keyword matching techniques. This involves manually configuring field correspondences or requiring users to provide highly structured input data for import and parsing using templates in financial software.

[0016] The inventors realized that existing financial data processing solutions rely heavily on static rule-based matching systems. When faced with heterogeneous financial data with diverse expressions in actual business, the semantic understanding is weak, and data matching is prone to failure due to differences in terminology or implicit formula-derived logic. At the same time, existing systems lack flexible fault tolerance and error correction mechanisms when faced with matching anomalies. When errors occur, it is often necessary to manually rerun the entire process, which in turn leads to low data matching accuracy and high manual intervention costs in existing financial data processing solutions.

[0017] Based on this, the embodiments of this application provide a financial data processing method, apparatus, device and storage medium, which aims to solve the technical problem that existing financial data processing solutions lack semantic understanding and fault tolerance mechanisms, resulting in low accuracy of heterogeneous data matching and high cost of manual intervention.

[0018] The financial data processing method, apparatus, device, and storage medium provided in this application are specifically described through the following embodiments. First, the financial data processing method in this application embodiment is described.

[0019] The financial data processing method provided in the application embodiments can be applied to, for example, Figure 1 In this application environment, the client communicates with the server via a network. The server can obtain the raw financial data to be processed from the client and extract the fields to be identified from the raw financial data; perform preliminary matching analysis on the fields to be identified with standard fields in the multidimensional financial terminology knowledge base to obtain candidate standard fields and their corresponding preliminary confidence levels; when the preliminary confidence level does not meet the preset high confidence level condition, extract the semantic embedding vector of the fields to be identified, and calculate the semantic similarity between the semantic embedding vector and the standard fields in combination with the context syntactic roles to determine the final matching confidence level; based on the final matching confidence level and the preset subject risk control rules, trigger the corresponding hierarchical response strategy to map the fields to be identified to the corresponding target standard fields to obtain the target financial reporting data, and feed back the target financial reporting data or the interactive prompts generated by the hierarchical response (such as alternative recommendations, alarm information, etc.) to the client. This application addresses complex business scenarios such as merging and intelligently filling in multi-source heterogeneous financial data. A multi-stage semantic matching and hierarchical response control scheme is utilized. First, a two-level linkage mechanism of preliminary fuzzy screening and deep semantic feature analysis accurately parses the deep meaning and syntactic features of complex expressions such as synonyms, abbreviations, and industry terms to output high-precision matching confidence. Then, this confidence is combined with risk control rules such as core business attributes to implement a dynamic intervention mechanism including automatic filling, alternative recommendations, expert review, and anomaly blocking. This effectively improves the matching accuracy of non-standard unfamiliar fields and significantly reduces the burden of manual full-scale review while accurately intercepting core business attribute conflicts or data logic inconsistencies. This solves the technical problem of low accuracy in heterogeneous data matching and high manual intervention costs in existing financial data processing solutions due to a lack of semantic understanding and fault tolerance mechanisms. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The following detailed description of specific embodiments further illustrates this application.

[0020] Figure 2 This is an optional flowchart of the financial data processing method provided in the embodiments of this application. Figure 2 The method may include, but is not limited to, steps S201 to S204.

[0021] Step S201: Obtain the initial financial data to be processed and extract the fields to be identified from the initial financial data.

[0022] Step S202: Perform a preliminary matching analysis between the field to be identified and the standard fields in the multidimensional financial terminology knowledge base to obtain candidate standard fields and their corresponding preliminary confidence levels.

[0023] Step S203: When the initial confidence level does not meet the preset high confidence level condition, extract the semantic embedding vector of the field to be identified, and calculate the semantic similarity between the semantic embedding vector and the standard field in combination with the context syntactic roles to determine the final matching confidence level.

[0024] Step S204: Based on the final matching confidence level and the preset subject risk control rules, trigger the corresponding hierarchical response strategy to map the field to be identified to the corresponding target standard field and generate target financial reporting data; wherein, the hierarchical response strategy includes automatic filling strategy, alternative recommendation strategy, expert review strategy and anomaly blocking strategy.

[0025] In some embodiments, to overcome the weak semantic understanding caused by the reliance on static fixed rules and preset mapping tables in traditional financial data entry tools, and to achieve accurate mapping and flexible fault tolerance of multi-source heterogeneous financial data, thereby solving the technical problems of low accuracy in heterogeneous data matching and high cost of manual intervention in complex business scenarios, the solution of this application combines a dynamic and scalable knowledge base with a multi-stage fusion semantic matching engine. This allows for step-by-step parsing of the non-standard fields to be identified, from lightweight to deep semantics, and incorporates a gradient mapping processing strategy based on risk control rules to obtain high-quality target financial reporting data, as described below.

[0026] In step S201 of some embodiments, the initial financial data typically originates from files exported from ERP systems of different companies, cross-departmental business reports, or third-party financial statements. Because these data come from diverse sources, have varied business models, and lack standardized specifications, they often contain numerous synonyms, abbreviations, or industry-specific terms. For example, the standard financial term "operating revenue" is abbreviated to "revenue," or "accounts payable" is non-standardized as "payables." The process of extracting the fields to be identified involves parsing the file structure and header metadata of the aforementioned initial financial data to extract the header names or indicator dimension labels that need to be standardized and mapped. These unstandardized fields to be identified form the underlying input foundation for subsequent multi-stage matching and mapping transformations.

[0027] In step S202 of some embodiments, the "multidimensional financial terminology knowledge base" is a structured data set with rich benchmark terms and mapping relationships. In the preliminary matching analysis stage, the system executes a lightweight, fast filtering mechanism. Specifically, it can utilize lightweight algorithms such as basic string fuzzy matching or text distance metrics to quickly measure the difference and overlap between the field to be identified and each standard field in the knowledge base at the literal character level. This allows for the initial identification of several candidate standard fields with high literal similarity from a massive number of entries, and the calculated matching score is used as the corresponding preliminary confidence level. This stage aims to use low-computational-consumption lightweight algorithms to quickly filter out a large number of completely irrelevant entries, thereby narrowing the search space and significantly improving the overall system's retrieval efficiency.

[0028] Please see Figure 3 In some embodiments, step S202 may include, but is not limited to, steps S301 to S302.

[0029] Step S301: Use edit distance and prefix wildcard algorithms to perform preliminary fuzzy matching and screening between the field to be identified and the standard field, and calculate the candidate standard field and the corresponding preliminary confidence level.

[0030] Step S302: If the initial confidence level meets the preset high confidence level condition, then the candidate standard field is directly used as the target standard field, and the initial confidence level is used as the final matching confidence level.

[0031] In some embodiments, to achieve rapid filtering and accurate matching of basic-level literal differences based on the characteristics of massive non-standard financial feature data, thereby solving the technical problem of excessive computational consumption and low processing efficiency of a single deep model in complex financial mapping scenarios, this application's solution encapsulates lightweight text distance metrics and prefix matching rules into preliminary verification logic. This allows for low-latency filtering and confidence assessment of massive knowledge base entries using surface-level matching algorithms. In step S301 of some embodiments, edit distance refers to the minimum number of single-character editing operations (including insertion, deletion, and replacement) required to transform one string into another; the prefix wildcard algorithm is a string matching rule used to detect whether the field to be identified starts with a specific character sequence. In the specific calculation process, the extracted field to be identified is compared with each standard field in the multidimensional financial terminology knowledge base, and candidate targets are initially located using edit distance and the prefix wildcard algorithm. Based on the comprehensive evaluation of the above algorithms, the overlap of text at the literal character level can be quickly measured and converted into a quantified surface similarity score, which serves as the corresponding preliminary confidence parameter (denoted as ). Finally, the system filters out the initial confidence levels. Knowledge base entries that meet the basic requirements will be output as candidate standard fields.

[0032] In step S302 of some embodiments, a quantified high-confidence determination condition is pre-configured, which is typically expressed as a set confidence threshold (e.g., denoted as ). (The default value can be set to 98%). The system obtains the preliminary confidence level of the output from the preceding steps. and compare it with the threshold Comparison and judgment. If a match is found and the confidence level exceeds the threshold (default 98%), it indicates that the current field to be identified and the corresponding candidate standard field have reached a very high degree of certainty in terms of literal composition. The candidate standard field is then directly established as the target standard field for the mapping output, and the current preliminary confidence level is adjusted accordingly. Assigned and used as the final match confidence variable (denoted as The process is streamlined, thus skipping subsequent complex semantic parsing nodes.

[0033] Through steps S301 to S302 described above, this embodiment combines the speed advantages of edit distance and prefix wildcard algorithms in literal feature comparison. This not only accurately identifies invalid terms with significant differences, greatly reducing the data retrieval space, but also achieves rapid mapping of highly similar standardized fields through a direct-access decision mechanism with a high confidence threshold. This design, while ensuring the accuracy of initial screening, greatly saves the overall system's underlying computing power, reserving sufficient computing resources for subsequent, unavoidable deep semantic analysis, and significantly improving the overall response speed of heterogeneous financial data processing.

[0034] In step S203 of some embodiments, the aforementioned high confidence condition can be digitally represented by setting a specific numerical threshold. When the initial confidence exceeds the set condition, it indicates that the literal matching is accurate enough, and the result can be directly adopted to save computing resources; when the initial confidence does not reach the condition, it indicates that the field to be identified may have complex semantic transitions or deep references, and the deep feature analysis process will be activated. In this process, the discrete text to be identified is converted into a high-dimensional continuous semantic embedding vector to represent its deep abstract semantics. Subsequently, the similarity between the semantic embedding vector and the corresponding vectors of each candidate standard field in the knowledge base is calculated in the vector space to obtain a deep semantic similarity score. In order to further filter out interference items that are semantically similar but have completely different business roles, natural language processing technology is also used to extract the contextual syntactic role of the field to be identified in the original sentence context, such as identifying whether it is a subject, object or numerical modifier in the syntactic dependency tree, and using it as a feature constraint to weight and correct the semantic similarity, thereby determining the most accurate final matching confidence.

[0035] Please see Figure 4 In some embodiments, step S203 may include, but is not limited to, steps S401 to S403.

[0036] Step S401: If the initial confidence level does not meet the preset high confidence level condition, then the semantic embedding vector of the field to be identified is generated based on the pre-trained financial domain-specific language model.

[0037] Step S402: Calculate the semantic similarity score between the semantic embedding vector and the candidate standard field, and extract the contextual syntactic role of the field to be identified in the original sentence.

[0038] Step S403: Filter out matching interference items based on contextual syntactic roles to determine the final matching confidence.

[0039] In some embodiments, to accurately identify complex expressions such as synonyms, abbreviations, and industry terms in financial data, and to achieve high-precision mapping of non-standard fields with significant literal differences but consistent deep semantics, thereby solving the technical problem of high matching failure rates caused by weak semantic understanding in traditional pure text matching schemes, the solution of this application introduces a pre-trained financial domain-specific language model and syntactic analysis technology. This model can transform discrete text into continuous high-dimensional feature space vectors and perform cross-validation based on contextual roles to output accurate and reliable final matching confidence, as described below.

[0040] In step S401 of some embodiments, the preliminary confidence parameters output by the preceding preliminary matching stage are obtained. And compare it with a pre-set high confidence threshold. Perform numerical comparison. When a determination is made... This indicates a significant difference in the surface string structure between the current field to be identified and the candidate standard fields, making it impossible to directly ascertain their mapping relationship based on shallow features. To delve deeper into the business meaning behind this field, a pre-trained financial domain-specific language model (e.g., a deep neural network model fine-tuned with massive amounts of professional financial text corpus) is invoked to process the field to be identified. This pre-trained financial domain-specific language model possesses feature extraction and attention mechanisms, enabling it to capture the industry-specific logic behind the text, mapping and converting the originally discrete text data of the field to be identified into a high-dimensional continuous vector representation, namely the semantic embedding vector. This semantic embedding vector not only preserves the literal information but also deeply integrates the abstract semantic representation of the word in the financial professional context, thus providing a high-dimensional feature data foundation for subsequent calculations.

[0041] In step S402 of some embodiments, after obtaining the semantic embedding vector of the field to be identified... Then, the high-dimensional standard vectors corresponding to the candidate standard fields in the knowledge base are retrieved synchronously (denoted as...). Subsequently, the system performs spatial measurements in the vector space and calculates... and The distance between them is used to obtain a quantified semantic similarity score. This score reflects the similarity of meaning between the two fields in a financial professional context. However, financial data often carries contextual dependence. To prevent semantic ambiguity caused by the polysemy of isolated words, natural language processing techniques are further used to perform syntactic dependency analysis on the original input data or original sentences containing the field to be identified. By analyzing the structural dependency relationships between words in the sentence, the contextual syntactic role parameter of the field to be identified in the original data context is extracted. The contextual syntactic role typically reflects the syntactic function of a field in specific business logic, such as identifying whether it functions as a subject, object, or numerical modifier—key structural features.

[0042] In step S403 of some embodiments, in actual financial data flow, it is common to encounter situations where individual fields, although having high individual semantic similarity scores, conflict with their corresponding candidate standard fields in actual business reference due to differences in their position and modification relationship in the original statement (e.g., structural differences between modifying the main business and modifying other businesses). These evaluation objects with high similarity scores but misaligned syntactic roles are considered matching interference items. In this execution step, the previously extracted contextual syntactic role parameters are... The evaluation incorporates key structural constraints. Based on a pre-defined syntactic mapping logic, the baseline syntactic attributes of candidate standard fields are examined in relation to the extracted contextual syntactic roles. Compatibility and matching. Utilize semantic similarity score Weighted corrections are applied to effectively eliminate matching interference items that, although geographically close in the vector space, contradict the actual syntactic relations and modified objects. After this rigorous filtering and verification based on syntactic features, the final matching confidence score is output, which is accurate and strongly supported by business logic. This confidence level data will be used as the core basis for determining the triggering of various graded fault-tolerant response strategies.

[0043] Through the above steps S401 to S403, the embodiments of this application break through the limitations of traditional literal character matching, upgrade the matching dimension to a higher-order semantic vector space, and introduce the original syntactic dependency relation as a cross-validation and interference filtering method. This effectively eliminates the interference caused by the polysemy of single words and the complexity of financial term combinations, greatly reduces the probability of mismatch caused by industry jargon or long-tail non-standard expressions, and significantly improves the accuracy of heterogeneous financial data mapping and the robustness of the overall system.

[0044] In step S204 of some embodiments, facing the inherent uncertainty in financial matching, this step implements four levels of dynamic feedback and risk control intervention. Specifically, the first level is an automatic filling strategy, that is, when the final matching confidence is extremely high, the mapping to the target standard field can be completed automatically without manual intervention; the second level is an alternative recommendation strategy, which generates alternative solutions on the interactive interface for users to confirm for matching results with medium confidence; the third level is an expert review strategy, which is strongly correlated with the preset account risk control rules. When the system detects that the current mapping operation involves a conflict of core business attributes (such as potential risks such as the reversal of asset and liability directions), even if the confidence level meets the standard, it will forcibly intercept the automatic mapping flow and transfer the corresponding data to the expert review queue to prevent systemic risks; the fourth level is an anomaly blocking strategy, which will forcibly block the current processing flow and output alarm information to the terminal when it is found that the data characteristics of the field to be identified have not passed the data logic consistency check (such as serious logical contradictions such as filling negative numbers into the accumulated depreciation column). Ultimately, after the screening and intervention of the above-mentioned multi-level safety valves, the fields to be identified are accurately and safely mapped to the target standard fields, and then the target financial reporting data that meets the compliance audit specifications is summarized and output. Please see Figure 5 In some embodiments, step S204 may include, but is not limited to, steps S501 to S504.

[0045] Step S501: If the final matching confidence is greater than or equal to the first confidence threshold, the automatic filling strategy is triggered, and the candidate standard field corresponding to the final matching confidence is directly used as the target standard field for filling.

[0046] In step S502, if the final matching confidence is less than the first confidence threshold but greater than or equal to the second confidence threshold, the alternative recommendation strategy is triggered, a recommendation prompt containing at least one alternative candidate standard field is generated and displayed on the interactive interface, and the target standard field is determined based on the received confirmation instruction.

[0047] Step S503: If a core business attribute conflict is detected between the field to be identified and the corresponding candidate standard field according to the subject risk control rules, the expert review strategy is triggered to intercept the current automatic mapping process and transfer the corresponding data to the expert review queue for manual determination of the target standard field.

[0048] Step S504: If the data features of the field to be identified fail the data logic consistency check associated with the candidate standard field, the abnormal blocking strategy is triggered according to the subject risk control rules, the current processing flow is forcibly blocked and an alarm message is output.

[0049] In some embodiments, to address the uncertainties and potential business risks inherent in financial data mapping and achieve adaptive fault tolerance and precise intervention in the data flow process, thereby solving the technical challenges of high costs and compliance risks associated with traditional manual verification, the solution in this application deeply integrates the confidence assessment of matching results with underlying financial risk control rules. The system can execute multi-level gradient intervention mechanisms, including automatic flow, interactive recommendation, and expert review, to output highly accurate and compliant target financial reporting data, as described below.

[0050] In step S501 of some embodiments, an internal quantitative evaluation system is pre-configured with a judgment parameter for characterizing extremely high matching confidence, namely the first confidence threshold. The final matching confidence score output by the preceding deep semantic analysis. With the first confidence threshold Perform numerical comparison. When a determination is made... When the match confidence score is high, it indicates that the current field to be identified and the candidate standard field are highly consistent in semantic features, and the probability of mapping error is extremely low. Based on this judgment result, an automatic fill strategy is triggered. That is, without any manual intervention or confirmation interaction, the candidate standard field pointed to by the final matching confidence score is directly established as the target standard field for output, and the automatic assignment and fill of the underlying data link is completed. This step is mainly for regular, high-frequency and standardized financial data. Through fully automated machine decision-making, it frees up human resources for basic data entry and improves the overall throughput of financial data processing.

[0051] In step S502 of some embodiments, in order to construct a fault-tolerant range with a gradient, an internal range slightly lower than [previous value] is also set. Second confidence threshold When determining the final match confidence level Falling in the range If the current matching result has a high semantic relevance, but has not yet reached the absolute safety level where the machine can make an autonomous decision, then there is a possibility of confusion due to multiple meanings of a single word or similar business concepts. At this point, an alternative recommendation strategy is triggered. Forced automatic mapping is no longer executed; instead, the field to be identified and its corresponding matching results containing at least one high-scoring candidate standard field are encapsulated, generating a visual recommendation prompt, which is then pushed to the interactive interface of an external terminal for display. Subsequently, the system enters a listening suspension state, awaiting manual review by business personnel. Once a confirmation instruction (such as an external click or selection) is received from the interactive interface, a unique target standard field is established based on the option indicated by the confirmation instruction. This avoids the risk of machine misjudgment and eliminates the tedious process of manually searching the standard library from scratch.

[0052] In step S503 of some embodiments, the subject risk control rules are a set of business logic security rules independent of the confidence calculation system. These rules are used to verify the compatibility of data in terms of their essential financial characteristics. In the matching process, not only is literal and semantic similarity considered, but the subject risk control rules are also continuously invoked for parallel verification. Specifically, the inherent attributes of the field to be identified are extracted and compared with the baseline attributes of the mapping target. Once a conflict of core business attributes, such as opposite fund flows or incorrect account levels, is detected, even if the final matching confidence score calculated above meets a high condition, an expert review strategy will be forcibly triggered. Under this strategy, the current automatic mapping or recommendation process is immediately intercepted and blocked. The data with suspected attribute conflicts, along with its contextual information, is packaged and transferred to a higher-authority expert review queue. This forces a human node with advanced financial qualifications to intervene and review the data, ultimately determining the target standard field through human decision-making. This step effectively prevents systemic financial direction errors caused by the semantic limitations of the algorithm.

[0053] In step S504 of some embodiments, in addition to the aforementioned attribute conflict, hard constraint tests are performed on the compliance of the underlying data using account risk control rules. Specifically, the underlying data characteristics (such as the positive or negative polarity of the value, data type, and amount magnitude) carried by the field to be identified are extracted, and the data logic consistency is verified using the baseline business logic preset by the candidate standard field. If it is found that there is an irreconcilable serious violation between the data characteristics of the field to be identified and the physical restrictions or financial common sense of the target field (such as mismatch between numerical features and text features, or abnormal filling of mutually exclusive accounts), it is determined that the mapping request has a logical vulnerability. At this time, the highest response level of the abnormal blocking policy is immediately triggered, the mapping request is discarded and the associated processing flow of the current batch is forcibly blocked; at the same time, an alarm message containing the error traceability source code and details of conflict variables is generated and output to the system console or the terminal of the operation and maintenance personnel.

[0054] In summary, by performing automatic filling for high-confidence data, introducing alternative recommendation interactions for fuzzy data, and applying expert interception and forced blocking for attribute conflict and logically abnormal data, this solution ensures high-concurrency, manual flow of massive amounts of standardized financial data while establishing a rigorous and controllable human-machine collaboration and abnormal circuit breaker mechanism to address potential business logic risks.

[0055] Please see Figure 6 In some embodiments, after step S204, steps S601 to S604 may be included, but are not limited to.

[0056] Step S601: If the field to be identified does not match a target standard field that meets the conditions in the multidimensional financial terminology knowledge base, then the field to be identified is marked as a term to be confirmed.

[0057] Step S602: Record the manual correction methods for the terms to be confirmed in the subsequent interactive processing flow.

[0058] Step S603: In response to receiving an active correction instruction for the term to be confirmed in the interactive interface, or detecting that the same term to be confirmed has been manually corrected in the same way in multiple consecutive independent processing tasks and the number of corrections has reached a second preset number threshold, a correction mapping relationship between the term to be confirmed and the manual correction method is generated.

[0059] Step S604: Dynamically write and update the multidimensional financial terminology knowledge base with the terms to be confirmed and their corresponding correction mapping relationships as official terms.

[0060] In some embodiments, to cope with the continuous evolution of financial business scenarios and the emergence of new and unknown terms, and to achieve adaptive growth of the knowledge base and dynamic expansion of mapping rules, thereby solving the technical problems of poor system scalability and high maintenance costs caused by traditional reliance on fixed mapping tables, the solution of this application, through the design of incremental learning based on human interactive feedback, can continuously track unknown fields and automatically extract reliable mapping relationships to achieve dynamic updates of the multidimensional financial terminology knowledge base, as described below.

[0061] In step S601 of some embodiments, during the actual processing of complex financial data, with the development of new business, it is inevitable to encounter entirely new fields or obscure non-standard expressions not yet included in the current knowledge base. If, after performing the aforementioned multi-level matching analysis, no target standard field satisfying the mapping conditions can be matched in the multi-dimensional financial terminology knowledge base, a new data structure instance can be created in the background cache for the unresolved field to be identified, and a specific status label can be assigned to it, marking it as a term to be confirmed. This action is designed to effectively contain and isolate unknown data that deviates from existing rules.

[0062] In step S602 of some embodiments, once the field to be identified is intercepted by the system and marked as a term to be confirmed, its subsequent mapping flow inevitably requires the intervention of external experts or business personnel with financial qualifications to complete the final filling task. During this period, the lifecycle of the term to be confirmed is continuously recorded. When the user manually assigns a standard field to the unidentified term to be confirmed in the system's interactive interface, the specific manual correction method for the term to be confirmed in this subsequent interactive processing flow is recorded (for example, recording which accounting subject level in the knowledge base the operator specifically mapped it to). These detailed records of manual correction actions and their corresponding contextual parameters constitute the raw sample data for the system's subsequent machine autonomous learning and rule extraction.

[0063] In step S603 of some embodiments, a dual verification learning mechanism is designed to ensure the security and anti-pollution of knowledge base evolution. The first evolution path is explicit direct learning, where the system directly responds to proactive correction instructions submitted by users with advanced management privileges in the interactive interface for the term to be confirmed (e.g., a finance manager proactively adds a new term as a synonym for a standard field). The second evolution path is statistical verification learning, which maintains a historical operation counter variable for each term to be confirmed. When scanning the historical operation log, it is detected that the same term to be confirmed has been manually corrected by business personnel in the same way in multiple consecutive independent processing batches or tasks, and the cumulative number of such consistent correction actions reaches or exceeds the second preset threshold (configurable to 3 times) set by the system, it can be determined that the correction behavior has the universality of financial operations. Once any of the above verification conditions are passed, the term to be confirmed can be structurally encapsulated and bound to the repeatedly verified manual correction method, thereby automatically generating a standardized correction mapping relationship.

[0064] Please see Figure 7 In some embodiments, steps S701 to S704 may be included before step S202.

[0065] Step S701: Identify whether there is a composite index relationship in the field to be identified.

[0066] Step S702: If a composite index relationship exists, the preset initial formula rule library is called to decompose and parse the composite index relationship for matching.

[0067] Step S703: During the processing of multiple independent financial reporting tasks, record the manual confirmation operation of the equivalence relationship submitted for the new combination expression.

[0068] Step S704: In response to the detection that the number of times the same novel combination expression is repeatedly confirmed as the same equivalence relation reaches a first preset number threshold, the novel combination expression and the corresponding same equivalence relation are added to the extended formula library.

[0069] In some embodiments, in order to cope with the complex calculation logic derived from basic accounts in financial statements, and to achieve accurate parsing and dynamic learning of composite indicators, thereby solving the technical problem that traditional systems can only handle single-field mappings and cannot understand derived formulas, leading to matching failures, the solution of this application, through the built-in initial formula rules and the introduction of incremental learning based on the frequency statistics of human interaction, can adaptively decompose and expand the relationships of composite indicators to obtain a financial feature set with strong logical extensibility, as described below.

[0070] In step S701 of some embodiments, in financial business practices, the initial financial data to be processed not only includes basic items of a single dimension (such as "accounts receivable"), but also often includes derived indicators composed of multiple basic items combined through mathematical operations. Composite indicator relationships refer to financial expression features that include arithmetic logic such as addition, subtraction, multiplication, and division, or specific summary hierarchy logic (e.g., the deduction relationship between "net profit" and "total profit" and "income tax expense"). After extracting the field to be identified, in the pre-stage of performing regular matching analysis, specific regular expressions are used to scan and structurally analyze the text characters and potential logical modifiers of the field to identify and determine whether the field to be identified implicitly or explicitly contains composite indicator relationships, thereby classifying it from conventional single fields.

[0071] In step S702 of some embodiments, the initial formula rule base is a built-in set of basic calculation logic, which pre-encapsulates standard calculation equations commonly used in the financial field (such as the baseline rule "Revenue = Main Business Revenue + Other Business Revenue"). When it is determined in the preceding step that the field to be identified has a composite indicator relationship, the initial formula rule base will be invoked. Specifically, the composite indicator relationship is passed as an input parameter to the parsing engine, which decomposes it downwards into multiple basic fields according to the baseline rules in the initial formula rule base, and performs feature alignment and parsing matching on these decomposed basic fields respectively. This decomposition action can transform complex derivative combination problems into known unit subject matching problems, ensuring the parsing accuracy in basic formula scenarios.

[0072] In step S703 of some embodiments, as the dimensions of enterprise financial analysis become more refined, new combined expressions that are not covered by the initial formula rule base often appear in actual business (for example, a non-standard definition of "Company A's total revenue" being equal to the sum of several specific business revenues). When the existing initial formula rule base cannot be used to automatically complete the parsing, the processing task will trigger an interactive feedback mechanism for business personnel to intervene. In the process of processing multiple independent financial reporting tasks across batches, once the business personnel manually specify the corresponding calculation equation or mapping combination for the unrecognized new combined expression in the terminal interactive interface, the new combined expression and its corresponding manual equivalence relation confirmation operation are recorded as a structured behavior record and cached in the operation log database.

[0073] In step S704 of some embodiments, to ensure that the new formula learned by the system has the universality and accuracy of financial logic, a consistency check is performed on the cached historical operation records. An internally defined constraint parameter is a first preset threshold number (configurable to 3 times), which is the cumulative number of times the same novel combination expression in the dynamic statistical log database is identified as having the exact same equivalence relation by business personnel in different independent tasks. In response to detecting that the cumulative number of occurrences reaches or exceeds the first preset threshold number, it is determined that the equivalence relation possesses industry consensus or a common mapping standard within the enterprise. Based on this determination, the novel combination expression and its corresponding identical equivalence relation are automatically converted and encapsulated into a standard calculation rule format and written into the extended formula library. This extended formula library will work collaboratively with the initial formula rule library in subsequent iterations, jointly participating in the formula parsing operation of new batches of data.

[0074] Through steps S701 to S704 above, the embodiments of this application effectively eliminate the technical problem that traditional systems cannot handle composite indicators by matching and decomposing the initial rules. By utilizing multi-independent task interaction verification based on frequency statistics, the underlying rule base is dynamically updated without source code reconstruction. This ensures the stability and security of the core parsing logic during the initial deployment and gives it the ability to continuously expand the combined semantic understanding as complex financial business evolves, thereby improving the flexibility and long-term applicability of the intelligent data entry system.

[0075] Please see Figure 8 In some embodiments, after step S204, there may be steps S801 to S803, among others.

[0076] Step S801: Extract the original input stream for the initial financial data, the intermediate decision path including initial matching and semantic matching, the manually modified records, and the final output format of the target financial data.

[0077] Step S802: Add timestamps, operator identification, and explanations of reasons for changes to the original input stream, intermediate decision paths, manual modification records, and final output format, respectively.

[0078] Step S803: Encapsulate the record information with the added identifier to generate a full-link operation log chain.

[0079] In step S801 of some embodiments, the business form and semantic structure of the data undergo multiple calculations and transformations throughout the complex financial data processing lifecycle. To achieve a retrospective and traceable review of the entire processing process, a snapshot backup of the initial financial data without any processing is first performed to form the original input stream. After the data enters the multi-stage matching engine, the intermediate decision path generated by the algorithm model is recorded in real time. This path records in detail the lightweight screening trajectory of the preliminary matching analysis stage and the feature alignment logic and confidence change process of the deep semantic matching stage. At the same time, for data nodes that trigger expert review or manual intervention, the specific modification actions of business personnel are collected synchronously and saved as manual modification records. Finally, when the data completes all verifications and is mapped to the output, the final output format of the target financial data after structured mapping is obtained.

[0080] In step S802 of some embodiments, simply extracting state data is insufficient to constitute rigorous audit evidence; it is also necessary to assign spatiotemporal attributes to these discrete state transition data. Specifically, all business node records extracted in the preceding steps are obtained, and the system's underlying time server is invoked to accurately attach a timestamp to each record to solidify the time points of various operations. Simultaneously, by reading the environment variables of the current process or the session credentials of the interactive client, information such as the employee ID of the automated algorithm module executing the action or the specific interactive personnel is extracted and bound to the corresponding record as an operator's identity identifier. In addition, a change reason explanation with business interpretation properties is generated based on the preset rule code that triggered the state change or the notes submitted manually on the interactive interface. By strongly correlating these dimensional attribute characteristics with the underlying state data, it is ensured that the flow of each piece of financial data has evidentiary attributes.

[0081] In step S803 of some embodiments, if the aforementioned discrete records with added time and identity attributes are stored in plaintext, they are easily tampered with under external interference, thus losing their validity as audit credentials. Therefore, in this step, these record information with added multi-dimensional identifiers are used as input and encapsulated as a whole using structured sequence technology. This encapsulation process aims to encapsulate the interrelated original inputs, intermediate logic, modification trajectories, and output results in a time series, forming a continuous, tamper-proof, end-to-end operation log chain. This chain record is then persistently written to and stored in the system's secure storage medium. Any attempt to modify historical conversion logic or conceal illegal actions will destroy the integrity structure of the log chain, thereby triggering system-level security blocking and warnings.

[0082] By combining steps S801 to S803 above, this embodiment of the application realizes full-cycle behavior tracking of financial data from original input to final mapping output. By extracting and encapsulating the timestamp, operator identity and reason for change in a structured manner, the objectivity, authenticity and immutability of the operation log are guaranteed. While improving the transparency of financial data governance, it effectively solves the technical problems of difficult post-event compliance audits and high accountability costs caused by the black box of business processing.

[0083] For complex business scenarios such as merging multi-source heterogeneous financial data and intelligent data entry, a mechanism combining preliminary fuzzy screening and deep semantic feature analysis is first used to accurately analyze the deep meaning and syntactic features of complex expressions such as synonyms, abbreviations, and industry terms to output high-precision matching confidence. Then, this confidence is combined with risk control rules such as core business attributes to implement a dynamic intervention mechanism including automatic filling, alternative recommendations, expert review, and anomaly blocking. This can effectively improve the matching accuracy of non-standard unfamiliar fields and significantly reduce the burden of manual full-scale review while accurately blocking reverse errors or anomalies that violate financial logic. This can solve the technical problems of existing financial data processing solutions, which lack semantic understanding and fault tolerance mechanisms, resulting in low accuracy of heterogeneous data matching and high cost of manual intervention.

[0084] Through steps S201 to S204 above, this application embodiment addresses complex business scenarios such as merging and intelligently filling in multi-source heterogeneous financial data. First, it uses a linkage mechanism of preliminary fuzzy coarse screening and deep semantic feature analysis to accurately parse the deep meaning and syntactic features of complex expressions such as synonyms, abbreviations, and industry terms to output a high-precision matching confidence score. Then, it combines this confidence score with risk control rules such as core business attributes to implement a dynamic intervention mechanism including automatic filling, alternative recommendations, expert review, and anomaly blocking. This can effectively improve the matching accuracy of non-standard unfamiliar fields and significantly reduce the burden of manual full-scale review while accurately intercepting reverse errors or anomalies that violate financial logic. It can solve the technical problem that existing financial data processing solutions lack semantic understanding and fault tolerance mechanisms, resulting in low accuracy of heterogeneous data matching and high cost of manual intervention.

[0085] As can be seen in the above scheme, please refer to Figure 9 This application also provides a financial data processing apparatus that can implement the above-described financial data processing method. The apparatus includes: The acquisition module is used to acquire the initial financial data to be processed and extract the fields to be identified from the initial financial data; The preliminary matching module is used to perform preliminary matching analysis between the field to be identified and the standard fields in the multidimensional financial terminology knowledge base to obtain candidate standard fields and their corresponding preliminary confidence levels. The determination module is used to extract the semantic embedding vector of the field to be identified when the initial confidence does not meet the preset high confidence condition, and calculate the semantic similarity between the semantic embedding vector and the standard field in combination with the context syntactic roles to determine the final matching confidence. The generation module is used to trigger corresponding hierarchical response strategies to map the fields to be identified to the corresponding target standard fields based on the final matching confidence level and preset subject risk control rules, thereby generating target financial reporting data. The hierarchical response strategies include automatic filling strategy, alternative recommendation strategy, expert review strategy, and anomaly blocking strategy.

[0086] In some embodiments, the preliminary matching module is specifically used for: The edit distance and prefix wildcard algorithms are used to perform preliminary fuzzy matching and screening between the field to be identified and the standard field, and the candidate standard fields and their corresponding preliminary confidence scores are calculated. If the initial confidence level meets the preset high confidence level condition, the candidate standard field is directly used as the target standard field, and the initial confidence level is used as the final matching confidence level.

[0087] In some embodiments, the determining module is specifically used for: If the initial confidence level does not meet the preset high confidence level condition, then a semantic embedding vector of the field to be identified is generated based on the pre-trained financial domain-specific language model. Calculate the semantic similarity score between the semantic embedding vector and the candidate standard field, and extract the contextual syntactic role of the field to be identified in the original sentence; Contextual syntactic roles are used to filter out match interference items in order to determine the final match confidence.

[0088] In some embodiments, the generation module is specifically used for: If the final match confidence score is greater than or equal to the first confidence score threshold, the autofill strategy is triggered, and the candidate standard field corresponding to the final match confidence score is directly used as the target standard field for filling. If the final matching confidence is less than the first confidence threshold but greater than or equal to the second confidence threshold, the alternative recommendation strategy is triggered, a recommendation prompt containing at least one alternative candidate standard field is generated and displayed on the interactive interface, and the target standard field is determined based on the received confirmation instruction. If a core business attribute conflict is detected between the field to be identified and the corresponding candidate standard field according to the subject risk control rules, the expert review strategy is triggered to intercept the current automatic mapping process and transfer the corresponding data to the expert review queue for manual determination of the target standard field. If the data characteristics of the field to be identified fail the data logic consistency check associated with the candidate standard field, the abnormal blocking strategy is triggered according to the subject risk control rules, the current processing flow is forcibly blocked and an alarm message is output.

[0089] In some embodiments, the preliminary matching module is further configured to: Identify whether there are composite indicator relationships in the fields to be identified; If a composite indicator relationship exists, the pre-set initial formula rule base is called to decompose, parse, and match the composite indicator relationship; During the processing of multiple independent financial reporting tasks, the manual confirmation of equivalence relationships submitted for new combined expressions is recorded. In response to the detection that the number of times the same novel combinatorial expression is repeatedly confirmed as the same equivalence relation reaches a first preset threshold, the novel combinatorial expression and the corresponding same equivalence relation are added to the extended formula library.

[0090] In some embodiments, the generation module is further configured to: If the field to be identified does not match a target standard field that meets the conditions in the multidimensional financial terminology knowledge base, the field to be identified will be marked as a term to be confirmed. Record the manual correction methods used for terms awaiting confirmation in subsequent interactive processing flows; In response to receiving an active correction instruction for the term to be confirmed in the interactive interface, or detecting that the same term to be confirmed has been manually corrected in the same way in multiple consecutive independent processing tasks and the number of corrections has reached a second preset number threshold, a correction mapping relationship between the term to be confirmed and the manual correction method is generated. The terms to be confirmed and their corresponding correction mappings are dynamically written into and updated as official terms in the multidimensional financial terminology knowledge base.

[0091] In some embodiments, the generation module is further configured to: Extract the raw input stream of the initial financial data, the intermediate decision path including initial matching and semantic matching, manually modified records, and the final output format of the target financial data; Add timestamps, operator identification, and explanations of reasons for changes to the original input stream, intermediate decision paths, manually modified trajectories, and final output formats, respectively. The recorded information with the added identifier is encapsulated to generate a full-link operation log chain.

[0092] This invention provides a financial data processing device. First, through a linkage mechanism of preliminary fuzzy screening and deep semantic feature analysis, it accurately analyzes the deep meaning and syntactic features of complex expressions such as synonyms, abbreviations, and industry terms to output a high-precision matching confidence score. Then, it combines this confidence score with risk control rules such as core business attributes to implement a dynamic intervention mechanism including automatic filling, alternative recommendations, expert review, and anomaly blocking. This can effectively improve the matching accuracy of non-standard unfamiliar fields and significantly reduce the burden of manual full-scale review while accurately blocking reverse errors or anomalies that violate financial logic. It can solve the technical problems of existing financial data processing solutions, which lack semantic understanding and fault tolerance mechanisms, resulting in low accuracy of heterogeneous data matching and high cost of manual intervention.

[0093] Specific limitations regarding target object detection can be found in the limitations of the financial data processing method described above, and will not be repeated here. Each module in the aforementioned financial data processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.

[0094] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 10 As shown. The computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a financial data processing method on the server side.

[0095] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 11 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements a client-side function or step based on a financial data processing method.

[0096] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Obtain the initial financial data to be processed and extract the fields to be identified from the initial financial data; A preliminary matching analysis is performed between the field to be identified and the standard fields in the multidimensional financial terminology knowledge base to obtain candidate standard fields and their corresponding preliminary confidence levels. When the initial confidence level does not meet the preset high confidence level condition, the semantic embedding vector of the field to be identified is extracted, and the semantic similarity between the semantic embedding vector and the standard field is calculated in combination with the context syntactic roles to determine the final matching confidence level. Based on the final matching confidence level and the preset subject risk control rules, the corresponding tiered response strategy is triggered to map the fields to be identified to the corresponding target standard fields and generate target financial reporting data. The tiered response strategy includes an automatic filling strategy, an alternative recommendation strategy, an expert review strategy, and an anomaly blocking strategy.

[0097] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0098] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0099] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0100] It should be noted that any AI models, software tools, or components not belonging to this company appearing in the embodiments of this application are merely illustrative examples and do not represent actual use. All user personal information involved in the embodiments of this application has been authorized (with the knowledge and consent) by the relevant parties or has been fully authorized by all parties, and the executing entity may obtain it through various legal and compliant means. The collection, storage, use, processing, transmission, provision, and disclosure of the information, data, and signals involved all comply with relevant laws and regulations and do not violate public order and good morals.

[0101] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A financial data processing method, characterized in that, include: Obtain the initial financial data to be processed, and extract the fields to be identified from the initial financial data; A preliminary matching analysis is performed on the field to be identified and the standard fields in the multidimensional financial terminology knowledge base to obtain candidate standard fields and their corresponding preliminary confidence levels. When the initial confidence level does not meet the preset high confidence level condition, the semantic embedding vector of the field to be identified is extracted, and the semantic similarity between the semantic embedding vector and the standard field is calculated in combination with the context syntactic roles to determine the final matching confidence level. Based on the final matching confidence level and the preset subject risk control rules, the corresponding hierarchical response strategy is triggered to map the field to be identified to the corresponding target standard field and generate target financial reporting data; wherein, the hierarchical response strategy includes an automatic filling strategy, an alternative recommendation strategy, an expert review strategy, and an anomaly blocking strategy.

2. The financial data processing method according to claim 1, characterized in that, The preliminary matching analysis between the field to be identified and the standard fields in the multidimensional financial terminology knowledge base is performed to obtain candidate standard fields and their corresponding preliminary confidence levels, including: The field to be identified and the standard field are initially fuzzy matched and filtered using edit distance and prefix wildcard algorithms, and the candidate standard fields and their corresponding preliminary confidence scores are calculated. If the preliminary confidence level meets the preset high confidence level condition, then the candidate standard field is directly used as the target standard field, and the preliminary confidence level is used as the final matching confidence level.

3. The financial data processing method according to claim 1, characterized in that, When the initial confidence level does not meet the preset high confidence level condition, the semantic embedding vector of the field to be identified is extracted, and the semantic similarity between the semantic embedding vector and the standard field is calculated in combination with the context syntactic roles to determine the final matching confidence level, including: If the initial confidence level does not meet the preset high confidence level condition, then the semantic embedding vector of the field to be identified is generated based on the pre-trained financial domain-specific language model. Calculate the semantic similarity score between the semantic embedding vector and the candidate standard field, and extract the contextual syntactic role of the field to be identified in the original sentence; Based on the contextual syntactic roles, interference items are filtered to determine the final match confidence.

4. The financial data processing method according to claim 1, characterized in that, The step of triggering a corresponding tiered response strategy based on the final matching confidence level and preset subject risk control rules to map the field to be identified to the corresponding target standard field and generate target financial reporting data includes: If the final matching confidence is greater than or equal to the first confidence threshold, the automatic filling strategy is triggered, and the candidate standard field corresponding to the final matching confidence is directly used as the target standard field for filling. If the final matching confidence is less than the first confidence threshold and greater than or equal to the second confidence threshold, the alternative recommendation strategy is triggered, a recommendation prompt containing at least one alternative candidate standard field is generated and displayed on the interactive interface, and the target standard field is determined based on the received confirmation instruction. If a core business attribute conflict is detected between the field to be identified and the corresponding candidate standard field according to the subject risk control rules, the expert review strategy is triggered to intercept the current automatic mapping process and transfer the corresponding data to the expert review queue for manual determination of the target standard field. If the data characteristics of the field to be identified fail the data logic consistency check associated with the candidate standard field, the abnormal blocking strategy is triggered according to the subject risk control rules, forcibly blocking the current processing flow and outputting alarm information.

5. The method according to claim 1, characterized in that, Before performing the preliminary matching analysis between the field to be identified and the standard fields in the multidimensional financial terminology knowledge base, the method further includes: Identify whether there is a composite index relationship in the field to be identified; If the composite index relationship exists, the preset initial formula rule base is called to decompose, parse and match the composite index relationship; During the processing of multiple independent financial reporting tasks, the manual confirmation of equivalence relationships submitted for new combined expressions is recorded. In response to the detection that the number of times the same novel combination expression is repeatedly confirmed as the same equivalence relation reaches a first preset threshold, the novel combination expression and the corresponding same equivalence relation are added to the extended formula library.

6. The financial data processing method according to claim 1, characterized in that, After generating the target financial reporting data, the method further includes: If the field to be identified does not match the target standard field that meets the conditions in the multidimensional financial terminology knowledge base, then the field to be identified is marked as a term to be confirmed. Record the manual correction methods used for the terms to be confirmed in subsequent interactive processing flows; In response to receiving an active correction instruction for the term to be confirmed in the interactive interface, or detecting that the same term to be confirmed has been manually corrected in the same way in multiple consecutive independent processing tasks and the number of corrections has reached a second preset number threshold, a correction mapping relationship between the term to be confirmed and the manual correction method is generated. The terms to be confirmed and their corresponding correction mapping relationships are dynamically written as official terms and the multidimensional financial terminology knowledge base is updated.

7. The financial data processing method according to claim 1, characterized in that, After obtaining the target financial reporting data, the method further includes: Extract the original input stream for the initial financial data, the intermediate decision path including initial matching and semantic matching, the manual modification records, and the final output format of the target financial data; Add timestamps, operator identification, and explanations of reasons for changes to the original input stream, the intermediate decision path, the manual modification record, and the final output format, respectively. The recorded information with the added identifier is encapsulated to generate a full-link operation log chain.

8. A financial data processing device, characterized in that, The device includes: The acquisition module is used to acquire the initial financial data to be processed and extract the fields to be identified from the initial financial data; The preliminary matching module is used to perform preliminary matching analysis between the field to be identified and the standard fields in the multidimensional financial terminology knowledge base to obtain candidate standard fields and their corresponding preliminary confidence levels. The determination module is used to extract the semantic embedding vector of the field to be identified when the initial confidence does not meet the preset high confidence condition, and calculate the semantic similarity between the semantic embedding vector and the standard field in combination with the context syntactic roles to determine the final matching confidence. The generation module is used to trigger a corresponding tiered response strategy to map the field to be identified to the corresponding target standard field based on the final matching confidence level and the preset subject risk control rules, thereby generating target financial reporting data; wherein, the tiered response strategy includes an automatic filling strategy, an alternative recommendation strategy, an expert review strategy, and an anomaly blocking strategy.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the financial data processing method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, characterized in that, when the computer program is executed by a processor, it implements the financial data processing method according to any one of claims 1 to 7.