Method and device for analyzing and checking reimbursement data and medium

By calculating the correlation between the reimbursement association fields and historical fields and combining the bill coding classification and semantic structure chart, the problem of time-consuming and labor-intensive reimbursement order verification and difficulty in matching rules is solved, and efficient and accurate financial data verification is achieved.

CN120278840AActive Publication Date: 2025-07-08SHANDONG LANGCHAO SMART CULTURAL TOURISM IND DEV CO LTD +1

Patent Information

Application Number
CN202510724704.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-07-08
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

In the financial management of the existing technology, reimbursement order verification is time-consuming and labor-intensive, and it is difficult to discover potential logical relationships between data, resulting in false reimbursement or incorrect payments, and it is difficult to dynamically match the review rules for different business scenarios.

Method used

By obtaining the correlation degree values of the reimbursement correlation field and the historical correlation field, filtering the historical data with high correlation degree for comparison, combining the bill encoding classification storage and financial semantic structure chart, multi-dimensional quantitative evaluation and semantic comparison are performed to generate verification results.

Benefits of technology

It improves the efficiency and accuracy of reimbursement order verification, reduces the subjectivity of manual review, ensures the authenticity and compliance of data, and promptly detects abnormal situations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278840A_ABST
    Figure CN120278840A_ABST
Patent Text Reader

Abstract

The invention provides a reimbursement data analysis and check method and device and a medium, and belongs to the technical field of enterprise reimbursement data processing, and the method comprises the steps: obtaining a reimbursement associated field of new financial accounting data, and determining a correlation degree value between the reimbursement associated field and a historical associated field in a financial data set; acquiring historical financial accounting data corresponding to the historical associated fields of which the association degree numerical values are greater than a matching degree threshold value, and taking the acquired historical financial accounting data as comparison financial accounting data; and according to the comparison financial bookkeeping data and the new financial bookkeeping data, determining a reimbursement bill checking result of the new financial bookkeeping data. According to the method, from acquisition of reimbursement associated fields, screening of comparison data to deep comparison based on the financial semantic structure chart, relevance and similarity of new and old data are comprehensively and quantitatively analyzed. While the checking efficiency is improved, the accuracy and reliability are enhanced, and the problems of low efficiency and error proneness of traditional manual checking are effectively avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of enterprise reimbursement data processing, and particularly relates to a method, device, and medium for analyzing and verifying reimbursement data. Background Art

[0002] In enterprise financial management and reimbursement processes, the verification of financial accounting data is a fundamental and crucial task. With the expansion of enterprise scale and the increase in business complexity, the financial system needs to process a large number of reimbursement vouchers every day. These data have a wide range of sources, diverse formats, and inconsistent structures, posing great challenges to verification.

[0003] In related technologies, when verifying reimbursement vouchers, it is necessary to manually or systemically traverse all historical financial data to find comparable objects, which is time-consuming and laborious when facing a large amount of data. For example, large enterprises process thousands of reimbursement vouchers every day. Comparing historical data one by one may lead to an audit cycle of up to several days, affecting the efficiency of capital turnover.

[0004] Some use office software to assist in verification, but only relying on methods such as amount matching and invoice number verification easily ignores potential logical relationships between data. Problems such as inconsistent expense items and actual business scenarios in reimbursement vouchers and incorrect data formats are difficult to detect, which may lead to false reimbursements or incorrect payments, causing economic losses to the enterprise. Existing technologies are difficult to dynamically match audit rules for different business scenarios, such as business travel expenses and procurement expenses. For example, when auditing business travel expenses, it is necessary to pay attention to the rationality of the itinerary, and when auditing procurement expenses, it is necessary to verify the supplier's qualifications. Unified rules cannot meet diverse needs, resulting in some abnormal reimbursement vouchers being missed in the audit. Summary of the Invention

[0005] The present invention provides a method for analyzing and verifying reimbursement data, which improves the audit efficiency compared with manual audits in related technologies; through multi-dimensional quantitative evaluation, it effectively avoids problems such as false reimbursements and data errors, and ensures the authenticity and compliance of enterprise financial data.

[0006] The method includes: S101: Obtain the reimbursement-related fields of the new financial accounting data, and determine the correlation degree value between the reimbursement-related fields and the historical related fields in the financial data set; S102: Obtain the historical financial accounting data corresponding to the historical related fields with a correlation degree value greater than the matching threshold, and use the obtained historical financial accounting data as the comparison financial accounting data; S103: Determine the reimbursement voucher verification result of the new financial accounting data according to the comparison financial accounting data and the new financial accounting data.

[0007] Preferably, determining the correlation degree value between the reimbursement-related fields and the historical related fields in the financial data set in step S101 specifically includes: Obtain the reimbursement-related fields of the new financial accounting data and analyze their matching degree with the historical related fields; Statistically analyze all the matching degree results, determine the highest matching degree and the lowest matching degree as the benchmarks for calculating the correlation degree values; Combine the highest matching degree, the lowest matching degree and the matching degree of each historical field to generate the corresponding correlation degree values; Sort according to the correlation degree values, and select the historical field data higher than the threshold as the comparison data for checking the reimbursement form of the new financial accounting data.

[0008] Preferably, the method also obtains the difference between the highest matching degree and the lowest matching degree as the global range parameter; For each historical related field, calculate the difference between its matching degree with the reimbursement-related field and the lowest matching degree as the local offset; Divide the local offset by the global range parameter to obtain the relative ratio as the correlation degree value; Filter and compare the financial accounting data according to the correlation degree values.

[0009] Preferably, in step S101, obtaining the reimbursement-related fields of the new financial accounting data specifically includes: Clean the invalid records of the new financial accounting data to obtain the preliminarily processed data; According to the data format specifications matching the reimbursement type, verify and correct the data format; Filter out the non-business content, retain the core business data, and generate the preprocessed reimbursement form; Encode the preprocessed data through a structured information parsing tool to generate reimbursement-related fields with a fixed length and clear semantics.

[0010] Preferably, the method further includes: Construct multiple classified storage units through the ticket coding location, and each classified storage unit only contains the historical financial accounting data corresponding to a specific ticket coding; Obtain the new ticket coding of the new financial accounting data; Traverse all the classified storage units, match the new ticket coding with the preset historical ticket coding in the classified storage units, and filter out the target storage unit; Extract the historical related fields and their corresponding historical financial accounting data from the target storage unit; Based on the semantic and syntactic analysis results of the reimbursement-related fields and the historical related fields, calculate the correlation degree values to complete the verification of the reimbursement form.

[0011] Preferably, step S103 further includes: Obtain the financial semantic structure diagram of the new financial accounting data and the historical financial semantic template of the comparison financial accounting data; Determine the verification result of the expense report for the new financial accounting data according to the financial semantic structure diagram and the historical financial semantic template.

[0012] Preferably, step S103 further includes: Classify the new financial accounting data into a preset business scenario classification system by analyzing the prefix of the bill code, keyword matching, and form title of the new financial accounting data; For the classified new financial data, identify the core business entities and their associated relationships therein; Map the extracted entities to a predefined financial semantic meta-model, unify the entity expression form, standardize the attribute value format, and construct a standardized financial semantic structure diagram; Based on the business scenario classification of the new financial data, retrieve the set of historical templates of the same type from the historical financial semantic template library to form a comparison sample pool; Compare the topological structures of the newly constructed financial semantic structure diagram and the historical template, count the matching ratio of the same entity types and relationship paths, and generate a structural similarity score; For the matching entity nodes, verify the consistency of their attribute values, and mark the abnormal levels for the nodes with differences; Combine the structural similarity score and the attribute consistency verification result to generate a multi-dimensional comparison report. When the similarity exceeds the preset threshold and there are no abnormalities in the key attributes, it is determined that the expense report passes the verification.

[0013] Preferably, step S102 also extracts the financial semantic structure diagram of the new financial accounting data; the financial semantic structure diagram shows the relationships between financial data in the form of nodes and edges, where the nodes represent different financial entities and the edges represent the associations between entities; Retrieve the historical financial semantic template; the historical financial semantic template stores the structured information of past financial data in a standardized format, also presented in the form of nodes and edges, and has been verified to comply with the established financial rules and reimbursement policies; Through the semantic comparison algorithm, compare the financial semantic structure diagram of the new financial accounting data with the historical financial semantic template item by item; During the comparison process, pay attention to the consistency and differences in node types, node attributes, edge connection relationships, and data flow paths; According to the preset matching rules and matching degree thresholds, calculate the matching degree between the two; When the matching degree reaches a certain threshold, it is considered that the new data highly matches a certain historical data. Based on this, combined with the reimbursement rules, judge the compliance and accuracy of the new data, so as to obtain the verification result of the expense report.

[0014] According to another embodiment of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the reimbursement data analysis and verification method when executing the program.

[0015] According to another embodiment of the present application, a storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the reimbursement data analysis and verification method are implemented.

[0016] It can be seen from the above technical solutions that the present invention has the following advantages: The reimbursement data analysis and verification method provided by the present invention establishes a quantitative association between new financial accounting data and historical data by extracting reimbursement related fields and calculating the correlation value. Accurately screen out historical data with high correlation as comparison samples to avoid blindly searching in all historical data, greatly improving data retrieval efficiency. Filter historical financial accounting data according to the correlation threshold, further narrow the data range, focus on the most valuable comparison data, and improve the pertinence of the verification.

[0017] Business scenarios are also classified based on bill codes, keywords, and form titles to quickly locate the business type to which new data belongs, laying the foundation for the precise application of subsequent audit rules. Avoid confusion of data in different business scenarios and reduce the problem of mismatching of audit rules. Identify core business entities and relationships, and achieve standardization through semantic metamodel mapping to eliminate differences in data representation and format. Unified data expression facilitates cross-data comparison and improves data consistency and standardization. Retrieve and compare historical templates of the same type, use historical compliance data to quickly determine the rationality of new data structures, and discover abnormal business structures in a timely manner. New data and historical data are also converted into visual semantic structure diagrams and templates, clearly presenting data logic in the form of nodes and edges. The semantic comparison algorithm compares each item from multiple dimensions such as nodes, edges, and data flow, combined with matching rules and threshold quantitative judgments, which not only ensures the comprehensiveness of the comparison, but also improves the efficiency and accuracy of the audit, and reduces the subjectivity and omissions of manual audits. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solution of the present invention, the accompanying drawings required for use in the description will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0019] Figure 1 A flowchart of the verification method for reimbursement data analysis; Figure 2 A flowchart of an embodiment of a method for analyzing and checking reimbursement data; Figure 3 Flowchart of another embodiment of the method for analyzing and verifying reimbursement data Figure 4 Schematic diagram of an electronic device Detailed implementation manners

[0020] The following will describe in detail the specific steps of the method for analyzing and verifying reimbursement data involved in this application. For the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are proposed to thoroughly understand the embodiments of this application. However, those skilled in the art should clearly understand that this application can also be implemented in other embodiments without these specific details.

[0021] It should be understood that when used in the specification of this application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0022] The statements such as "in one embodiment" or "in some embodiments" described in this application mean that the specific features, structures, or characteristics described in the embodiment are included in one or more embodiments of this application. Thus, the statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" and the like that appear in different parts of this application do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways.

[0023] In the embodiments of the present invention, computer program code for performing the operations of the present disclosure can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include but are not limited to object-oriented programming languages - such as Java, Smalltalk, C++, and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code can be executed completely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or completely on a remote computer. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, alternatively, can be connected to an external computer (exemplarily, by using an Internet service provider to connect through the Internet).

[0024] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0025] Please refer to Figure 1 The figure shows a flowchart of a method for analyzing and verifying reimbursement data in a specific embodiment. The method includes: S101: Obtain the reimbursement-related fields of the new financial accounting data, and determine the correlation degree value between the reimbursement-related fields and the historical related fields in the financial dataset.

[0026] In some embodiments, key information is extracted from the newly submitted financial accounting data to form reimbursement-related fields, which include but are not limited to the name of the reimburser, reimbursement date, expense item, amount, invoice number, etc. Then, these newly extracted reimbursement-related fields are compared one by one with the historical related fields stored in the financial dataset. During the comparison process, for each historical related field, analyze its similarity or matching degree with the new field in each dimension, and finally obtain a quantified correlation degree value to represent the degree of closeness between the new field and each historical field.

[0027] In some specific embodiments, the specific process further includes: Step 1011: Clean the invalid records in the new financial accounting data to obtain the preliminarily processed data.

[0028] Step 1012: Match the data format specifications according to the reimbursement type, and verify and correct the data format, such as unifying the currency unit and date format.

[0029] Step 1013: Filter out non-business content, retain the core business data, and generate a preprocessed reimbursement form.

[0030] Step 1014: Encode the preprocessed data through a structured information parsing tool to generate reimbursement-related fields with a fixed length and clear semantics.

[0031] In this embodiment, a structured information parsing tool is used to encode the preprocessed reimbursement form to generate reimbursement-related fields with a fixed length and standardized semantics. This process is based on predefined data format specifications to ensure the uniformity of reimbursement-related fields in different financial accounting scenarios. For example, for a certain type of reimbursement form, the tool extracts key information such as amount, date, and supplier name, and converts them into standardized codes, such as numeric identifiers or hash values, to form fields of length L. This method solves the problem of difficult field matching caused by data heterogeneity in traditional reimbursement forms by eliminating redundancy and format differences in the original data, thereby improving the efficiency and accuracy of subsequent verification.

[0032] In this embodiment, by obtaining the reimbursement-related fields and calculating the correlation degree values, it is possible to quickly locate the historical data that may be related to the new financial accounting data, narrow the scope of subsequent comparison, and improve the efficiency of data processing. At the same time, the quantified correlation degree values provide a clear basis for subsequent screening, avoiding the subjectivity and arbitrariness of human judgment.

[0033] S102: Obtain the historical financial accounting data corresponding to the historical correlation fields with correlation degree values greater than the matching degree threshold, and use the obtained historical financial accounting data as the comparison financial accounting data.

[0034] In this embodiment, a matching degree threshold is set to distinguish which historical correlation fields have a high enough correlation with the new financial accounting data. According to the magnitude of the correlation degree values, historical correlation fields with correlation degrees greater than this matching degree threshold are screened out from the financial dataset, and the historical financial accounting data corresponding to the historical correlation fields is extracted and used as the comparison financial accounting data. By comparing with the correlation degree values, historical data with low correlation degrees is automatically filtered out, and the part that may be highly correlated with the new data is retained as the basis for subsequent comparison.

[0035] S103: Determine the reimbursement form verification result of the new financial accounting data according to the comparison financial accounting data and the new financial accounting data.

[0036] Compare and analyze the comparison financial accounting data obtained in step S102 with the new financial accounting data. The comparison content covers all aspects of the data, such as the reasonableness of the reimbursement reason, the consistency of the expense amount, the authenticity of the invoice information, etc. According to the pre-set verification rules and standards, judge and evaluate the comparison results, and finally determine whether the reimbursement form of the new financial accounting data meets the requirements, and obtain the reimbursement form verification result, such as passing, failing, or requiring further supplementary information, etc.

[0037] For example, check whether the reimbursement amount is within the budget, whether the invoice is within the validity period and not reimbursed repeatedly, etc. According to the inspection results, make a logical judgment in combination with the rule standards to obtain the final verification conclusion. It can effectively discover the problems and anomalies in the new reimbursement form, ensuring the authenticity, legality and compliance of the reimbursement form.

[0038] Furthermore, as a refinement and extension of the specific implementation manner of the above embodiment, S101 further includes: constructing a plurality of classified storage units through the location of the bill code, and each classified storage unit only contains historical financial accounting data corresponding to a specific bill code; Obtain the new bill code of the new financial accounting data; match the corresponding classified storage unit and filter out the historical associated field data consistent with the bill code. This method avoids the inefficiency of traditional full-volume data retrieval through the classified storage design of the bill code, significantly improving the response speed of the reimbursement form verification. At the same time, the accurate matching mechanism between the classified storage unit and the new bill code ensures that the data range of the correlation analysis focuses on the historical data with consistent business logic.

[0039] It should be noted that the historical financial accounting data is classified according to the location rule of the bill code (for example, the first 3 digits represent the bill type, and the middle 4 digits represent the invoicing area). For example, the accounting data corresponding to all invoices with the first 3 digits being "FP0" is stored in the same unit, forming a classified storage unit of "FP0 bill type" to achieve structured storage of data and improve data retrieval efficiency.

[0040] The system automatically extracts the bill code field from the newly submitted financial accounting data. For example, in the electronic invoice data, the unique bill code is obtained by identifying the characters after a fixed position (such as the 5th - 15th digits at the head of the file) or a specific identifier. Access each classified storage unit in turn, and compare the new bill code with the preset coding rule in the unit. If the code completely matches or conforms to a specific matching logic, that is, the first few digits are the same and it is regarded as the same type of bill, then the classified storage unit is marked as the target storage unit. After determining the target storage unit, the system reads all the historical financial accounting data in the unit and extracts the key associated fields (such as the reimburser, amount, expense item, etc.). These fields correspond to the reimbursement associated field structure of the new data, facilitating subsequent comparison. Conduct a dual analysis of the associated fields of the new data and the historical data: semantic analysis to judge whether the field meanings are consistent (such as "travel expenses" and "business trip expenses" are regarded as the same type); syntactic analysis to check whether the field formats match (such as whether the amount is in digital format and whether the date conforms to the standard format). According to the analysis results, comprehensively evaluate the degree of tightness of the association between the fields and obtain a quantitative correlation value. Through dual verification of semantics and syntax, the verification accuracy is improved.

[0041] In an embodiment of the present invention, based on determining the correlation degree value between the reimbursement associated field in step S101 and the historical associated fields in the financial dataset, a possible embodiment will be given below to non - restrictively elaborate on its specific implementation.

[0042] Obtain the reimbursement associated field of the new financial accounting data and analyze its matching degree with the historical associated fields. Count all the matching degree results, determine the highest matching degree and the lowest matching degree as the benchmarks for calculating the correlation degree value. Combine the highest matching degree, the lowest matching degree and the matching degree of each historical field to generate the corresponding correlation degree value. Sort according to the correlation degree value and select the historical field data higher than the threshold as the comparison data for checking the reimbursement form of the new financial accounting data.

[0043] The method of this embodiment also obtains the difference between the highest matching degree and the lowest matching degree as the global range parameter; for each historical associated field, calculate the difference between its matching degree with the reimbursement associated field and the lowest matching degree as the local offset; divide the local offset by the global range parameter to obtain the relative ratio as the correlation degree value; screen and compare the financial accounting data according to the correlation degree value.

[0044] It should be noted that the global range parameter is obtained by calculating the difference between the highest matching degree and the lowest matching degree, which characterizes the matching degree distribution range of all historical associated fields. This parameter is used to quantify the overall fluctuation range of historical data. The local offset is calculated for each historical associated field, which is the difference between the matching degree of the reimbursement associated field and the lowest matching degree, characterizing the relative position of the current field in the overall range. The relative ratio is obtained by dividing the local offset by the global range parameter, which is the relative ratio between the reimbursement associated field and the historical associated field. This ratio reflects the normalized position of the current field's matching degree in the overall distribution, thereby generating the correlation degree value.

[0045] It can be seen that reimbursement-related fields, such as key information like the reimburser, reimbursement amount, expense item, invoice number, etc., are extracted from the newly submitted financial accounting data. Then, these new fields are compared one by one with the related fields in the historical financial data to analyze the matching degree between the two, and the matching degree result corresponding to each historical field is obtained. Subsequently, all the matching degree results are statistically analyzed to find the highest and lowest matching degree values, and these two values will be used as the benchmarks for subsequent calculation of the correlation degree. Then, by combining the highest matching degree, the lowest matching degree, and the matching degree of each historical field itself, the corresponding correlation degree value is generated through a specific calculation method, and this value quantitatively reflects the degree of closeness of the association between the new field and each historical field. Finally, the correlation degree values of all historical fields are sorted, a threshold is set, and the historical fields with correlation degree values higher than this threshold and their corresponding financial accounting data are selected as the comparison data for checking against the new financial accounting data, thus completing the preliminary data screening work for the reimbursement form check.

[0046] For example, fields such as "reimburser's name", "amount value", "invoice number string", etc. are extracted from the electronic reimbursement form. Then, for each historical related field, different matching rules are set according to the field type: for text fields (such as the reimburser's name), a character similarity algorithm is used to determine whether they are the same or similar; for numerical fields (such as the amount), an allowable error range (such as ±5%) is set to judge the closeness of the numerical values; for numbered fields (such as the invoice number), exact matching is required. By comparing the new fields with the historical fields one by one according to these rules, the matching degree of each historical field is obtained.

[0047] Of course, in this embodiment, the formula "(matching degree of a certain historical field - lowest matching degree) / (highest matching degree - lowest matching degree)" can also be used to calculate the relative proportion, and this proportion is used as the correlation degree value. This process maps the original matching degree to the interval of 0 - 1. The closer the value is to 1, the closer the association between the historical field and the new field. The matching degree results of different magnitudes and dimensions are converted into a unified standard quantitative index, eliminating the interference caused by data differences and improving the data processing efficiency.

[0048] In this embodiment, the difference between the highest matching degree and the lowest matching degree is first calculated. This difference reflects the overall fluctuation range of the matching degrees of all historical fields and is used as a global range parameter. Then, for each historical related field, the difference between its matching degree and the lowest matching degree is calculated to obtain a local offset. This offset represents the degree of deviation of this field from the overall matching level. Finally, the local offset is divided by the global range parameter to obtain a relative proportion, and this proportion is the correlation degree value, which is used to measure the association strength between this historical field and the new field. This avoids the influence of a single extreme value on the overall judgment and further improves the accuracy of screening and comparing data.

[0049] In an embodiment of the present application, in order to obtain the reimbursement form verification result more accurately, the present application adopts a reimbursement form verification mechanism, and step S103 further includes: Obtain the financial semantic structure diagram of the new financial accounting data and the historical financial semantic template for comparing the financial accounting data. Determine the reimbursement form verification result of the new financial accounting data according to the financial semantic structure diagram and the historical financial semantic template.

[0050] It should be noted that the financial semantic structure diagram is a structured semantic network, where nodes represent business entities (such as "reimburser", "amount"), and edges represent the relationships between entities (such as "belong to", "associate with"). Working process: Identify the key entities and relationships in the reimbursement form through semantic analysis, construct a directed graph structure, and label each node with the entity type (such as "personnel", "expense item") and attribute value (such as "Zhang San", "500 yuan").

[0051] The historical financial semantic template is to perform cluster analysis on historical reimbursement forms, extract high-frequency entities and relationship patterns, and form a standardized template library. Each template represents a typical reimbursement business (such as "business trip reimbursement", "office supplies procurement").

[0052] Business scenario classification is based on analyzing the prefix of the bill code (such as "TRAVEL-" represents business trip expenses) or keywords (such as "air ticket", "accommodation fee"), and classifying the reimbursement form into a preset business scenario (such as "business trip", "procurement", "conference").

[0053] Business entity normalization is to map synonymous entities to standard terms (such as "computer", "electronic device"); normalize the numerical format (such as "¥5,000", "5000.00"); remove format symbols (such as "Invoice number: FP20250525" → "FP20250525") to ensure the consistency of entity expression.

[0054] The business scenario recognition in this embodiment determines the business scenario (such as "business trip expenses", "office supplies procurement") to which the new financial accounting data belongs by analyzing the prefix of the bill code or keywords. The semantic structure construction is based on the identified business scenario, extracts the entities (such as "reimburser", "amount", "supplier") and their relationships in the new financial data, and constructs a financial semantic structure diagram. Template matching is to select the historical template that matches the current business scenario from the historical financial semantic template library to form a comparison sample set.

[0055] In this embodiment, the business scenario classification of the new financial accounting data is determined according to the bill code of the new financial accounting data. Based on the business scenario classification, the target parsing library is determined. The new financial accounting data and the comparison financial accounting data are parsed based on the target parsing library to obtain the financial semantic structure diagram of the new financial accounting data and the historical financial semantic template of the comparison financial accounting data. When the target parsing library parses the new financial accounting data and the comparison financial accounting data, business entity standardization is required to ensure the consistency of the structure in the semantics.

[0056] For each node in the financial semantic structure diagram and the historical template in this embodiment, the type (such as "personnel", "expense item") and the standardized attribute value (such as "Zhang San", "500.00") are extracted.

[0057] The node type and the attribute value are combined to generate a hash value, which is used as the unique identifier of the node. The node hash values of the new financial data and the historical template are compared, and the ratio of the same nodes is counted as the structure matching degree. The historical template with the highest matching degree is selected as the comparison result. If the matching degree exceeds the threshold (such as 95%), it is determined that the reimbursement form is compliant.

[0058] On the basis of the above embodiment, in order to further improve the reliability of the reimbursement data analysis and verification method provided in the above embodiment, the following is a more specific implementation manner. The specific method includes: S901, preprocess the new financial accounting data to obtain the preprocessed financial accounting data.

[0059] S902, perform coding processing on the preprocessed financial accounting data based on the structured information parsing tool to obtain the reimbursement-related fields of the new financial accounting data.

[0060] S903, obtain the new bill code of the new financial accounting data.

[0061] S904, select the financial data set from the classification storage unit according to the new bill code and the historical bill codes of the historical financial accounting data corresponding to each historical-related field in the classification storage unit. The historical bill codes of the historical financial accounting data corresponding to each historical-related field in each classification storage unit are the same; the historical bill codes of the historical financial accounting data corresponding to each historical-related field in the financial data set are the same as the new bill code.

[0062] S905, determine the matching degree between the reimbursement-related field and the historical-related field in the financial data set.

[0063] S906, select the highest matching degree and the lowest matching degree from each matching degree.

[0064] S907, use the difference between the highest matching degree and the lowest matching degree as the global range parameter.

[0065] S908. For each historical associated field, calculate the difference between its matching degree and the minimum matching degree with the reimbursement associated field as the local offset.

[0066] S909. Divide the local offset by the global range parameter to obtain the relative ratio as the association degree value.

[0067] S910. Obtain the historical financial accounting data corresponding to the historical associated fields with the association degree value greater than the matching degree threshold, and use the obtained historical financial accounting data as the comparison financial accounting data.

[0068] S911. Obtain the financial semantic structure diagram of the new financial accounting data and the historical financial semantic template of the comparison financial accounting data.

[0069] S912. Determine the verification result of the reimbursement form for the new financial accounting data according to the financial semantic structure diagram and the historical financial semantic template.

[0070] As Figure 2 shown, for step S103 of this application, in addition to the above method, there is also another implementation method as follows: Step S1031: Classify the new financial accounting data into a preset business scenario classification system by analyzing the prefix of the bill code, keyword matching, and form title of the new financial accounting data. Step S1032: Identify the core business entities and their associated relationships in the classified new financial data. Step S1033: Map the extracted entities to a predefined financial semantic meta-model, unify the entity expression form, and standardize the attribute value format to construct a standardized financial semantic structure diagram. Step S1034: Based on the business scenario classification of the new financial data, retrieve the set of historical templates of the same type from the historical financial semantic template library to form a comparison sample pool. Step S1035: Compare the topological structures of the newly constructed financial semantic structure diagram and the historical template, count the matching ratio of the same entity types and relationship paths, and generate a structure similarity score. Step S1036: Verify the consistency of the attribute values for the matching entity nodes, and mark the abnormal levels for the nodes with differences. Step S1037: Combine the structure similarity score and the attribute consistency verification result to generate a multi-dimensional comparison report. When the similarity exceeds the preset threshold and there are no abnormalities in the key attributes, it is determined that the reimbursement form passes the verification.

[0071] It can be seen that by analyzing the invoice code prefix, keyword matching, and form title of the new financial accounting data, it is classified into a preset business scenario classification system, laying a foundation for subsequent processing. For the classified data, the core business entities and their association relationships are identified, and the key elements and interactions of the data are clarified. Then, the extracted entities are mapped to a predefined financial semantic meta-model to unify the entity expression form and standardize the attribute value format, constructing a standardized financial semantic structure diagram for structured comparison. Based on the business scenario classification of the new financial data, a set of historical templates of the same type is retrieved from the historical financial semantic template library to form a comparison sample pool, providing historical data references. By comparing the topological structures of the newly constructed financial semantic structure diagram and the historical templates, the matching ratio of the same entity types and relationship paths is statistically calculated to generate a structural similarity score, quantifying the similarity degree. For the matching entity nodes, the consistency of their attribute values is verified, and the nodes with differences are marked with abnormal levels to identify potential problems. Finally, combining the structural similarity score and the attribute consistency verification results, a multi-dimensional comparison report is generated. When the similarity exceeds the preset threshold and there are no abnormalities in the key attributes, it is determined that the reimbursement form passes the verification, giving a clear conclusion.

[0072] In this way, mapping the extracted entities to a predefined financial semantic meta-model and standardizing the attribute value format can unify the data expression form, eliminate the comparison errors caused by inconsistent formats, and retrieving a set of historical templates of the same type from the historical financial semantic template library to form a comparison sample pool can provide historical reference data matching the new data business scenario, increasing the pertinence and effectiveness of the comparison. By comparing the topological structures of the newly constructed financial semantic structure diagram and the historical templates and generating a structural similarity score, the similarity degree between the new and old data can be quantified, providing an objective basis for judging the compliance and accuracy of the reimbursement form. Verifying the consistency of the attribute values of the matching entity nodes and marking the abnormal levels for the different nodes can accurately identify the specific differences between the new and old data, and timely discover possible errors or abnormal situations. Combining the structural similarity score and the attribute consistency verification results to generate a multi-dimensional comparison report can comprehensively present the comparison situation of the new and old data, giving a clear verification conclusion. When the similarity and key attributes meet the conditions, it is determined that the reimbursement form passes the verification, effectively improving the automation degree of reimbursement form processing and the accuracy of decision-making.

[0073] As Figure 3 shown, as a way of this application, step S102 further includes the following implementation method: Step S4011: Extract the financial semantic structure diagram of the new financial accounting data; the financial semantic structure diagram shows the relationship between financial data in the form of nodes and edges, where the nodes represent different financial entities and the edges represent the associations between entities; Step S4012: Retrieve a historical financial semantic template; the historical financial semantic template stores structured information of past financial data in a standardized format, also in the form of nodes and edges, and is verified to comply with established financial rules and reimbursement policies; Step S4013: using a semantic comparison algorithm, the financial semantic structure diagram of the new financial accounting data is compared item by item with the historical financial semantic template; Step S4014: During the comparison process, attention is paid to the consistency and difference of node types, node attributes, edge connection relationships, and data flow paths; Step S4015: Calculate the degree of match between the two according to the preset matching rules and matching threshold. When the matching degree reaches a certain threshold, it is considered that the new data is highly matched with a certain historical data. Based on this, combined with the reimbursement rules, the compliance and accuracy of the new data are judged, thereby obtaining the reimbursement form verification result.

[0074] For example, a financial semantic structure graph is extracted from new financial accounting data. This involves in-depth analysis of the data, identifying key financial entities such as income, expenditure, and accounts as nodes, and relationships between these entities such as capital flow and ownership as edges.

[0075] Step S4012 retrieves historical financial semantic templates from the existing historical data storage. These templates are not only standardized in format, but also consistent with the company's financial rules and reimbursement policies, and are also constructed in the form of nodes and edges. Then in step S4013, a specially designed semantic comparison algorithm is used to compare the semantic structure diagram of the new data with the historical templates one by one. This algorithm will carefully check the node type, such as different types of income or expenditure categories, node attributes such as amount, time, etc., and edge connection relationships, such as which account the funds flow from and to which account and whether the path of data flow is consistent.

[0076] In step S4014, the focus of the comparison work is to find out the consistency or difference between the new and old data in terms of node type, attributes, edge connection relationship and data flow path. Finally, step S4015 calculates the degree of match between the new and old data according to the pre-set matching rules and matching threshold. If the matching degree reaches or exceeds the set threshold, it is considered that the new data is highly matched with a certain historical data. At this time, combined with the company's reimbursement rules, the compliance and accuracy of the new data can be judged, thereby obtaining the final result of the reimbursement form verification. The degree of match is quantified by preset matching rules and thresholds, making the audit results more objective and repeatable. When the new data is highly matched with the historical data, the final judgment is made in combination with the reimbursement rules, which not only utilizes historical experience, but also ensures compliance with current policy requirements, ensures that the audit results of the reimbursement form are reliable, and reduces the financial risks of the enterprise.

[0077] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0078] As Figure 4 shown, the present application also provides an electronic device, including a display module 103, a memory 102, a processor 101, and a computer program stored on the memory and executable on the processor 101. When the processor 101 executes the program, the steps of the reimbursement data analysis and verification method are implemented.

[0079] In the embodiments of the present invention, the electronic device includes, but is not limited to, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of the present application described herein and / or claimed.

[0080] In the embodiments of the present application, the processor 101 can be implemented by using at least one of an application specific integrated circuit, a programmable logic device, a field programmable gate array, a processor, a controller, a microcontroller, a microprocessor, and an electronic unit designed to execute the functions described herein. In some cases, such an implementation can be implemented in a controller. For a software implementation, an implementation of a process or function can be implemented with a separate software module that allows the execution of at least one function or operation. The software code can be implemented by a software application (or program) written in any suitable programming language. The software code can be stored in the memory and executed by the controller.

[0081] The display module 103 is used to display information input by the user or information provided to the user. The display module 103 may include a display panel, and the display panel can be configured in the form of a liquid crystal display, an organic light emitting diode, etc.

[0082] The memory 102 can be used to store software programs and various data. The memory 102 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid state storage devices.

[0083] The present application also provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the reimbursement data analysis and verification method are implemented.

[0084] The storage medium may adopt any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0085] In the storage medium, the readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries the readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0086] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for analyzing and verifying reimbursement data, characterized in that, The method includes: S101: Obtain the reimbursement associated fields of the new financial accounting data, and determine the association degree value between the reimbursement associated fields and the historical associated fields in the financial dataset; S102: Obtain the historical financial accounting data corresponding to the historical associated fields with the association degree value greater than the matching threshold, and use the obtained historical financial accounting data as the comparison financial accounting data; S103: Determine the reimbursement form verification result of the new financial accounting data according to the comparison financial accounting data and the new financial accounting data.

2. The reimbursement data analysis and verification method according to claim 1, characterized in that The specific determination of the association degree value between the reimbursement associated fields and the historical associated fields in step S101 includes: Obtain the reimbursement associated fields of the new financial accounting data, and analyze their matching degree with the historical associated fields; Count all the matching degree results, determine the highest matching degree and the lowest matching degree as the basis for calculating the association degree value; Combine the highest matching degree, the lowest matching degree and the matching degree of each historical field to generate the corresponding association degree value; Sort according to the association degree value, and select the historical field data higher than the threshold as the comparison data for verifying the reimbursement form of the new financial accounting data.

3. The reimbursement data analysis and verification method according to claim 2, wherein The method also obtains the difference between the highest matching degree and the lowest matching degree as the global range parameter; For each historical associated field, calculate the difference between its matching degree with the reimbursement associated field and the lowest matching degree as the local offset; Divide the local offset by the global range parameter to obtain the relative ratio as the association degree value; Filter the comparison financial accounting data according to the association degree value.

4. The reimbursement data analysis and verification method according to claim 1, wherein In step S101, obtaining the reimbursement associated fields of the new financial accounting data specifically includes: Clean the invalid records of the new financial accounting data to obtain the preliminarily processed data; Verify and correct the data format according to the data format specification matching the reimbursement type; Filter out the non-business content, retain the core business data, and generate the preprocessed reimbursement form; Encode the preprocessed data through a structured information parsing tool to generate the reimbursement associated fields with a fixed length and clear semantics.

5. The reimbursement data analysis and verification method according to claim 1, wherein The method also includes: Construct multiple classified storage units through the bill coding location, and each classified storage unit only contains the historical financial accounting data corresponding to a specific bill coding; Obtain the new bill coding of the new financial accounting data; Traverse all the classified storage units, match the new bill coding with the preset historical bill coding in the classified storage units, and screen out the target storage unit; Extract the historical associated fields and their corresponding historical financial accounting data from the target storage unit; Calculate the association degree value based on the semantic and syntactic analysis results of the reimbursement associated fields and the historical associated fields to complete the reimbursement form verification.

6. The reimbursement data analysis and verification method according to claim 1, wherein Step S103 also includes: Obtain the financial semantic structure diagram of the new financial accounting data and the historical financial semantic template of the comparison financial accounting data; Determine the verification result of the expense report for the new financial accounting data according to the financial semantic structure diagram and the historical financial semantic template.

7. The expense data analysis and verification method according to claim 1, characterized in that Step S103 further includes: Classify the new financial accounting data into a preset business scenario classification system by analyzing the prefix of the bill code, keyword matching, and form title of the new financial accounting data; For the classified new financial data, identify the core business entities and their associated relationships therein; Map the extracted entities to a predefined financial semantic meta-model, unify the entity expression form, standardize the attribute value format, and construct a standardized financial semantic structure diagram; Based on the business scenario classification of the new financial data, retrieve the set of historical templates of the same type from the historical financial semantic template library to form a comparison sample pool; Compare the topological structures of the newly constructed financial semantic structure diagram and the historical template, count the matching ratio of the same entity types and relationship paths, and generate a structural similarity score; For the matching entity nodes, verify the consistency of their attribute values, and mark the abnormal levels for the nodes with differences; Generate a multi-dimensional comparison report by combining the structural similarity score and the attribute consistency verification result. When the similarity exceeds the preset threshold and there are no abnormalities in the key attributes, it is determined that the expense report passes the verification.

8. The expense data analysis and verification method according to claim 1, characterized in that Step S102 also extracts the financial semantic structure diagram of the new financial accounting data; the financial semantic structure diagram shows the relationships between financial data in the form of nodes and edges, where the nodes represent different financial entities and the edges represent the associations between the entities; Retrieve the historical financial semantic template; the historical financial semantic template stores the structured information of past financial data in a standardized format, also presented in the form of nodes and edges, and has been verified to comply with the established financial rules and reimbursement policies; Compare the financial semantic structure diagram of the new financial accounting data with the historical financial semantic template item by item through a semantic comparison algorithm; During the comparison process, pay attention to the consistency and differences in node types, node attributes, edge connection relationships, and data flow paths; Calculate the matching degree between the two according to the preset matching rules and matching degree threshold; When the matching degree reaches a certain threshold, it is considered that the new data highly matches a certain historical data. Based on this, combined with the reimbursement rules, judge the compliance and accuracy of the new data to obtain the verification result of the expense report.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the expense data analysis and verification method according to any one of claims 1 to 8.

10. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the expense data analysis and verification method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Financial reimbursement method, system and device and storage medium

    CN110288451A

  • Multifunctional enterprise financial accounting system

    CN114140092A

  • Enterprise financial reimbursement management system and method

    CN118446822A

  • Data warehouse construction method and system, electronic equipment and storage medium

    CN118568183A

  • Financial account closing method and device based on big data, equipment and medium

    CN119624676A

Cited By

  • Financial document processing method and device and electronic equipment

    CN120564213A