Reimbursement data management system and method based on AI artificial intelligence
By introducing AI technology into the reimbursement data management system, including multimodal data collection, intelligent auditing and compliance judgment, optimized approval process, data analysis and prediction, and blockchain storage, the problems of inefficiency and difficulty in ensuring data accuracy and reliability of the traditional reimbursement data management model are solved, and efficient and intelligent reimbursement data management is achieved.
Patent Information
- Application Number
- CN202411971345.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-27
AI Technical Summary
The traditional reimbursement data management model has shortcomings inefficiency, difficulty in ensuring data accuracy and reliability, in-depth analysis of the information behind the data, and difficulty in adapting to diversified reimbursement scenarios and dynamic management policies.
The reimbursement data management system is developed using AI-based artificial intelligence technology, including multimodal data acquisition and fusion module, audit and compliance determination module, approval process optimization module, analysis and prediction module, and data storage and blockchain module. Through deep learning, natural language understanding, reinforcement learning, machine learning and blockchain technology and other means, intelligent reimbursement data management is realized.
It significantly improves the efficiency, quality and safety of reimbursement data management, can deeply analyze data, identify potential laws and abnormal consumption patterns, optimize cost control strategies, enhance data accuracy and reliability, and adapt to diverse reimbursement scenarios and dynamic management policies.
Smart Images

Figure CN120047255A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of reimbursement data management, and specifically provides a reimbursement data management system and method based on AI artificial intelligence. Background Art
[0002] In the complex operation environment of modern enterprises, the reimbursement business, as a key part of financial management, the traditional management mode has been difficult to meet the increasing data processing requirements and efficient management requirements.
[0003] The traditional reimbursement process highly relies on manual operations. Employees need to spend a lot of time and energy sorting out paper bills and filling out cumbersome reimbursement forms. Financial personnel have to check the authenticity and compliance of a large amount of reimbursement information and bills one by one. This not only leads to a long reimbursement cycle and low efficiency, but also manual operations are extremely prone to problems such as data entry errors and audit deviations due to factors such as fatigue and negligence, seriously affecting the accuracy and reliability of financial data, and may further lead to adverse consequences such as distorted financial statements and incorrect decisions.
[0004] Traditional reimbursement data management has serious limitations in data analysis. It can only perform simple classification and summary statistics, and cannot deeply explore the rich information hidden behind the data, such as the potential laws of expense expenditures, the intelligent identification of abnormal consumption patterns, the precise optimization strategies for cost control, and the internal correlation relationships with other business data of the enterprise. It is difficult to meet the urgent needs of enterprise senior management for comprehensive, in-depth, and forward-looking data insights, greatly restricting the enterprise's adaptability and innovation and development potential in the market competition.
[0005] In addition, with the continuous expansion of enterprise scale, the continuous improvement of business complexity, and the increasingly changeable market environment, the traditional reimbursement management system is unable to cope with diverse reimbursement scenarios, flexibly adapt to the dynamic adjustment of enterprise internal management policies, and effectively prevent various fraudulent reimbursement behaviors. There is an urgent need for an innovative and more intelligent reimbursement data management solution to comprehensively improve the efficiency, quality, and security of reimbursement data management, and promote the development of enterprise financial management towards intelligence and refinement. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a reimbursement data management system and method based on AI artificial intelligence to solve the deficiencies of the prior art described in the background art, comprehensively improve the efficiency, quality, and security of reimbursement data management, and promote the development of enterprise financial management towards intelligence and refinement.
[0007] To solve the above technical problems, the embodiments of the present invention provide the following technical solutions:
[0008] On the one hand, an embodiment of the present invention provides a reimbursement data management system based on AI artificial intelligence, including:
[0009] A multimodal data acquisition and fusion module for providing a reimbursement application entry for access. The reimbursement application entry is used to select manual input of reimbursement information, upload invoice pictures and other supporting documents, and voice input of reimbursement reasons.
[0010] An audit and compliance determination module for constructing an invoice authenticity identification model based on deep learning algorithms, interacting with the tax department database and third-party verification platforms to verify invoice authenticity; applying natural language understanding technology and machine learning algorithms to construct a reimbursement reason analysis model; adopting a reimbursement item classification and audit model combining rule engines and machine learning to accurately classify and automatically audit reimbursement items.
[0011] An approval process optimization module for using reinforcement learning algorithms to determine the optimal approval path and approval node sequence based on the enterprise organizational structure, business processes, risk control strategies, and historical approval data.
[0012] An analysis and prediction module for performing multi-dimensional analysis of reimbursement data based on a big data processing framework and data mining algorithms; applying association rule mining, clustering analysis, and classification algorithms to mine potential association relationships between reimbursement data and other business data; based on a machine learning prediction model, using historical data to predict future reimbursement needs, expense trends, and cost changes, and generating visual reports and results.
[0013] A data storage and blockchain module for constructing a distributed and immutable data storage ledger using blockchain technology, encrypting and storing reimbursement data and related audit and approval information, and using blockchain smart contracts to execute reimbursement process automation rules and simultaneously execute data verification logic.
[0014] Optionally, the audit and compliance determination module includes an invoice authenticity identification sub-module, a reimbursement reason analysis sub-module, and a reimbursement item classification sub-module, where:
[0015] The invoice authenticity identification sub-module is used for deep learning and training of a large amount of invoice sample data. The invoice sample data includes the printing format, anti-counterfeiting marks, and code rules of invoices, and the verification data output by deep learning algorithms is interactively verified with the authoritative invoice database of the tax department and a third-party invoice verification platform.
[0016] The reimbursement reason analysis sub-module is used for constructing an intelligent reimbursement reason analysis model by applying natural language understanding technology and machine learning algorithms, performing semantic understanding and logical analysis on the reimbursement reason descriptions submitted by employees, and intelligently judging the reasonableness of reimbursement reasons in combination with the enterprise's internal detailed reimbursement policy knowledge base, industry standard specifications, and historical reimbursement data case library.
[0017] The reimbursement item classification sub-module adopts an intelligent classification and auditing model for reimbursement items that combines a rule engine and machine learning. According to the pre-set reimbursement item classification rules and the classification patterns automatically learned through machine learning, it accurately classifies reimbursement items according to the reimbursement item classification function, and for different types of reimbursement items, it conducts automated audits based on the corresponding audit criteria and threshold ranges.
[0018] Optionally, the approval process optimization module is used to utilize the reinforcement learning algorithm, based on the enterprise organizational structure, business processes, risk control strategies, and historical approval data. Specifically:
[0019] Let the state space be S, which includes the status information such as the amount, expense type, applicant's department, and historical reimbursement credit record of the reimbursement application; the action space be A, including different operations in the approval process; the reward function be R(s,a), which determines the reward value according to the approval result and business objectives; the state transition probability be P(s t+1 |s t ,a t ), which represents the probability of transferring to state s t after taking action a t in state s t+1 . The value function is V(s), and the update formula of the value function is based on the Bellman equation:
[0020]
[0021] where η is the learning rate and γ is the discount factor.
[0022] Optionally, the analysis and prediction module uses association rule mining, clustering analysis mining, and classification algorithms to mine the potential association relationships between reimbursement data and other business data. Specifically:
[0023] (1) The specific method of mining the potential association relationships between reimbursement data and other business data through association rule mining is as follows:
[0024] Integrate reimbursement data with other relevant business data, where the other relevant business data includes sales data, procurement data, inventory data, and human resources data;
[0025] Clean the data, handle missing values, outliers, and duplicate data, and perform encoding and conversion to convert categorical variables into numerical types;
[0026] Select the Apriori algorithm or the FP-Growth algorithm for association rule mining, find the frequent item sets and association rules between reimbursement data and other business data, and analyze the mined association rules to understand the business meaning of the association rules;
[0027] (2) The specific method for mining potential correlation relationships between reimbursement data and other business data through cluster analysis is as follows:
[0028] Determine the features for cluster analysis, including key indicators in reimbursement data and relevant features in other business data, and perform standardization processing on the selected features to make different features have the same dimension and data distribution;
[0029] Adopt clustering algorithms such as K-Means clustering algorithm, DBSCAN algorithm or hierarchical clustering algorithm for cluster analysis, calculate cluster evaluation indicators, evaluate the clustering effect, and adjust the clustering algorithm parameters or select a more suitable clustering algorithm according to the evaluation results;
[0030] Analyze the clustering results, assign business meanings to each cluster, and provide decision-making suggestions for the enterprise according to the clustering results;
[0031] (3) The specific method for mining potential correlation relationships between reimbursement data and other business data through classification algorithms is as follows:
[0032] Determine the classification task objective, select and extract features related to the classification task, and adopt a feature selection algorithm to select features with higher importance for classification model training;
[0033] Select a classification algorithm, use the training data set to train the selected classification algorithm, and adjust the algorithm parameters to optimize the model performance;
[0034] Use the test data set to evaluate the trained classification model, calculate the evaluation indicators, and apply the trained classification model with good evaluation indicators to the actual reimbursement data management to classify and predict new reimbursement applications or data.
[0035] Optionally, in the reimbursement reason analysis sub-module, set the reimbursement reason rationality judgment function as:
[0036] G(R) = α × M p (R) + β × M s (R) + γ × M c (R)
[0037] Where M p (R) represents the matching degree evaluation function of R based on the reimbursement policy knowledge base, M s (R) represents the evaluation function of R based on industry standard specifications, M c (R) represents the similarity evaluation function of R based on the historical reimbursement data case base, α, β, and γ are the corresponding weight coefficients, and α + β + γ = 1. When G(R) ≥ T G , the reimbursement reason is reasonable; otherwise, it is unreasonable; where T G is the preset rationality threshold, and R is the reason described by the employee.
[0038] Optionally, in the reimbursement item classification sub-module, the reimbursement item classification function is specifically as follows:
[0039]
[0040] Among them, Ln represents the nth category of reimbursement items, and K n (X) represents the classification discrimination function of the nth category of reimbursement items, and T n is the corresponding classification threshold.
[0041] Optionally, the data storage and blockchain module uses blockchain technology to build a distributed and tamper-proof data storage ledger, specifically as follows:
[0042] Select the blockchain architecture according to factors such as enterprise scale, data privacy requirements, and performance requirements
[0043] Set full nodes and light nodes. The full nodes store the complete ledger data and participate in transaction verification and block packaging, establish a secure P2P network connection, and set the node communication protocol;
[0044] Use asymmetric encryption algorithms to encrypt reimbursement data and audit and approval information, and select secure hash algorithms to perform hash processing on the data;
[0045] Build smart contracts to define the submission rules for reimbursement applications, embed data verification logic in the smart contracts, and deploy the written smart contracts to the nodes of the blockchain network;
[0046] Define the block structure. Each block contains a block header and a block body. Adopt a hierarchical storage strategy, store the data frequently accessed recently in a high-speed storage medium, and store historical data in a large-capacity and low-cost storage.
[0047] Optionally, the data storage and blockchain module uses blockchain smart contracts to execute the automation rules for the reimbursement process, specifically as follows:
[0048] Divide the smart contract into multiple functional modules, including a reimbursement application module, an approval process module, a data verification module, and a notification module;
[0049] Design the state transition logic of the smart contract based on the state machine model, and the state transition logic conforms to the predefined business rules and processes;
[0050] When an employee submits a reimbursement application, the smart contract automatically checks whether the data format meets the requirements, and verifies whether the reimbursement application complies with the relevant business rules according to the enterprise reimbursement policy;
[0051] The smart contract automatically determines the approval level and approval order based on factors such as the amount of the reimbursement application, the type of expense, and the department where the employee is located;
[0052] The smart contract automatically obtains the information of the corresponding approver from the enterprise organizational structure database or the predefined list of approvers, and notifies the approver to process the reimbursement application through the system.
[0053] Optionally, the data storage and data verification logic of the blockchain module specifically include:
[0054] Reimbursement amount is reconciled with the budget. By interacting with the enterprise budget management system or database through the smart contract, the budget quota information of the department, project, or individual where the employee is located is obtained. When the reimbursement application is submitted, the contract calculates the total amount reimbursed by the employee or project during the current period and compares it with the remaining budget;
[0055] Compliance check of reimbursement items. According to the reimbursement item classification standards preset by the enterprise, the smart contract checks whether the reimbursement items submitted by the employee are correctly classified, and verifies whether the reimbursement items comply with the enterprise's compliance policies and relevant laws and regulations.
[0056] On the other hand, the embodiment of the present invention also provides a working method of the reimbursement data management system based on AI artificial intelligence as described above, including:
[0057] Provide an accessible reimbursement application entry, which is used to select manual input of reimbursement information, upload invoice pictures and other supporting documents, and voice input of reimbursement reasons;
[0058] Build an invoice authenticity recognition model based on deep learning algorithms, interact with the tax department database and third-party verification platforms to verify the authenticity of invoices; use natural language understanding technology and machine learning algorithms to build a reimbursement reason analysis model; adopt a reimbursement item classification and review model based on the combination of rule engine and machine learning to accurately classify and automatically review reimbursement items;
[0059] Use reinforcement learning algorithms to determine the optimal approval path and the order of approval nodes based on the enterprise organizational structure, business processes, risk control strategies, and historical approval data;
[0060] Based on the big data processing framework and data mining algorithms, conduct multi-dimensional analysis of reimbursement data; use association rule mining, clustering analysis, and classification algorithms to mine the potential association relationships between reimbursement data and other business data; based on the machine learning prediction model, use historical data to predict future reimbursement needs, expense trends, and cost changes, and generate visual reports and results;
[0061] Build a distributed and tamper - proof data storage ledger using blockchain technology, encrypt and store reimbursement data and related review and approval information, and use blockchain smart contracts to execute the automation rules of the reimbursement process while implementing data verification logic.
[0062] The beneficial effects of the above - mentioned technical solutions of the present invention are as follows:
[0063] 1. The multi - modal data acquisition and fusion module of the present invention supports multiple terminals and input methods, docks with the enterprise internal system to obtain multi - source data and automatically fuses and verifies it. It can also provide real - time feedback on the reimbursement status, facilitating employee operation and ensuring data accuracy. The approval process optimization module uses reinforcement learning algorithms to automatically adjust the approval level according to the characteristics of reimbursement applications, simplifies the small - amount regular reimbursement process, adds approval links for major - risk reimbursements, and predicts approval opinions, significantly improving the reimbursement efficiency.
[0064] 2. The invoice authenticity identification sub - module of the present invention ensures the authenticity and legality of invoices through deep learning and multi - platform interactive verification. The reimbursement reason analysis sub - module uses natural language understanding and machine learning technologies, combines multi - aspect knowledge bases and case bases, and accurately judges the reasonableness of reimbursement reasons. The reimbursement item classification sub - module is based on a rule engine and a machine learning model, accurately classifies reimbursement items and automatically audits them according to standards.
[0065] 3. The analysis and prediction module of the present invention uses a variety of data mining algorithms to explore the correlation between reimbursement data and other business data, provides cross - departmental business insights, and helps optimize resource allocation and business processes. Based on a machine learning prediction model, it uses historical data to predict future reimbursement needs, etc., providing a scientific basis and forward - looking guidance for enterprise financial budgets, etc.
[0066] 4. The data storage and blockchain module of the present invention builds a distributed and tamper - proof storage ledger using blockchain technology, encrypts and stores data, and ensures data confidentiality, integrity, and tamper - proofness through reasonable architecture selection, node setting, encryption algorithms, and smart contracts. It uses blockchain smart contracts to execute the automation rules of the reimbursement process, strictly implements data verification logic, including verifications such as invoice authenticity, amount compliance, and item classification accuracy, and manages user permissions at the same time to ensure the compliance and transparency of the process. Brief Description of the Drawings
[0067] Figure 1 It is the process and principle block diagram of the reimbursement data management system based on AI artificial intelligence of the present invention;
[0068] Figure 2 It is the architecture diagram of the reimbursement data management system based on AI artificial intelligence of the present invention;
[0069] Figure 3This is the principle block diagram of the audit and compliance determination module of the reimbursement data management system based on AI artificial intelligence of the present invention. Detailed implementation manners
[0070] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.
[0071] As Figure 1 、 Figure 2 shown, a reimbursement data management system based on AI artificial intelligence, the specific working method steps of the reimbursement data management system of the present invention are as follows:
[0072] The multimodal data acquisition and fusion module 101 is used to provide an access reimbursement application entry, and the reimbursement application entry is used to select manual input of reimbursement information, upload invoice pictures and other supporting documents, and voice input of reimbursement reasons. It fully supports access by various terminal devices (including but not limited to computers, smart phones, tablets, etc.). Employees can flexibly select various ways to input reimbursement information through this module, such as manually inputting detailed information such as reimbursement reasons, amounts, dates, uploading invoice pictures, itinerary receipts, and other relevant supporting documents. At the same time, they can also use the advanced voice input function to describe the reimbursement reasons. The system automatically converts the voice into structured text data by means of voice recognition and natural language processing technologies. It can achieve deep seamless docking with various internal business systems of the enterprise (such as travel reservation systems, office supplies procurement systems, enterprise resource planning systems, human resources systems, financial systems, etc.), automatically obtain structured data related to reimbursement, such as itinerary information, order details, employee department and rank information, budget information, etc., and perform real-time intelligent fusion and preliminary verification on these multi-source data and the data manually input or converted by voice by employees to ensure the integrity, consistency and accuracy of the data. It has a powerful information feedback mechanism, and shows the accurate status of the reimbursement application, such as to be submitted, under review, approved, rejected, modification required due to non-approval, etc., as well as detailed processing results and targeted prompt information to employees and approvers in real time, ensuring that users can always understand the progress of the reimbursement.
[0073] As Figure 3 shown, the audit and compliance determination module 102 includes an invoice authenticity identification sub-module 1021, a reimbursement reason analysis sub-module 1022, and a reimbursement item classification sub-module 1023, wherein:
[0074] The invoice authenticity identification sub-module 1021 is used for deep learning and training of a large amount of invoice sample data. The invoice sample data includes the printing format, anti-counterfeiting marks, and code rules of the invoice. The verification data output by the deep learning algorithm is interactively verified with the authoritative invoice database of the tax department and the third-party invoice verification platform to ensure the authenticity and legality of each reimbursement invoice. When verifying the authenticity of invoices through the deep learning algorithm, first collect a large amount of invoice sample data, including invoices of different types (such as special VAT invoices, ordinary invoices, etc.) and different sources (issued in different regions and industries), to ensure the diversity and representativeness of the samples. For the printing format of the invoice, extract feature information such as the overall layout, text typesetting, font style, and color distribution of the invoice, and convert it into structured data that can be processed by the deep learning algorithm. For example, divide and label the text areas, table areas, etc. on the invoice, and record their positions, sizes, contents, etc. For the anti-counterfeiting marks, use image recognition technology to extract features such as the patterns, textures, and color changes of the anti-counterfeiting marks, such as the watermarks, fluorescent fibers, security lines, etc. on the invoice, and quantify and encode these features. Analyze the code rules of the invoice, extract key code information such as the invoice code, number, issuing date, verification code, etc., and perform normalization processing to ensure the uniformity of the data format.
[0075] Then, construct and train the deep learning model. For example, select the convolutional neural network (CNN) as the basic architecture to construct the invoice authenticity identification model. A neural network with a multi-branch structure can also be designed. One branch is specifically used to process the feature data related to the printing format, one branch processes the anti-counterfeiting mark feature data, and another branch processes the code rule data. Finally, the results of the three branches are fused to comprehensively judge the authenticity of the invoice.
[0076] By dividing the invoice sample data into a training set, a validation set, and a test set, use the training set to train the constructed deep learning model. After training, verify the authenticity through the test set or the invoice to be verified. The verification data output by the deep learning model includes the above-mentioned verification data related to the printing format, anti-counterfeiting marks, and code rules. For example, through the learning of the overall layout of the invoice sample data by the deep learning model, output the matching degree between the invoice to be verified and the standard invoice layout template. For example, calculate the difference values of the relative positions and size ratios of each element on the invoice (such as the invoice title, purchaser information, seller information, invoice content area, total amount of price and tax area, etc.) and the standard template, and express the matching degree as a percentage. A matching degree of more than 95% can initially judge that the layout conforms to the standard format.
[0077] The reimbursement reason analysis sub-module 1022 is used to build an intelligent reimbursement reason analysis model by applying natural language understanding technology and machine learning algorithms, perform semantic understanding and logical analysis on the reimbursement reason descriptions submitted by employees, and combine the detailed reimbursement policy knowledge base within the enterprise, industry standard specifications, and historical reimbursement data case base to intelligently judge the reasonableness of the reimbursement reason. In the reimbursement reason analysis sub-module, the reimbursement reason reasonableness judgment function is set as:
[0078] G(R) = α × M p (R) + β × M s (R) + γ × M c (R)
[0079] Among them, M p (R) represents the matching degree evaluation function of R based on the reimbursement policy knowledge base, M s (R) represents the evaluation function of R based on industry standard specifications, M c (R) represents the similarity evaluation function of R based on the historical reimbursement data case base. α, β, and γ are the corresponding weight coefficients, and α + β + γ = 1. When G(R) ≥ T G , the reimbursement reason is reasonable; otherwise, it is unreasonable. Among them, T G is the preset reasonableness threshold, and R is the reason described by the employee.
[0080] (1) The matching degree evaluation function of R based on the reimbursement policy knowledge base includes:
[0081] Keyword matching and semantic analysis: Perform word segmentation on the reason R described by the employee, extract the key nouns, verbs, adjectives, etc. in it, such as keywords like "travel expenses", "conference expenses", "training", "purchase", "urgent", etc. Then search for these keywords in the reimbursement policy knowledge base and count the number and frequency of keyword matches. For example, if the reimbursement policy stipulates the reimbursement scope and standards for specific items, when the employee's description contains keywords related to these regulations, increase the matching degree score. Use semantic understanding technology to analyze the semantic similarity between the reason R described by the employee and the terms in the reimbursement policy knowledge base. Convert the employee's description into a semantic vector representation and calculate the cosine similarity with the semantic vectors of the terms in the knowledge base. For example, for travel expense reimbursement, if the employee's description is "transportation and accommodation expenses incurred due to business expansion needs to visit customers outside the city", and the semantic similarity with the travel expense in the reimbursement policy "necessary transportation and accommodation expenses incurred due to company business activities going out" is relatively high, then increase the matching degree.
[0082] Rule matching and logical judgment: According to the rules in the reimbursement policy knowledge base, determine whether the reasons described by the employee meet the basic conditions and procedures for reimbursement. For example, for some expense reimbursements, prior applications or specific approval procedures may be required. If the employee's description does not mention or does not comply with these rules, the matching score is reduced. Check whether the expense details in the reasons described by the employee comply with the expense types and standards specified in the reimbursement policy. Such as the transportation expense standards in business trip expenses (such as restrictions on the class of air tickets, regulations on the seat types of train tickets, etc.), accommodation expense standards (the upper limit of accommodation expenses divided by region and level), etc. If it exceeds the standard range, the matching degree is correspondingly reduced.
[0083] (2) Evaluation function for R based on industry standard specifications, including:
[0084] Compliance check: According to the general industry expense reimbursement standards and specifications, check whether the expense items in the reasons R described by the employee are compliant. For example, in some industries, there may be strict per capita consumption standards and entertainment scope restrictions for business entertainment expenses. If the business entertainment expenses described by the employee exceed the industry standards (such as too high per capita consumption, the entertainment objects do not meet the business-related scope), the evaluation score is reduced. For professional service fees (such as legal consultation fees, audit fees, etc.), check whether the service provider has the corresponding qualifications and industry recognition. If the qualifications of the service provider do not meet the industry specifications, record the violation situation and reduce the evaluation value.
[0085] Rationality judgment: Analyze whether the expense expenditures in the reasons described by the employee are reasonable from the perspective of industry practices. For example, for office supplies procurement expenses, compare with the industry average procurement price level. If the office supplies price reimbursed by the employee is significantly higher than the normal market price and there is no reasonable reason (such as special customization requirements, etc.), it is considered unreasonable and the evaluation score is reduced. Consider the expense expenditure range of similar business activities in the industry. If the expenses described by the employee are significantly different from the industry normal level (such as too high or too low), further verify its rationality. For example, the average expense of similar training activities in an industry is 200 yuan per person per day, while the training expense reimbursed by the employee is 300 yuan per person per day and no sufficient reasonable explanation is provided (such as inviting well-known experts to give lectures but no relevant certificates are provided), then it is evaluated as unreasonable.
[0086] (3) Similarity evaluation function for R based on the historical reimbursement data case base
[0087] Text similarity calculation: Use text similarity algorithms (such as edit distance algorithm, cosine similarity algorithm, etc.) to calculate the similarity between the employee's described reason R and the cases in the historical reimbursement data case library. Convert the employee's description and historical cases into text vector representations, and measure the similarity between the two by calculating the distance or similarity between the vectors. For example, the edit distance algorithm calculates the minimum number of edit operations (inserting, deleting, replacing characters, etc.) required to convert one text into another. The smaller the edit distance, the higher the similarity.
[0088] Feature matching and weight assignment: Extract key features from the employee's described reason R and historical reimbursement cases, such as expense type, business activity content, occurrence location, time, etc. Assign different weights to different features according to their impact on the rationality of reimbursement. For example, the weight of the expense type is relatively high, and the weight of the business activity content is the second. Calculate the feature matching degree. For the matching features, accumulate scores according to their weights. If the employee's described travel expense reimbursement has a high matching degree with the travel expense reimbursement in the historical case in terms of key features such as expense type, destination, and business trip time, the similarity score is increased; on the contrary, if there are many differences (such as different expense types, destinations that do not conform to business logic, etc.), the similarity evaluation value is decreased.
[0089] Anomaly detection and adjustment: If the similarity between the employee's described reason R and all cases in the historical reimbursement data case library is relatively low (lower than the set threshold, such as 30%), the anomaly detection mechanism is triggered. Further analyze the details in the employee's description to check for abnormal situations (such as excessive expenses without reasonable explanations, uncommon business activities, etc.). If anomalies are found, reduce the similarity evaluation score and may mark it as a high-risk reimbursement item that requires manual review. Regularly update the historical reimbursement data case library according to the enterprise's business development and market changes to ensure that the similarity evaluation function can adapt to new business models and expense expenditure situations. For example, as the enterprise's business expands into new fields, promptly incorporate reimbursement cases related to the new fields to improve the accuracy of similarity evaluation.
[0090] The reimbursement item classification sub-module 1023 adopts an intelligent classification and review model for reimbursement items that combines a rule engine and machine learning. According to the pre-set reimbursement item classification rules and the classification patterns automatically learned through machine learning, accurately classify the reimbursement items according to the reimbursement item classification function, and conduct automated reviews for different types of reimbursement items according to the corresponding review criteria and threshold ranges. The reimbursement item classification function is specifically:
[0091]
[0092] Among them, Ln represents the nth type of reimbursement item, and K n (X) represents the classification discriminant function of the nth type of reimbursement item, and Tn is the corresponding classification threshold.
[0093] The approval process optimization module 103 is used to determine the optimal approval path and the order of approval nodes based on the enterprise organizational structure, business processes, risk control strategies, and historical approval data by using the reinforcement learning algorithm, automatically simplify or increase the approval levels according to the characteristics of the reimbursement application, and perform automated transfer of the approval process, recording of approval information, and overdue reminder. The reinforcement learning algorithm is specifically as follows:
[0094] Let the state space be S, which includes the state information such as the amount size, expense type, department where the applicant is located, and historical reimbursement credit record of the reimbursement application; the action space be A, including different operations in the approval process, such as approval, rejection, transfer to the next-level approval, etc.; the reward function be R(s,a), which determines the reward value according to the approval result and business objectives; the state transition probability be P(s t+1 |s t ,a t ), which represents the probability of transferring to state s t after taking action a t in state s t+1 . The value function is V(s), and the update formula of the value function is based on the Bellman equation:
[0095]
[0096] where η is the learning rate and γ is the discount factor.
[0097] For small and regular reimbursement applications, the system can automatically simplify the approval process, directly perform quick approval by the direct supervisor, and intelligently predict the supervisor's approval opinion based on the supervisor's approval historical data and behavior patterns. If the prediction is passed and the accuracy rate reaches a certain threshold, some approval links can be automatically skipped and directly enter the financial review or reimbursement payment stage, significantly improving the reimbursement efficiency; while for reimbursement applications involving major project investments, high-cost expenditures, or high risks, the approval levels and professional review links are automatically increased, such as cost-benefit analysis by financial experts, compliance review by the legal department, and risk assessment by the audit department, to ensure that the approval process is rigorous and risk controllable.
[0098] The analysis and prediction module 104 performs multi-dimensional analysis on the reimbursement data based on the big data processing framework and data mining algorithms, and uses association rule mining, clustering analysis, and classification algorithms to mine the potential association relationships between the reimbursement data and other business data. Specifically:
[0099] I. Association rule mining
[0100] First, perform data preparation and preprocessing. Integrate reimbursement data with other relevant business data, such as sales data (sales amount, number of sales orders, etc.), procurement data (procurement amount, procurement item categories, etc.), inventory data (inventory turnover rate, inventory level, etc.), human resources data (number of employees, employee performance, etc.), etc. Ensure that the data is in the same data warehouse or data lake, and perform data cleaning to handle missing values, outliers, and duplicate data.
[0101] Then, encode and transform the data. Convert categorical variables into numerical types to facilitate the processing of association rule mining algorithms. For example, perform one-hot encoding on reimbursement item categories (such as travel expenses, office supplies expenses, etc.).
[0102] Next, select the algorithm and set parameters. Choose the Apriori algorithm or the FP-Growth algorithm for association rule mining. For the Apriori algorithm, set appropriate minimum support (min_support) and minimum confidence (min_confidence) parameters. The minimum support determines the lowest frequency threshold for the occurrence of frequent item sets, and the minimum confidence measures the reliability of the association rules. For example, initially, min_support = 0.05 (indicating that the item set appears in the dataset at least 5%), and min_confidence = 0.6 (indicating that the confidence of the association rule is at least 60%), and then adjust and optimize according to the mining results. For the FP-Growth algorithm, mainly focus on its memory usage and performance optimization parameters to ensure that the algorithm can efficiently process large-scale data.
[0103] Finally, perform association rule mining and analysis. Run the association rule mining algorithm to find the frequent item sets and association rules between reimbursement data and other business data. For example, it may be found that there is an association between "higher travel expense reimbursement amount" and "frequent sales business expansion activities (increase in the number of sales orders)", with a support of 0.1 (i.e., these two situations occur simultaneously in 10% of the data) and a confidence of 0.7 (indicating that when the travel expense reimbursement amount is higher, there is a 70% probability of frequent sales business expansion activities).
[0104] Analyze the mined association rules and understand their business implications. In addition to the above association between sales and travel expenses, it may also be found that "increase in office supplies procurement amount" is related to "new employee recruitment (increase in the number of employees)", or "increase in equipment maintenance costs" is associated with "decrease in production output (possibly indicating equipment failures affecting production)", etc. These association rules can provide cross-departmental business insights for the enterprise and help optimize resource allocation and business processes.
[0105] Visualize the association rules, such as using association rule graphs or matrix graphs, to intuitively present the relationship between reimbursement data and other business data. For example, in an association rule graph, nodes represent reimbursement items or business metrics, edges represent association relationships, and the thickness or color of the edges can represent the association strength (support or confidence).
[0106] II. Cluster Analysis
[0107] First, determine the features for cluster analysis, including key metrics in reimbursement data (such as reimbursement amount, reimbursement frequency, reimbursement item category, etc.) and relevant features in other business data (such as sales performance, inventory turnover rate, employee performance score, etc.). Select representative and discriminative features to avoid the curse of dimensionality caused by too many features. Standardize the selected features so that different features have the same dimension and data distribution. For example, use the Z-Score standardization method to convert feature values into a standard normal distribution with a mean of 0 and a standard deviation of 1, ensuring that the clustering algorithm will not be biased due to differences in feature scales.
[0108] Second, use clustering algorithms such as the K-Means algorithm, DBSCAN algorithm, or hierarchical clustering algorithm for cluster analysis. For the K-Means algorithm, preliminarily determine the number of clusters K based on business understanding and data distribution. For example, if you want to divide enterprise departments or employees into different expense behavior pattern groups, you can first try different values such as K = 3 or K = 4, and then select the optimal K value through evaluation metrics.
[0109] Calculate clustering evaluation metrics, such as the Silhouette Coefficient, Calinski-Harabasz index, etc., to evaluate the clustering effect. The Silhouette Coefficient measures the average ratio of the distance of each data point to its own cluster center to the distance to other cluster centers, and its value ranges from -1 to 1. The closer it is to 1, the better the clustering effect. The Calinski-Harabasz index evaluates the clustering quality by calculating the ratio of the between-class scatter and the within-class scatter, and the larger the ratio, the better.
[0110] Adjust the clustering algorithm parameters or select a more suitable clustering algorithm according to the evaluation results to obtain the best clustering effect. For example, if it is found that the K-Means algorithm has an unsatisfactory clustering effect in some cases (such as having many outliers or irregular cluster shapes), the DBSCAN algorithm can be tried. This algorithm can automatically discover dense regions and noise points in the data, has few assumptions about the data distribution shape, and is more suitable for dealing with clusters of complex shapes.
[0111] Analyze the clustering results and assign business meanings to each cluster. For example, through clustering analysis, enterprise employees can be divided into several groups. One group may be "employees with high reimbursement and high performance", who have relatively high reimbursement amounts but at the same time bring high business value to the enterprise (such as outstanding sales performance); another group may be "employees with low reimbursement and low performance", and further attention needs to be paid to their work performance and resource utilization; it is also possible to discover an "abnormal reimbursement employee group", whose reimbursement behavior is significantly different from other groups, and there may be potential risks of irregular reimbursement, which requires key auditing.
[0112] Provide decision-making suggestions for the enterprise based on the clustering results. For example, for the group of "employees with high reimbursement and high performance", more flexible reimbursement policies or incentive measures can be considered to encourage them to continue creating value for the enterprise; for the "abnormal reimbursement employee group", strengthen internal auditing and expense control, standardize the reimbursement process, and prevent waste and loss of enterprise resources. At the same time, the clustering results can also help the enterprise discover the similarities and differences between different business departments or business processes, providing a reference basis for optimizing the organizational structure and reengineering business processes.
[0113] III. Application of Classification Algorithms
[0114] First, determine the classification task objectives, such as predicting whether there is a fraud risk in reimbursement applications (classified into fraud and normal categories) or judging whether the reimbursement expenses exceed the budget (classified into over-budget and not over-budget categories), etc. Label the data according to the objectives to construct a training data set and a test data set. The labeling process can combine the internal audit results of the enterprise, financial rules, and the experience of business experts, etc.
[0115] Conduct feature engineering, select and extract features related to the classification task. In addition to the features of the reimbursement data itself (such as reimbursement amount, item category, time, etc.), features of other business data can also be considered as aids, such as employee credit records (from human resources data), supplier reputation (from procurement data), market volatility (from external market data), etc. Screen and combine the features to remove redundant and irrelevant features, and improve the accuracy and efficiency of the classification model.
[0116] Feature selection algorithms can be used, such as the Chi-Square Test, Information Gain, etc. to evaluate the importance of features, and select features with higher importance for the training of the classification model. For example, through the Chi-Square Test, it is found that the two features of "the deviation degree of the reimbursement amount from the average reimbursement amount" and "whether the supplier is a new supplier" have relatively high importance for predicting the risk of reimbursement fraud.
[0117] Select appropriate classification algorithms, such as decision tree algorithms (C4.5, CART, etc.), support vector machines (SVM), naive Bayes classifiers, neural networks (such as multi-layer perceptrons MLP), etc. Different algorithms have their own advantages, disadvantages and applicable scenarios. For example, decision tree algorithms are easy to understand and interpret and are suitable for processing data with hierarchical structures; SVM performs well in dealing with small sample and high-dimensional data; neural networks have powerful non-linear modeling capabilities and are suitable for complex classification tasks, but have relatively high computational complexity and poor interpretability.
[0118] Use the training dataset to train the selected classification algorithm and adjust the algorithm parameters to optimize the model performance. For example, for decision tree algorithms, the depth of the tree and the splitting node selection criteria (such as information gain ratio, Gini index, etc.) can be adjusted; for SVM algorithms, the kernel function type (such as linear kernel, polynomial kernel, radial basis function kernel, etc.) and the penalty parameter C can be adjusted; for neural networks, parameters such as the number of hidden layers, learning rate, and number of iterations can be adjusted. Select the best parameter combination through methods such as cross-validation to improve the generalization ability of the classification model.
[0119] Use the test dataset to evaluate the trained classification model and calculate evaluation metrics such as accuracy, precision, recall, F1 value, etc. Accuracy represents the proportion of the number of samples predicted correctly by the model to the total number of samples; precision measures the proportion of the number of samples predicted as positive and actually positive to the number of samples predicted as positive by the model; recall represents the proportion of the number of samples actually positive and predicted as positive by the model to the number of actual positive samples; the F1 value is the harmonic mean of precision and recall, comprehensively considering the influence of both. Judge whether the model performance meets the business requirements according to the evaluation metrics. If the performance is not ideal, further optimize the model or adjust the feature engineering.
[0120] Apply the trained and well-performing classification model to the actual reimbursement data management to classify and predict new reimbursement applications or data. For example, when a reimbursement application is submitted, the system automatically uses the classification model to predict whether there is a fraud risk or over-budget risk in the application and take corresponding measures according to the prediction results. If the prediction is a high risk, the system can automatically trigger a manual review process or remind relevant departments to pay key attention, thus effectively preventing enterprise financial risks and improving the intelligent level and decision-making support ability of reimbursement data management. At the same time, continuously collect new data and regularly update and optimize the classification model to adapt to the changes in the enterprise business environment and reimbursement behavior patterns.
[0121] Next, based on the machine learning prediction model, use historical data to predict future reimbursement needs, expense trends, and cost changes, and generate visual reports and results;
[0122] Based on machine learning prediction models, such as time series prediction models, regression analysis models, etc., using the time series characteristics of historical reimbursement data and the influencing factors of relevant business data, accurately predict future reimbursement needs, expense trends, cost changes, etc. Let the historical reimbursement data sequence be y1, y2, …, yt (t is the time point), and the prediction model be M F , and the prediction function be Y t+h (h is the prediction step), then:
[0123] Then Y t+h = M F (y1, y2, …, yt)
[0124] For a simple autoregressive integrated moving average (ARIMA) model, its form is
[0125]
[0126] where B is the lag operator, and θ(B) are polynomials, Δ d is the differencing operator, and ε t is a white noise sequence. Determine the model parameters by fitting the historical data, and then realize the prediction of future reimbursement data. The prediction results can provide a scientific basis and forward-looking guidance for enterprise financial budget preparation, cost control strategy formulation, resource planning and allocation, etc., helping the enterprise make preparations in advance and reduce business risks.
[0127] The data storage and blockchain module 105 is used to build a distributed and tamper-proof data storage ledger using blockchain technology, and encrypt and store reimbursement data and related audit and approval information. Among them, the data storage and blockchain module builds a distributed and tamper-proof data storage ledger using blockchain technology, specifically as follows:
[0128] I. Blockchain architecture selection and network construction
[0129] First, comprehensively consider factors such as enterprise scale, data privacy requirements, and performance requirements to select a blockchain architecture. For large enterprises with reimbursement data management involving multi-department collaboration, the consortium blockchain architecture is more suitable. It can ensure a certain degree of decentralization and meet the enterprise's requirements for data privacy and permission management. For example, the internal financial department, each business department, and the audit department of the enterprise can be used as nodes of the consortium blockchain to jointly maintain the security and stability of the ledger.
[0130] Then, compare different consortium blockchain platforms, such as Hyperledger Fabric, Corda, etc. Hyperledger Fabric provides a high degree of modularity and scalability, supports multiple consensus algorithms (such as PBFT, Raft, etc.), and can be flexibly configured according to the actual situation of the enterprise. Its channel concept allows data isolation and sharing between different business scenarios or departments, which is very suitable for the data interaction requirements of different processes and departments in reimbursement data management.
[0131] Next, determine the node types and distribution. Set up full nodes and light nodes. Full nodes store the complete ledger data and participate in transaction verification and block packaging, such as the nodes in the finance department; light nodes can only store some key data for query verification, such as the nodes in the departments where ordinary employees are located. Reasonably distribute the nodes to ensure the reliability of the network and redundant backup of data.
[0132] Finally, configure the network parameters. Establish a secure P2P network connection, set the communication protocol between nodes (such as the TLS protocol) to ensure the security of data transmission. At the same time, configure the node joining and leaving mechanisms, conduct identity authentication through the CA to issue digital certificates, and strictly manage node permissions to prevent illegal nodes from accessing.
[0133] II. Design and Implementation of Data Encryption Scheme
[0134] First, use an asymmetric encryption algorithm (such as ECC or RSA) to encrypt the reimbursement data and review and approval information. For the reimbursement data submitted by employees, use the public key of the recipient (such as the finance node) to encrypt it to ensure that only the finance department can decrypt and view it using the private key, guaranteeing the confidentiality of the data during transmission and storage.
[0135] For sensitive fields (such as employee identity information, bank account numbers, etc.), adopt additional encryption measures, such as the AES symmetric encryption algorithm combined with a key management system. Generate a random symmetric key to encrypt the sensitive fields, and then use the asymmetric encryption algorithm to encrypt the symmetric key and store it together with the encrypted data to ensure the high security of sensitive information.
[0136] Then, select a secure hash algorithm such as SHA-256 to perform hash processing on the data. Before storing the data, calculate the hash value of each data block (such as a reimbursement record) and store the hash value in the blockchain ledger. Any data modification will cause the hash value to change, and by comparing the hash values, it is possible to quickly detect whether the data has been tampered with.
[0137] Finally, construct a Merkle Tree structure. Use the hash values of multiple reimbursement data blocks as leaf nodes to construct the Merkle Tree, and store the hash value of the root node in the block header. By verifying the hash values on the Merkle Tree path, the integrity of the data blocks can be efficiently verified, while reducing the workload of data storage and verification.
[0138] III. Development and Deployment of Smart Contract Functions
[0139] First, automate the reimbursement process. Build a smart contract to define the submission rules for reimbursement applications, such as data format verification, mandatory field checks, etc.; automatically trigger the approval process, determine the approval level and approvers according to preset rules (such as amount size, expense type, etc.); implement the automatic recording and feedback of approval results to improve the efficiency and transparency of the reimbursement process.
[0140] Data verification and permission management. Embed data verification logic in the smart contract, such as invoice authenticity verification (by interacting with an external invoice verification platform), compliance check of reimbursement amount (comparing with budget data), verification of the accuracy of reimbursement item classification, etc. At the same time, manage user permissions to ensure that only authorized personnel can perform specific operations, such as submitting reimbursements, approving reimbursements, etc.
[0141] Write the smart contract code using a suitable programming language (such as Solidity or Chaincode language based on Hyperledger Fabric). Follow security coding specifications, conduct strict code reviews and tests to prevent data security issues caused by contract vulnerabilities.
[0142] Deploy the written smart contract to the nodes of the blockchain network. Before deployment, conduct sufficient tests, including unit tests, integration tests, and simulation environment tests, to ensure that the contract can execute correctly in various situations. After deployment, establish a contract upgrade mechanism so that the contract can be updated in a timely manner when business requirements change or vulnerabilities are discovered.
[0143] IV. Ledger Data Storage Structure and Index Optimization
[0144] Define the block structure. Each block contains a block header (including version number, timestamp, hash value of the previous block, hash value of the Merkle tree root, etc.) and a block body (including a series of reimbursement transaction data). The reimbursement transaction data details record the reimbursement order number, employee information, reimbursement item details, amount, invoice information, review comments, approval results, and time, etc., to ensure the integrity and traceability of the data.
[0145] Optimize the data storage method. Adopt a hierarchical storage strategy, store frequently accessed data recently in high-speed storage media (such as memory or SSD), and store historical data in large-capacity and low-cost storage (such as HDD). At the same time, perform partition storage according to data characteristics, such as partitioning by time or by business type, to improve data query and management efficiency.
[0146] Establish multi-dimensional indexes. Establish indexes (such as B-tree indexes or hash indexes) for common query conditions (such as reimbursement form numbers, employee numbers, reimbursement time ranges, approval statuses, etc.) to accelerate data query speed. At the same time, establish composite indexes, such as (employee number, reimbursement time) indexes, to meet the requirements of multi-condition queries.
[0147] Optimize the query algorithm. Adopt technologies such as paged query and caching query results to avoid performance degradation caused by loading a large amount of data at one time. For complex queries (such as those involving multi-table associations or statistical analysis), use the smart contract function of the blockchain for pre-computation and data aggregation, store the results in the contract, and directly obtain them during query to improve the response speed.
[0148] The data storage and blockchain module 105 uses the smart contract of the blockchain to execute the automation rules of the reimbursement process, specifically as follows:
[0149] I. Smart contract architecture design
[0150] First, divide the smart contract into multiple functional modules, such as the reimbursement application module, approval process module, data verification module, notification module, etc. Each module focuses on specific business logic to improve the maintainability and scalability of the contract. For example, the reimbursement application module is responsible for receiving and initially verifying the reimbursement application data submitted by employees; the approval process module determines the approval path and processes approval operations according to preset rules; the data verification module strictly verifies all aspects of the reimbursement data (such as invoice authenticity, amount compliance, etc.); the notification module is used to send reimbursement status updates and important event reminders to relevant personnel.
[0151] Second, communicate and interact with data between each module through standardized interfaces to ensure data consistency and smooth process. For example, after the reimbursement application module receives complete and initially verified application data, it passes the data to the approval process module through the interface; when the data verification module needs to verify data, it obtains relevant data from other modules and returns the verification results to the calling module.
[0152] Secondly, design the state transition logic of the smart contract based on the state machine model. The reimbursement process usually includes multiple states, such as "application submitted", "under approval", "approved", "rejected", "reimbursed", etc. The smart contract realizes legal transitions between states according to different events triggered (such as employees submitting reimbursement applications, approvers performing approval operations, etc.). For example, when an employee successfully submits a reimbursement application and the data verification passes, the contract state changes from "initial" to "application submitted"; after the approver approves the reimbursement application, the state changes to "approved", and the subsequent process continues based on this state.
[0153] The smart contract can ensure that each state transition complies with predefined business rules and process requirements, preventing illegal or incorrect state changes. For example, a reimbursement application can only enter the "reimbursed" state when all necessary review and approval links are completed and passed; if data problems or violations are found during the approval process, the contract state should change to "rejected" and record the reason for rejection.
[0154] II. Execution of Reimbursement Process Automation Rules
[0155] When an employee submits a reimbursement application, the smart contract automatically checks whether the data format meets the requirements, such as whether the date format is correct, whether the amount data is a legal value, whether the required fields are complete, etc. For example, it is required that the reimbursement date must be in the "YYYY - MM - DD" format, the reimbursement amount must be a positive number and retain two decimal places, and the required information such as the employee's name and department cannot be empty. If the data format is incorrect or incomplete, the contract rejects the application and prompts the employee to modify it.
[0156] Secondly, verify whether the reimbursement application complies with relevant business rules according to the enterprise reimbursement policy. For example, check whether the reimbursement item is within the scope allowed by the enterprise (such as some enterprises may limit the reimbursement of entertainment expenses); determine whether there is a duplicate application for reimbursement of the same item (by comparing with historical reimbursement data); verify whether the reimbursement amount exceeds the budget limit of the employee or project (interact with the budget management module or database to obtain budget information). If the business rules are violated, the contract automatically marks the application as abnormal and notifies the relevant personnel.
[0157] The smart contract automatically determines the approval level and approval order according to factors such as the amount of the reimbursement application, the type of expense, and the department where the employee is located. For example, it is set that small - amount reimbursements (such as amounts less than 1000 yuan) only require approval by the direct superior; medium - amount reimbursements (1000 yuan to 5000 yuan) need to be approved by the department manager and the financial supervisor; large - amount reimbursements (more than 5000 yuan) require additional approval by the company's senior leadership. At the same time, according to the type of expense (such as travel expenses, office supplies expenses, business entertainment expenses, etc.), different professional approvers or departments may be involved.
[0158] The contract automatically obtains the information of the corresponding approver from the enterprise organizational structure database or the predefined list of approvers, and reminds the approver to process the reimbursement application in a timely manner through system notifications (such as emails, enterprise internal messaging systems, etc.). After the approver logs in to the system, the smart contract displays the detailed information of the reimbursement application and auxiliary review opinions (such as data verification results, reference of historical reimbursement records, etc.) on the interface, facilitating the approver to make accurate decisions. After the approver completes the approval operation (approve, reject, or return for modification), the contract automatically records the approval result and opinions, and pushes the reimbursement application to the next approval link (if any) or ends the approval process according to the approval process rules.
[0159] In addition, set the maximum approval time for each approval link, such as 2 working days. If the approver fails to process the reimbursement application within the specified time, the contract automatically triggers a reminder mechanism (such as sending a reminder notice), and after a certain period of overdue (such as 1 working day overdue), it is processed according to the preset rules, such as automatically transferring the application to the superior approver or marking it as abnormal for manual intervention.
[0160] III. Implementation of Data Verification Logic
[0161] Through the interaction between the smart contract and the enterprise budget management system or database, obtain the budget quota information of the employee's department, project or individual, including annual budget, monthly budget, and budget allocation of different expense types. When the reimbursement application is submitted, the contract calculates the amount that has been reimbursed cumulatively by the employee or project in the current period and compares it with the remaining budget. If the reimbursement amount plus the cumulatively reimbursed amount exceeds the budget quota, the contract marks the reimbursement application as "over budget" and operates according to the over-budget handling process stipulated by the enterprise, such as prompting the employee to adjust the reimbursement amount, requiring an additional explanation for budget overrun, or initiating a special approval process (such as joint approval of the over-budget part by the department manager and the financial manager).
[0162] According to the reimbursement item classification standard preset by the enterprise, the smart contract checks whether the reimbursement items submitted by the employee are correctly classified. For example, judge whether an expense should be classified as travel expense or business entertainment expense to ensure the accuracy of expense statistics and analysis. If the classification is inaccurate, the contract notifies the employee to modify the reimbursement item classification. At the same time, it also verifies whether the reimbursement items comply with the enterprise's compliance policies and relevant laws and regulations. For example, check whether the business entertainment expense exceeds the per capita standard or the entertainment scope stipulated by the enterprise; verify whether the purchase of office supplies comes from the enterprise's designated suppliers or within the reasonable procurement channels; ensure that the itinerary arrangement and expense details of travel expenses comply with the enterprise's travel policies (such as transportation mode selection, accommodation standards, etc.). If any non-compliance is found, the contract marks the reimbursement application as "non-compliant", details the non-compliant items and reasons, and notifies the relevant personnel for handling.
[0163] In summary, the reimbursement data management system and method based on AI artificial intelligence provided by the present invention have many significant beneficial effects, can effectively solve many problems of the traditional reimbursement management mode, improve the efficiency, quality and security of enterprise reimbursement data management, and promote the development of enterprise financial management towards intelligence and refinement. Specifically as follows:
[0164] 1. Improve efficiency and user experience
[0165] The multi-modal data acquisition and fusion module supports multiple terminals and input methods, docks with the enterprise internal system to obtain multi-source data and automatically fuses and verifies it, and can also provide real-time feedback on the reimbursement status, facilitating employee operation and ensuring data accuracy. The approval process optimization module uses reinforcement learning algorithms to automatically adjust the approval level according to the characteristics of reimbursement applications, simplifies the small-amount regular reimbursement process, adds approval links for major risk reimbursements, and predicts approval opinions at the same time, significantly improving the reimbursement efficiency.
[0166] 2. Enhance audit accuracy and compliance
[0167] The invoice authenticity identification sub-module verifies through deep learning and multi-platform interaction to ensure the authenticity and legality of invoices. The reimbursement reason analysis sub-module uses natural language understanding and machine learning technologies, combines multi-faceted knowledge bases and case bases, and accurately judges the rationality of reimbursement reasons. The reimbursement item classification sub-module is based on a rule engine and a machine learning model, accurately classifies reimbursement items and automatically audits them according to standards.
[0168] 3. Mine data value to assist decision-making
[0169] The analysis and prediction module uses a variety of data mining algorithms to mine the correlation between reimbursement data and other business data, provides cross-departmental business insights, and helps optimize resource allocation and business processes. Based on a machine learning prediction model, historical data is used to predict future reimbursement needs, etc., providing a scientific basis and forward-looking guidance for enterprise financial budgets, etc.
[0170] 4. Ensure data security and transparency
[0171] The data storage and blockchain module uses blockchain technology to build a distributed and immutable storage ledger, encrypts and stores data, and ensures data confidentiality, integrity and immutability through reasonable architecture selection, node setting, encryption algorithms and smart contracts. The blockchain smart contract is used to execute the automation rules of the reimbursement process, strictly enforce the data verification logic, including verification of invoice authenticity, amount compliance and item classification accuracy, etc., and manage user permissions at the same time to ensure the compliance and transparency of the process.
[0172] The above are the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A reimbursement data management system based on AI artificial intelligence, characterized in that: include: A multimodal data collection and fusion module is used to provide an accessible reimbursement application portal, where the reimbursement application portal is used to select manual input of reimbursement information, upload invoice pictures and other supporting documents, and voice input of reimbursement reasons; The audit and compliance determination module is used to build an invoice authenticity recognition model based on a deep learning algorithm, and interact with the tax department database and a third-party verification platform to verify the authenticity of invoices; use natural language understanding technology and machine learning algorithms to build a reimbursement reason analysis model; use a reimbursement item classification and audit model based on a combination of rule engines and machine learning to classify and automatically audit reimbursement items; The approval process optimization module is used to use reinforcement learning algorithms to determine the optimal approval path and approval node sequence based on the enterprise organizational structure, business processes, risk control strategies and historical approval data; The analysis and prediction module conducts multi-dimensional analysis of reimbursement data based on the big data processing framework and data mining algorithms; uses association rule mining, cluster analysis, and classification algorithms to mine the potential correlation between reimbursement data and other business data; uses historical data based on machine learning prediction models to predict future reimbursement needs, expense trends, and cost changes, and generates visual reports and results; The data storage and blockchain module is used to build a distributed, tamper-proof data storage ledger using blockchain technology, encrypt and store reimbursement data and related review and approval information, use blockchain smart contracts to execute reimbursement process automation rules, and execute data verification logic at the same time.
2. The AI-based reimbursement data management system according to claim 1 is characterized in that: The audit and compliance determination module includes an invoice authenticity identification submodule, a reimbursement reason analysis submodule, and a reimbursement item classification submodule, wherein: The invoice authenticity identification submodule is used for deep learning and training of massive invoice sample data, including the printing format, anti-counterfeiting mark, and code rules of the invoice. The verification data output by the deep learning algorithm is interactively verified with the authoritative invoice database of the tax department and the third-party invoice verification platform; The reimbursement reason analysis submodule is used to build an intelligent reimbursement reason analysis model using natural language understanding technology and machine learning algorithms. It performs semantic understanding and logical analysis on the reimbursement reason descriptions submitted by employees, and intelligently judges the rationality of reimbursement reasons by combining the company's detailed reimbursement policy knowledge base, industry standards and historical reimbursement data case library; The reimbursement item classification submodule adopts an intelligent classification and review model for reimbursement items based on a combination of rule engine and machine learning. According to the pre-set reimbursement item classification rules and the classification pattern automatically learned through machine learning, the reimbursement items are accurately classified according to the reimbursement item classification function, and different types of reimbursement items are automatically reviewed according to the corresponding review standards and threshold ranges.
3. The AI-based reimbursement data management system according to claim 1 is characterized in that: The approval process optimization module is used to utilize a reinforcement learning algorithm based on the enterprise organizational structure, business process, risk control strategy and historical approval data, specifically: Suppose the state space is S, which contains the state information of the amount of the reimbursement application, the expense type, the applicant's department, and the historical reimbursement credit record; the action space is A, which includes the different operations in the approval process; the reward function is R(s,a), which determines the reward value based on the approval result and business goal; the state transition probability is P(s t+1 |s t ,a t ), indicating that in state s t Take action a t Then transfer to state s t+1 The probability of, the value function is V(s), then the update formula of the value function is based on the Bellman equation: Among them, η is the learning rate and γ is the discount factor.
4. The AI-based reimbursement data management system according to claim 1 is characterized in that: The analysis and prediction module uses association rule mining, cluster analysis mining, and classification algorithms to mine the potential association between reimbursement data and other business data, specifically: (1) The potential relationship between reimbursement data and other business data is mined through association rules: Integrate reimbursement data with other relevant business data, including sales data, procurement data, inventory data, and human resources data; Clean the data, handle missing values, outliers and duplicate data, and encode and transform the categorical variables into numerical types; Use Apriori algorithm or FP-Growth algorithm to mine association rules, find frequent item sets and association rules between reimbursement data and other business data, analyze the mined association rules, and understand the business meaning of the association rules; (2) The potential correlation between reimbursement data and other business data is mined through cluster analysis. Specifically: Identify the features used for cluster analysis, including key indicators in reimbursement data and relevant features in other business data, and standardize the selected features so that different features have the same dimension and data distribution; Use K-Means clustering algorithm, DBSCAN algorithm or hierarchical clustering algorithm to perform clustering analysis, calculate clustering evaluation indicators, evaluate clustering effect, adjust clustering algorithm parameters or select a more appropriate clustering algorithm based on the evaluation results; Analyze clustering results, assign business significance to each cluster, and provide decision-making recommendations to enterprises based on clustering results; (3) The potential relationship between reimbursement data and other business data is mined through classification algorithms: Determine the classification task objectives, select and extract features related to the classification task, and use feature selection algorithms to select features with higher importance for classification model training; Select a classification algorithm, train the selected classification algorithm using a training data set, and adjust algorithm parameters to optimize model performance; Use the test data set to evaluate the trained classification model, calculate the evaluation index, apply the trained classification model with good evaluation index to the actual reimbursement data management, and perform classification prediction on new reimbursement applications or data.
5. The AI-based reimbursement data management system according to claim 2 is characterized in that: In the reimbursement reason analysis submodule, the reimbursement reason rationality judgment function is assumed to be: G(R)=α×M p (R)+β×M s (R)+γ×M c (R) Among them, M p (R) represents the matching evaluation function of R based on the reimbursement policy knowledge base, M s (R) represents the evaluation function of R based on industry standard specifications, M c (R) represents the similarity evaluation function of R based on the historical reimbursement data case library, α, β, γ are the corresponding weight coefficients, and α+β+γ=1. When G(R)≥T G , the reimbursement reason is reasonable; otherwise, it is unreasonable; among them, T G is the preset rationality threshold, and R is the employee's description of the reason.
6. The AI-based reimbursement data management system according to claim 2 is characterized in that: In the reimbursement item classification submodule, the reimbursement item classification function is specifically: Among them, Ln represents the nth category of reimbursement items, K n (X) represents the classification discriminant function of the nth category of reimbursement items, T n is the corresponding classification threshold.
7. The AI-based reimbursement data management system according to claim 1 is characterized in that: The data storage and blockchain module uses blockchain technology to build a distributed, tamper-proof data storage ledger, specifically: Choose a blockchain architecture based on enterprise size, data privacy needs, and performance requirements Set up full nodes and light nodes. Full nodes store complete ledger data and participate in transaction verification and block packaging, establish secure P2P network connections, and set up inter-node communication protocols; Use asymmetric encryption algorithms to encrypt reimbursement data and review and approval information, and select secure hash algorithms to hash the data; Build smart contracts to define the submission rules for reimbursement applications, embed data verification logic in the smart contracts, and deploy the written smart contracts to the nodes of the blockchain network; Define the block structure, each block contains a block header and a block body, and adopt a layered storage strategy to store recently frequently accessed data in high-speed storage media and historical data in large-capacity, low-cost storage.
8. The AI-based reimbursement data management system according to claim 1 is characterized in that: The data storage and blockchain module uses blockchain smart contracts to execute the reimbursement process automation rules, specifically: Divide the smart contract into multiple functional modules, including reimbursement application module, approval process module, data verification module, and notification module; Design the state transition logic of the smart contract based on the state machine model, which complies with pre-defined business rules and processes; When an employee submits an expense reimbursement application, the smart contract automatically checks whether the data format meets the requirements and verifies whether the expense reimbursement application complies with relevant business rules based on the company's expense reimbursement policy. Smart contracts automatically determine the approval level and order based on the amount of the reimbursement application, the type of expense, and the employee’s department; The smart contract automatically obtains the information of the corresponding approver from the enterprise organizational structure database or the predefined list of approvers, and reminds the approver to process the reimbursement application through system notification.
9. The AI-based reimbursement data management system according to claim 1 is characterized in that: The data verification logic of the data storage and blockchain module specifically includes: The reimbursement amount is checked against the budget. Through smart contracts, the company's budget management system or database is interacted with to obtain the budget information of the employee's department, project or individual. When the reimbursement application is submitted, the contract calculates the cumulative reimbursement amount of the employee or project in the current cycle and compares it with the remaining budget. Compliance check of reimbursement items: Based on the reimbursement item classification standards pre-set by the enterprise, the smart contract checks whether the reimbursement items submitted by employees are correctly classified, and verifies whether the reimbursement items comply with the enterprise's compliance policies and relevant laws and regulations.
10. A working method of the reimbursement data management system based on AI artificial intelligence as claimed in any one of claims 1 to 9, characterized in that: include: Providing an accessible reimbursement application portal, wherein the reimbursement application portal is used to select manual input of reimbursement information, upload of invoice pictures and other supporting documents, and voice recording of reimbursement reasons; Build an invoice authenticity recognition model based on deep learning algorithms, and interact with the tax department database and third-party verification platform to verify the authenticity of invoices; use natural language understanding technology and machine learning algorithms to build a reimbursement reason analysis model; use a reimbursement item classification and review model based on a combination of rule engine and machine learning to accurately classify and automatically review reimbursement items; Using reinforcement learning algorithms, based on the company's organizational structure, business processes, risk control strategies and historical approval data, we determine the optimal approval path and approval node sequence; Based on the big data processing framework and data mining algorithms, we conduct multi-dimensional analysis of reimbursement data; we use association rule mining, cluster analysis, and classification algorithms to mine the potential correlation between reimbursement data and other business data; based on machine learning prediction models, we use historical data to predict future reimbursement needs, expense trends, and cost changes, and generate visual reports and results; Blockchain technology is used to build a distributed, tamper-proof data storage ledger, which encrypts and stores reimbursement data and related review and approval information. Blockchain smart contracts are used to execute reimbursement process automation rules and data verification logic at the same time.
Citation Information
Patent Citations
Expense account risk prediction method and device, terminal equipment and storage medium
CN108364106A
Intelligent charge control reimbursement method capable of automatically matching according to application scene
CN112669133A
Data processing method and device, computer equipment and storage medium
CN113610504A
Financial reimbursement management system based on big data
CN115170269A
Mobile intelligent general reimbursement system, control method, equipment and terminal
CN116542790A
Cited By
Automatic process processing method for accounting data
CN120256507A
Method and device for analyzing and checking reimbursement data and medium
CN120278840A
Comprehensive financial auditing system based on big data
CN120430879A
Customs declaration intelligent dispatching and cooperative processing system
CN120471403A
Intelligent bill checking and signing method and system and medium
CN120635929A