An AI-based financial reimbursement method, device and electronic device
Through the financial accounting method of artificial intelligence, semantic analytical models and automated audits are used to solve the problem of error-prone traditional manual entry, and an efficient and accurate financial accounting process is achieved to adapt to the reimbursement needs of different companies.
Patent Information
- Application Number
- CN202510221307.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-02-27
AI Technical Summary
Traditional manual financial accounting methods are prone to errors, time-consuming and inefficient. The differences in reimbursement processes and form formats of different companies make it difficult for employees to adapt, increasing the complexity of accounting and the probability of errors.
Using an artificial intelligence-based financial accounting method, reimbursement data is obtained through semantic analytical models, match form templates and automatically fill in data, and combine with automatic review process to realize the automated processing of reimbursement forms.
减少了人工出错概率,缩短了报账周期,提高了工作效率,适应企业报销政策变化和多样化业务场景,降低了整体财务报账流程难度。
Smart Images

Figure CN119722360B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a financial reimbursement method, device and electronic device based on artificial intelligence. Background Art
[0002] Financial reimbursement is an important part of enterprise or organizational financial management. It mainly refers to the process in which employees or departments, in order to reimburse expenses incurred due to business activities, sort out and record relevant expense information according to the established financial regulations and procedures of the enterprise, and submit it to the financial department for review and processing, ultimately realizing the expense reimbursement and financial bookkeeping.
[0003] With the expansion of enterprise scale and the increase in business complexity, financial management has become increasingly important. And with the increase in the types of financial reimbursement attachments, the traditional manual entry of business documents for financial reimbursement leads to errors easily during the entry and form filling process, long form filling and review time and low efficiency. At the same time, due to differences in reimbursement processes and form formats among different enterprises, employees often face great challenges when adapting to new reimbursement requirements, which further increases the complexity of reimbursement and the possibility of errors. Summary of the Invention
[0004] The purpose of the present invention is to solve the problems in the prior art, and a financial reimbursement method, device and electronic device based on artificial intelligence are proposed.
[0005] In order to achieve the above purpose, the present invention adopts the following technical solutions:
[0006] A financial reimbursement method based on artificial intelligence, comprising the following steps:
[0007] Obtain target reimbursement data input by a target user;
[0008] Input the target reimbursement data into a semantic parsing model to obtain fill-in form intention data output by the semantic parsing model; the semantic parsing model is trained by sample reimbursement data and its corresponding fill-in form intention label results;
[0009] Match the fill-in form intention data with a form template library to obtain a target reimbursement template that matches the fill-in form intention data;
[0010] Based on the fill-in form intention data, perform data filling operations on the target reimbursement template to obtain a target reimbursement form;
[0011] Review the target reimbursement form, and perform financial reporting based on the review confirmation result.
[0012] A financial reimbursement device based on artificial intelligence, comprising:
[0013] An acquisition unit for acquiring target reimbursement data input by a target user;
[0014] A model output unit for inputting the target reimbursement data into a semantic parsing model to obtain fill-in form intention data output by the semantic parsing model; the semantic parsing model is trained by sample reimbursement data and its corresponding fill-in form intention label results;
[0015] A matching unit for matching based on the fill-in form intention data with a form template library to obtain a target reimbursement template matched by the fill-in form intention data;
[0016] A writing unit for performing data filling operations on the target reimbursement template based on the fill-in form intention data to obtain a target reimbursement form;
[0017] An upload unit: for auditing the target reimbursement form and performing financial reporting based on the audit confirmation result.
[0018] The present invention also provides an electronic device, including a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, the steps of the above-mentioned financial reimbursement method based on artificial intelligence are implemented.
[0019] The present invention also provides a non-transitory readable storage medium, on which a program is stored. The program is characterized in that when the program is executed by a processor, the steps of the above-mentioned financial reimbursement method based on artificial intelligence are implemented.
[0020] The present invention has the following advantages compared with the prior art:
[0021] A financial reimbursement method, device, and electronic device based on artificial intelligence provided by the present invention first input target reimbursement data into a semantic parsing model for semantic parsing to obtain a fill-in form intention, and then match the fill-in form intention with a form database to obtain a corresponding target reimbursement template. Compared with the traditional method, it does not require manual entry one by one, which not only avoids problems that are prone to occur during the manual entry of documents, but also for new and old employees, it can avoid the problem of unfamiliarity with the reimbursement process and differences in reimbursement forms, and can flexibly match corresponding reimbursement templates according to different reimbursement data to adapt to changes in enterprise reimbursement policies and diverse business scenarios; then, through automatic data filling of the target reimbursement form and auditing and reporting the filled target reimbursement form, the automation of form data is realized, greatly shortening the reimbursement cycle, further reducing the probability of manual errors, and reducing the difficulty of the overall financial reimbursement process. Description of the Drawings
[0022] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0023] Figure 1 Schematic flowchart of the financial reimbursement method based on artificial intelligence provided by the embodiment of the present invention;
[0024] Figure 2 Schematic structural diagram of the financial reimbursement device based on artificial intelligence provided by the embodiment of the present invention;
[0025] Figure 3 Schematic structural diagram of the electronic device proposed by the present invention. Detailed implementation manners
[0026] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0027] The following will describe Figure 1 - Figure 3 a financial reimbursement method, device, and electronic device based on artificial intelligence of the present invention.
[0028] Figure 1 is a schematic flowchart of a financial reimbursement method based on artificial intelligence provided by the present invention. As Figure 1 shown, the method includes:
[0029] Step 101, obtaining target reimbursement data input by a target user.
[0030] Specifically, in the process of obtaining target reimbursement data, the target user enters the reimbursement data through a dedicated financial reimbursement software platform, the enterprise's internal information system, etc., so as to obtain the target reimbursement data. And the target reimbursement data includes voice information, text information or picture information input by the target user within a preset time. The preset time is a scheduled input time set according to requirements. The combination of multiple input methods and quick functions reduces the time and effort for users to manually fill in cumbersome information and improves the user experience. Among them, the preset time refers to the scheduled input time set according to requirements, so that when the target user enters the reimbursement data, only part of the data needs to be entered to proceed with the subsequent reimbursement process, without having to enter all of it, in order to reduce the input recognition difficulty and improve the overall input efficiency.
[0031] Furthermore, the voice information is the audio data input during the input process. When the system inputs it, it converts it into text through the built-in voice conversion technology and then proceeds with the subsequent input. The text information includes text description information (such as the purpose of the expense, itinerary details, etc.), numerical information (such as amount, quantity, etc.), date data (such as the expense occurrence date, reimbursement date, etc.), and relevant document information (such as contracts, etc.). The picture information includes relevant voucher pictures and invoice photos, etc. After the picture information is input into the system, the text content in the picture is extracted and converted into an editable text format through optical character recognition technology.
[0032] Even further, during the data acquisition process, the system will also perform preliminary format verification and integrity checks on the input data, such as checking whether the date format conforms to the specification of "YYYY-MM-DD", whether the amount is a number and reserved to the appropriate decimal places, whether there is content filled in the required fields, etc. If a format error or incompleteness is found, the system will promptly pop up a prompt box to inform the user to correct it, so as to ensure that the data format in the subsequent process is basically correct and improve the data quality and processing efficiency.
[0033] Step 102, input the target reimbursement data into the semantic parsing model to obtain the fill-in form intention data output by the semantic parsing model; the semantic parsing model is trained by sample reimbursement data and its corresponding fill-in form intention label results.
[0034] Specifically, when training the semantic parsing model, first collect a large amount of historical reimbursement data covering various types of enterprise reimbursement situations (such as travel expenses, office supplies procurement expenses, business entertainment expenses, etc. in different departments and business scenarios), and then label the corresponding fill-in form intention label results for each sample reimbursement data according to the established enterprise reimbursement rules. For example, for a travel expense reimbursement sample, its fill-in form intention labels may include business trip, transportation method and expense details, accommodation arrangement and expenses, and meal subsidy standards, etc.
[0035] Further, the sample reimbursement data needs to be preprocessed before model training, including text tokenization, stop word removal, and extraction of key feature vectors (such as expense type, amount range, payment method, etc.). Using the feature vectors as input and the form filling intention label as output, deep learning algorithms (such as RNN, LSTM, or models based on attention mechanisms, etc.) are used for training, and the parameters are continuously adjusted during training to minimize the loss function between the predicted form filling intention and the true label. Specifically, it is as described in steps 1021 - 1023.
[0036] Furthermore, in one embodiment, when the target reimbursement data is input into the trained semantic parsing model, the model will perform feature extraction and semantic understanding on it. For example, for the above-mentioned reimbursement data of the salesperson, the model will identify the expense type of business trip expenses, extract the amount information of various expenses such as transportation, accommodation, and catering, as well as key intention information such as the purpose of the business trip of business negotiation, and output a form filling intention data structure similar to {expense type: business trip expenses, transportation expenses: [round-trip air ticket amount, in-city transportation amount], accommodation expenses: [hotel name, number of accommodation days, amount], business purpose: customer negotiation, catering expenses: [amount]}. It should be noted that not all form filling data structures adopt the above form, and it can be adjusted according to actual needs and company requirements.
[0037] Step 103: Match the form filling intention data with the form template library to obtain the target reimbursement template that matches the form filling intention data.
[0038] Specifically, during the construction of the form template library, the enterprise designs and establishes it according to its own financial reimbursement system, tax regulations requirements, and business needs. Each template is customized for different expense types and business scenarios, covering from the form header information (such as company name, reimbursement number, reimbursement date, etc.) to the detailed expense item details (corresponding required and optional fields are set according to different expense types, such as transportation mode, itinerary start and end locations, number of accommodation days and standards, catering subsidy standards, etc. in the business trip expense template; item name, specification, quantity, unit price, supplier information, etc. in the office supplies procurement template), as well as information related to the approval process (clearly defining the positions and authorities of approval personnel at all levels, such as expenses within the approval authority of the department manager, key information related to financial compliance approved by the financial person in charge, and reimbursement of large amounts or special situations approved by the company's senior leadership, etc.). The reimbursement templates are stored in the database in a structured format for convenient and quick retrieval and call.
[0039] Furthermore, match the form filling intention data with the form template library to obtain the target reimbursement template that matches the form filling intention data. The specific process is as described in steps 1031 - 1035.
[0040] Step 104, based on the form filling intention data, the target reimbursement template is filled with data to obtain the target reimbursement form.
[0041] Specifically, according to the matched target reimbursement template, the information in the form filling intention data is accurately filled in according to the structure and field requirements of the template. For text-type fields, such as the description of the purpose of the expense, the name of the business project, etc., the corresponding text content in the form filling intention data is directly filled in the corresponding position of the template to ensure that the information is complete and accurate. For numeric fields, such as the amount and quantity of each expense, strict data format verification and conversion (such as converting the amount string in the form filling intention data to a numeric type and retaining the appropriate number of decimal places according to the template requirements) are performed before filling in the template. At the same time, for some fields that need to be calculated according to the company's financial rules, such as the conversion of the amount including tax and the amount excluding tax, and the calculation of the actual reimbursement amount according to the reimbursement ratio, the system will automatically call the pre-set calculation formula for calculation and filling. For example, if the company's reimbursement policy stipulates that certain expenses are reimbursed at a rate of 80%, the system will automatically calculate the actual reimbursement amount based on the original expense amount in the form filling intention data and fill it in the corresponding field.
[0042] In addition, for some fields involving associated information, such as department information in the form filling intention data, the budget code or person in charge information of the department is automatically associated and filled into the template to improve the information integrity and accuracy of the reimbursement form, so that it complies with the company's financial management standards and approval process requirements, and generates a complete, standardized and accurate target reimbursement form, including all necessary header information, expense details, approval signature column and other content, which is convenient for subsequent review and financial processing. Through the automated data filling process, the problems of typos, omissions or inconsistent formats that may occur in manual filling are avoided, the accuracy and consistency of the data in the reimbursement form are guaranteed, the quality of financial data is improved, and reliable basic data is provided for financial review and statistical analysis. At the same time, the form filling intention data is quickly and automatically filled into the target reimbursement template, which greatly shortens the time to generate the reimbursement form. Compared with the traditional method of manual filling or manual copying and pasting of data, it significantly improves work efficiency and speeds up the entire financial reimbursement process.
[0043] Step 105, review the target reimbursement form and make a financial report based on the review and confirmation result.
[0044] Specifically, during the audit process, first, a comprehensive automated audit is conducted on the target reimbursement form based on the financial rules and audit criteria preset by the enterprise. For example, for the audit of expense rationality, by comparing with the enterprise's internal expense standard database, it is judged whether each expense is within a reasonable range. Taking business trip expenses as an example, according to the business trip subsidy standards for employees at different levels, check transportation expenses (such as the price ceiling corresponding to the flight ticket class, regulations on train seat types, etc.), accommodation expenses (the upper limit of accommodation standards in different cities), and meal subsidies to see if they comply with company policies.
[0045] Furthermore, for the audit of data integrity, ensure that all required fields in the reimbursement form (such as the name of the reimburser, department, expense details, invoice number, etc.) are accurately filled to avoid rejection during the audit or difficulties in subsequent financial processing due to missing information. At the same time, verify the invoice information, including querying the authenticity of the invoice (by connecting to the invoice query interface of the tax department, inputting key information such as the invoice number, code, and invoice date to obtain the authenticity verification result of the invoice), checking the consistency between the invoice title and the enterprise name, and reconciling the invoice amount with the reimbursement amount.
[0046] For the audit of logical consistency, check whether the logical relationships between the data in the reimbursement form are correct. For example, the total reimbursement amount should be equal to the sum of the amounts of each expense detail; if there is an expense sharing situation, check whether the sum of the sharing ratios is 100% and whether the calculated amount after sharing is accurate. For expense reimbursements involving multiple projects or departments, verify the reasonableness and accuracy of the expense allocation to ensure that the expenses borne by each project or department are consistent with the actual business situation.
[0047] Meanwhile, for the reimbursement forms marked as abnormal or high-risk during the automated audit process, as well as some complex business scenarios (such as expense reimbursements related to major projects, expense handling of special business matters, etc.), the system automatically transfers them to the manual audit link. Financial personnel can view the detailed information of the reimbursement form in the system, including the original target reimbursement data, the data of the filling intention after semantic analysis, the matching target reimbursement template, and the audit doubts and risk prompt information given by the system, and make further audit judgments based on their professional knowledge and experience.
[0048] Furthermore, an audit record and feedback function is also set up. The system details the results of both automated and manual audits, including different audit statuses and opinions such as audit passed, audit failed (specific reasons for failure need to be clearly listed, such as cost overrun, invoice missing, data logic error, etc.), and materials need to be supplemented. The audit result information is fed back to the target user in a timely manner, enabling them to clearly understand the processing status of the reimbursement application and make corresponding modifications and improvements according to the feedback. At the same time, the audit result data is also stored in the enterprise's financial audit database for subsequent statistical analysis and audit traceability. After the audit is passed, the target reimbursement form and approval record are uploaded to facilitate subsequent process operations. Thus, by fully implementing the automated audit rules, a large number of reimbursement forms can be processed quickly, significantly shortening the audit cycle and improving the audit efficiency. Compared with the traditional method of manually reviewing reimbursement forms one by one, accurate comparison and verification of various data can be carried out in a short time, avoiding audit errors and omissions caused by human factors and significantly improving the accuracy of the audit.
[0049] The present invention relates to the field of artificial intelligence, and provides a financial reimbursement method, device and electronic device based on artificial intelligence. In the present invention, the proposed financial reimbursement method based on artificial intelligence, compared with the traditional method, does not require manual entry one by one, which not only avoids the problem of easy errors when manually entering documents, but also enables new and old employees to avoid the troubles caused by unfamiliarity with the reimbursement process and form differences. Moreover, it can flexibly match corresponding reimbursement templates according to different reimbursement data, adapting to the changes in enterprise reimbursement policies and diverse business scenarios. Then, through automatic data filling of the target reimbursement form and submitting the filled target reimbursement form for review, the automation of form data is achieved, greatly shortening the reimbursement cycle, further reducing the probability of human errors, and lowering the overall difficulty of the financial reimbursement process.
[0050] In one embodiment, in step 102, the semantic parsing model is obtained by fusing a first target model, a second target model, and a third target model.
[0051] Among them, the training steps of the first target model in step 1021 include:
[0052] Obtain the first sample data; the first sample data includes the structured data in the historical financial reimbursement data; input the first sample data into the first pre-trained model to obtain the first prediction result output by the first pre-trained model; obtain the first loss value based on the first prediction result according to the first loss function of the first pre-trained model; adjust the first optimization function in the first pre-trained model based on the first loss value; predict the prediction result of the first sample data based on the adjusted first pre-trained model until the loss value output by the first loss function is less than or equal to the first preset value, the square of the difference between adjacent loss values is less than 1, and the difference between at least three consecutive adjacent loss values is greater than zero, to obtain the first target model. Among them, the first preset value can be set to 0.01.
[0053] Specifically, for the first sample data, extract the historical financial reimbursement records from the enterprise's financial database, preprocess these records, and screen out the structured data part. Structured data has a clear format and fixed semantics. For example, date (following a specific date format, such as "YYYY-MM-DD"), amount (exact numerical type, possibly including decimal places and currency units), employee number (usually a number or a specific coding format), expense category code (pre-defined classification code, such as 1 representing travel expenses, 2 representing office supplies expenses, etc.). These data can be efficiently extracted and organized through database query statements and data cleaning rules to form the first sample data set.
[0054] Furthermore, the first pre-trained model can be a multi-layer perceptron MLP, a deep neural network, etc. Taking the multi-layer perceptron MLP as an example, in a model based on MLP, the input layer receives the feature values of the structured data, and through the neurons in the hidden layer for weighted summation and non-linear transformation (using an activation function, such as the ReLU function), finally outputs the prediction result (such as 0 or 1 for a binary classification problem, indicating non-compliance or compliance) at the output layer.
[0055] Furthermore, the first loss function is:
[0056] ;
[0057] Among them, N represents the number of samples of the first pre-trained data, and M represents the dimension of the semantic representation of professional vocabulary; represents the value of the true label (the correct semantic representation of the professional vocabulary after standardization processing) of the i-th sample in the first pre-trained data on the j-th dimension, represents the value of the prediction output of the first model for the i-th sample on the j-th dimension, Denote the weight coefficient. The first half of the first loss function adopts the cross-entropy function, which can reflect the inaccuracy of the model prediction. When the probability distribution predicted by the model is closer to the probability distribution of the true label, the value of the cross-entropy is smaller. The second half introduces a constraint on the overall difference between the prediction result and the true label as a whole, and adjusts the importance of this part in the loss function through the weight coefficient to adjust the importance of this part in the loss function. Since in the financial reimbursement data, there may be some situations where although there is a certain deviation in the classification probability (the cross-entropy may not change much), the overall prediction result has a large difference in value from the true label. For example, in a certain dimension of predicting the reimbursement amount, the sum difference between the model prediction value and the true value is large. This difference is reflected by the absolute value part (the second half), and the influence of this part in the loss function is controlled by the weight coefficient, so as to prevent the model from only focusing on the classification probability and ignoring the accuracy of the overall value.
[0058] By substituting the prediction result into this formula, the first loss value is calculated. This value reflects the degree of difference between the model prediction result and the true label. The larger the loss value, the more inaccurate the model prediction is.
[0059] Furthermore, during the model training process, the adjustment of the first optimization function is continuously repeated until three stopping conditions are finally met to meet the training requirements. Among them, the loss value output by the first loss function is less than or equal to the first preset value, which indicates that the model has reached a certain accuracy requirement. Continuing training may not bring significant performance improvement and may even lead to overfitting. The square of the difference between adjacent loss values is less than 1. This is to avoid overly drastic fluctuations or unstable situations during the model training process. Let the loss value of the t-th iteration be , and the loss value of the (t - 1)-th iteration be , then the square of the difference between adjacent loss values is . If this condition is not met, it means that the model may have abnormal changes under the current parameter update, and it is necessary to further adjust the optimization strategy or check whether there are problems with the data. For at least three consecutive loss values, the difference between adjacent values is greater than zero, which means that the model is still continuously improving in several consecutive iterations, and the loss value is steadily decreasing, rather than falling into a local minimum or oscillating. In summary, setting multiple stopping conditions can comprehensively consider the accuracy, stability, and improvement trend of the model, avoid premature stopping of model training (underfitting) or overtraining (overfitting), and ensure that the obtained first target model has good generalization ability and performance. Through fine control of the training process, a model most suitable for processing structured financial reimbursement data can be obtained under limited training time and resources, improving the applicability and reliability of the model.
[0060] Step 1022, the training steps of the second target model, including:
[0061] Obtain the second sample data; the second sample data includes semi-structured data in the historical financial reimbursement data; input the second sample data into the second pre-trained model to obtain the second prediction result output by the second pre-trained model; obtain the second loss value based on the second prediction result according to the second loss function of the second pre-trained model; adjust the second optimization function in the second pre-trained model based on the second loss value; predict the prediction result of the second sample data based on the adjusted second pre-trained model until the loss value output by the second loss function is less than or equal to the second preset value and the product of adjacent loss values is greater than 0 and less than 1, to obtain the second target model. Among them, the second preset value can be set to 0.05.
[0062] Specifically, the semi-structured data is also obtained from the historical financial reimbursement data, such as expense detail lists (listing the names, amounts, and brief descriptions of various expenses in a specific format, but the format is not as strictly fixed as structured data), approval process records (including approvers, approval opinions, and approval times, but there may be some non-standard expressions and format differences), and some text data with simple tags or classifications. By using text parsing techniques and data extraction tools, these semi-structured data are converted into a format suitable for model input. For example, the expense detail list is converted into a table form, and the approval process record is extracted as a list of key information pairs (approver - approval opinion), etc., to form the second sample data set.
[0063] Furthermore, in the stopping condition of the second target model, the loss value output by the second loss function is less than or equal to the second preset value, which indicates that the prediction error of the model on the semi-structured data has been reduced to an acceptable range, and the model has reached a certain accuracy requirement. The product of adjacent loss values is greater than 0 and less than 1, which is to ensure the stability and continuous improvement of the model during the training process, and also prevent the model from oscillating or prematurely converging to a local minimum during the training process, ensuring that the model still has good generalization ability and adaptability to new data while reaching a certain accuracy.
[0064] Furthermore, the second pre-trained model can adopt a recurrent neural network (RNN), a long short-term memory network (LSTM), and a gated recurrent unit (GRU). Taking the recurrent neural network (RNN) as an example, for the expense detail list data, each expense item is sequentially input into the RNN model. The model gradually updates its internal state based on the information of the previous expense item and the characteristics of the current item, and finally outputs the prediction result for the entire expense detail, such as predicting whether the expense is reasonable or whether there are abnormal situations.
[0065] Furthermore, the second loss function is:
[0066] ;
[0067] Among them, represents the number of samples of the second pre-training data, represents the dimension of the semantic representation of the converted professional vocabulary, represents the value of the true label of the professional vocabulary converted from the original spoken vocabulary of the i-th sample in the second pre-training data on the j-th dimension, represents the value of the predicted output of the i-th sample in the second pre-training data on the j-th dimension, is an indicator function, indicating whether the second model successfully integrates the spoken vocabulary into the existing semantic system for the i-th sample (if the integration is successful, ; otherwise ) represents the penalty coefficient for failed integration. The second loss function measures the loss in different ways according to whether the model successfully integrates the spoken vocabulary, so as to guide the model to better process the semi-structured financial reimbursement data containing spoken vocabulary. When it is the case that this part calculates the loss, which helps the model to make the predicted semantic representation of the professional vocabulary close to the true semantic representation of the professional vocabulary in terms of probability distribution , which can prompt the model to select a more accurate professional vocabulary representation during prediction and improve the accuracy of semantic understanding. When it is the case that a fixed penalty term
[0068] Step 1023, the training steps of the third target model, include:
[0069] Obtain the third sample data; the third sample data includes unstructured data in historical financial reimbursement data; input the third sample data into the third pre-trained model to obtain the third prediction result output by the third pre-trained model; obtain the third loss value based on the third prediction result according to the third loss function of the third pre-trained model; adjust the third optimization function in the third pre-trained model based on the third loss value; predict the prediction result of the third sample data based on the adjusted third pre-trained model until the loss value output by the third loss function is less than or equal to the third preset value and the reciprocal of the product of the phase loss values is greater than the fourth preset value to obtain the third target model. Among them, the third preset value can be set to 0.1. The third pre-trained model can be a language model based on the Transformer architecture (such as GPT, BERT and their variants). Input the preprocessed third sample data into the model, and the model models the relationship between each word in the text and other words through its multi-head attention mechanism, so as to capture the semantic information and context dependence of the text.
[0070] Specifically, the unstructured data is collected from historical financial reimbursement records, such as expense purpose descriptions (for example, due to discussing business with customers, entertaining customers at XX restaurant, discussing cooperation details, during which catering expenses and transportation expenses are incurred, aiming to promote project cooperation), special situation descriptions (due to the temporary cancellation of the flight, the cost has increased due to repurchasing the ticket, and relevant certificates have been retained after communicating with the airline), and text content in attachments (such as clause descriptions in contracts, discussion items in meeting minutes, etc.). Then, natural language processing techniques are used, such as text cleaning tools to remove noise (such as redundant punctuation marks, special characters, garbled codes, etc.), lexical analysis tools for word segmentation (splitting the text into individual words) and part-of-speech tagging (tagging the part of speech of each word, such as nouns, verbs, adjectives, etc.), and then key entities (such as people, places, organizations, amounts, times, etc.) are extracted through entity recognition techniques to form a structured text representation as the third sample data.
[0071] Furthermore, in the stopping condition of the third target model, the loss value output by the third loss function is less than or equal to the third preset value, which indicates that the prediction error of the model on the unstructured data has been reduced to an acceptable range, and the model has reached a certain accuracy requirement. The reciprocal of the product of the phase loss values is greater than the fourth preset value, which is to ensure the stability and continuous improvement of the model during the training process. The fourth preset value can be adjusted according to the actual training requirements, and it can prevent the model from oscillating or converging to a local minimum prematurely during the training process, ensuring that the model still has good generalization ability and adaptability to new data while reaching a certain accuracy.
[0072] The third preset loss function is:
[0073] ;
[0074] ;
[0075] ;
[0076] Among them, is mainly used for processing the correction of fuzzy words, is mainly used for processing the semantic analysis of special words, and balances the importance of these two parts in the overall loss function through the weight coefficient to effectively calculate the loss and optimize the model for different characteristics (fuzzy words and special words) in unstructured financial reimbursement data respectively. If is close to 1, grinding pays more attention to the correction of fuzzy words. If is close to 0, it pays more attention to the semantic analysis of special words. K represents the number of fuzzy words in the third pre-training data, is an indicator function, indicating whether the predicted correction of the fuzzy words in the i-th sample by the third model is correct. (If it is correct, ; otherwise ), represents the penalty coefficient when the correction of fuzzy words is incorrect, is the weight coefficient, represents the probability value that the predicted correction result of the fuzzy words in the i-th sample by the third model is converted into correct; L represents the number of samples in the fourth pre-training data, and P represents the dimension of the semantic labels of special words, represents the value of the true semantic label of the special words in the i-th sample in the fourth pre-training data on the j-th dimension, represents the value of the predicted parsing result of the special words in the i-th sample by the third model on the j-th dimension, represents a preset minimum value, which is used to prevent the denominator in the logarithmic operation from being zero.
[0077] Furthermore, by combining the two parts of fuzzy word correction and special word semantic analysis, the loss function can comprehensively evaluate the performance of the model when processing unstructured financial reimbursement data. During the training process of the model, it is necessary to strive to accurately correct fuzzy words (through the optimization of the part) and correctly parse the semantics of special words (through the optimization of the part). It can improve the processing ability of the model for complex word situations in unstructured data, enabling the model to better adapt to the diversity and complexity of financial reimbursement data in actual business.
[0078] It should be noted that the first target model, the second target model, and the third target model are trained for structured, semi-structured, and unstructured data respectively, and their respective stopping conditions ensure good performance when processing corresponding types of data. During fusion, this accuracy can be preserved and integrated. For example, the high-precision judgment of the first target model on structured data (such as accurately identifying amounts, dates, etc.), the precise grasp of the logical relationships of semi-structured data by the second target model (such as expense details and approval processes), and the accurate understanding of the semantics of unstructured data by the third target model (such as expense purpose descriptions) complement each other. When processing a complete financial reimbursement data, the fusion model can accurately analyze from multiple perspectives, avoiding errors caused by the deficiencies of a single model in processing a certain type of data, thereby improving the overall accuracy. In addition, different stopping conditions ensure the stability of each model during training. During fusion, this stability can prevent the occurrence of unstable training after model fusion. Finally, the three models are trained for different types of data respectively, and the stopping conditions consider the generalization ability of the models. After fusion, the fusion model can inherit this generalization advantage and better adapt to the financial reimbursement data of different enterprises.
[0079] In one embodiment, steps 1031-1035 are as follows:
[0080] In step 1031, key features are extracted from the form filling intention data and the form template library to obtain a first form filling feature vector and a form feature vector respectively.
[0081] Specifically, for the form filling intention data, natural language processing techniques and data mining algorithms are used to extract key features such as expense type (travel expenses, office supplies expenses, business entertainment expenses, etc.), amount range (through numerical recognition and range division, such as 0-1000 yuan, 1001-5000 yuan, etc.), business association (involved project names, departments, customers, etc.), time information (expense occurrence time, reimbursement time, etc.), and convert them into vector form, that is, the first form filling feature vector.
[0082] Furthermore, for each template in the form template library, key features are extracted, such as the expense type applicable to the template (determined by the template name and preset expense type labels), the preset amount interval in the template (the amount filling range set in the template according to different expense standards of enterprise finance), the business department or project type associated with the template (obtained from the template usage instructions or preset association information), the time format requirements of the template (such as specific date filling formats and time limits), etc., and these features are converted into vector form to obtain the form feature vector.
[0083] By extracting key features and converting them into vector form, complex text data and template information can be quantitatively represented, facilitating efficient calculation and comparative analysis.
[0084] Step 1032: Obtain the historical reimbursement feature vector based on the first form-filling feature vector and the form feature vector, and calculate the historical usage frequency of each template in the form template library in the historical reimbursement records.
[0085] Specifically, first collect historical reimbursement records related to the current form-filling intention data and form templates from the enterprise's historical reimbursement database. For each historical reimbursement record, extract the corresponding form-filling feature vector and the form template information used (i.e., the first form-filling feature vector and the form feature vector), conduct aggregation analysis on the form-filling feature vectors, such as using the K-Means clustering algorithm to cluster similar form-filling feature vectors into different categories, where each category represents a historical reimbursement characteristic. Then, count the number of occurrences of each form template in each historical reimbursement characteristic category and divide it by the total number of reimbursement records in that category to obtain the usage frequency of each template under different historical reimbursement characteristics.
[0086] Furthermore, in one embodiment, the goal of the K-Means clustering algorithm is to divide the data set , ( which can be the form-filling feature vector of the historical reimbursement record) into clusters such that the sum of squares within the clusters is minimized. Its iterative update process is as follows:
[0087] First, randomly initialize cluster centers .
[0088] Repeat the following steps until convergence:
[0089] For each data point , calculate its distance (such as the Euclidean distance) from each cluster center : , where and represent the and th eigenvalue of and respectively, and n represents the dimension of the feature vector. Assign to the cluster where the nearest cluster center is located. For each cluster , update its cluster center: ; where represents the data of the data points in cluster .
[0090] After the clustering converges, for each cluster , count the number of times each form template is used , then the usage frequency of the template in the cluster is: ; where LL represents the total number of templates in the form template library. Combine the usage frequencies of each template in each cluster to form a historical reimbursement feature vector for subsequent similarity weighted calculation.
[0091] Step 1033: Based on the historical usage frequency of the template, weight the original similarity to obtain the weighted similarity.
[0092] Specifically, use the cosine similarity to calculate the original similarity between the first form filling feature vector and the form feature vector. The calculation formula is: , where represents the dot product operation of and represents the modulus length of the vector , represents the modulus length of the vector ; is the first form filling feature vector = ( ), is the form feature vector = ( ). Then, use the template historical usage frequency in the historical reimbursement feature vector obtained in the previous step as the weight to weight the original similarity. The historical reimbursement feature vector (R is the number of categories of historical reimbursement characteristics), then the weighted similarity , where represents the original similarity calculated under the jth historical reimbursement characteristic (the original feature vector needs to be appropriately adjusted or recalculated according to different historical reimbursement characteristics). In this way, the template that better matches the historical reimbursement situation gets a greater weight in the similarity calculation, increasing its probability of being selected.
[0093] Step 1034: Based on the weighted similarity, determine the clusters to be selected among multiple clustering sets; the clusters to be selected are the top preset number of clusters among multiple clustering sets.
[0094] Specifically, all templates in the form template library are clustered according to clustering methods (such as hierarchical clustering, DBSCAN, etc.) to form multiple clustering sets. The basis for clustering can be the similarity of the key feature vectors of the templates (for example, templates with similar features such as expense type, amount range, business association, etc. are clustered into one category). Then, according to the weighted similarity calculated in the previous step, the templates in each cluster are sorted, and the top preset number of clusters with higher weighted similarity are selected as the clusters to be selected. The top preset number of clusters can be set as needed. For example, using the hierarchical clustering method, the form template library is divided into 5 clustering sets, and each clustering set contains several templates. For each clustering set, calculate the weighted similarity between the internal templates and the fill-in form intention data, and sort them from high to low according to the similarity. If it is preset to select the top 2 clusters as the clusters to be selected, then select the clusters where the top 2 templates with the highest similarity in each clustering set are located. Finally, 2 clusters to be selected are obtained. The templates in the two clusters to be selected will be used as the candidate set for further screening the target reimbursement templates, narrowing the scope of subsequent searches and improving the efficiency of template selection.
[0095] Step 1035, obtain the overall similarity between the second fill-in form feature vector in the cluster to be selected and each template, and based on the overall similarity, determine the target reimbursement template; the second fill-in form feature vector is the vector after dimensionality reduction of the first fill-in form feature vector.
[0096] Specifically, use dimensionality reduction techniques (such as principal component analysis PCA, singular value decomposition SVD, etc.) to perform dimensionality reduction processing on the first fill-in form feature vector to obtain the second fill-in form feature vector. The purpose of dimensionality reduction is to reduce the dimension of the feature vector while retaining the main information, reduce the computational complexity, and at the same time remove some possible noise features. For example, through the PCA algorithm, project the first fill-in form feature vector onto a low-dimensional subspace to obtain the second fill-in form feature vector , and the specific process is as follows:
[0097] First, calculate the covariance matrix of the first fill-in form feature vector , (where Q is the number of samples. That is, the number of historical fill-in form data);
[0098] Then, perform eigenvalue decomposition on the covariance matrix to obtain the eigenvalues and the corresponding eigenvectors , and sort the eigenvectors in descending order of eigenvalues;
[0099] Select the first w eigenvectors (w is the dimension after dimensionality reduction) to form the projection matrix ;
[0100] The second fill-in form feature vector is calculated through the following formula: 。
[0101] After that, for each template in the cluster to be selected, calculate its overall similarity with the second fill-in form feature vector. The overall similarity here can comprehensively consider various factors. In addition to the previous feature vector similarity (such as cosine similarity), some business rule similarities (such as the matching degree between the preset approval process in the template and the implicit approval requirements in the fill-in form intention) and text semantic similarities (perform semantic analysis on the text description in the template and the expense purpose description in the fill-in form intention and calculate the similarity) can also be added. For example, use the weighted average method to calculate the overall similarity ; among them, represents the feature vector of the template, represents the cosine similarity calculation function, represents the business rule similarity calculation function, represents the text semantic similarity calculation function, 、 、 represent the corresponding weight coefficients, which are set according to the actual business needs and data characteristics of the enterprise. Finally, select the template with the highest overall similarity as the target reimbursement template, complete the template matching and selection process, and provide an accurate template basis for subsequent data filling operations
[0102] Figure 2 is a schematic structural diagram of a financial reimbursement device based on artificial intelligence provided by the present invention, as shown in Figure 2 shown. The device includes: an acquisition unit 10 for acquiring target reimbursement data input by a target user; a model output unit 20 for inputting the target reimbursement data into a semantic parsing model to obtain fill-in form intention data output by the semantic parsing model; the semantic parsing model is trained by sample reimbursement data and its corresponding fill-in form intention label results; a matching unit 30 for matching based on the fill-in form intention data with a form template library to obtain a target reimbursement template matching the fill-in form intention data; a writing unit 40 for performing data filling operations on the target reimbursement template based on the fill-in form intention data to obtain a target reimbursement form; an upload unit 50: for auditing the target reimbursement form and performing financial reporting based on the audit confirmation result
[0103] Figure 3 is a schematic structural diagram of an electronic device provided by the present invention, as shown in Figure 3As shown, the electronic device may include: a processor 310, a communications interface 320, a memory 330, and a communication bus 340. Among them, the processor 310, the communications interface 320, and the memory 330 communicate with each other through the communication bus 340. The processor 310 can call the logical instructions in the memory 330 to execute an artificial intelligence-based financial reimbursement method, device, and electronic device. The method includes: obtaining target reimbursement data input by a target user; inputting the target reimbursement data into a semantic parsing model to obtain fill-in form intention data output by the semantic parsing model; the semantic parsing model is trained by sample reimbursement data and its corresponding fill-in form intention label results; matching the fill-in form intention data with a form template library to obtain a target reimbursement template that matches the fill-in form intention data; performing a data filling operation on the target reimbursement template based on the fill-in form intention data to obtain a target reimbursement form; auditing the target reimbursement form and performing financial reporting based on the audit confirmation result.
[0104] In addition, when the logical instructions in the above-mentioned memory 330 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0105] On the other hand, the present invention also provides a program product. The program product includes a program stored on a non-transitory readable storage medium. The program includes program instructions. When the program instructions are executed, they can execute an artificial intelligence-based financial reimbursement method provided by the above-mentioned various methods. The method includes: obtaining target reimbursement data input by a target user; inputting the target reimbursement data into a semantic parsing model to obtain fill-in form intention data output by the semantic parsing model; the semantic parsing model is trained by sample reimbursement data and its corresponding fill-in form intention label results; matching the fill-in form intention data with a form template library to obtain a target reimbursement template that matches the fill-in form intention data; performing a data filling operation on the target reimbursement template based on the fill-in form intention data to obtain a target reimbursement form; auditing the target reimbursement form and performing financial reporting based on the audit confirmation result.
[0106] In another aspect, the present invention also provides a non-transitory readable storage medium, on which a program is stored. When the program is executed by a processor, it is configured to execute a financial reimbursement method based on artificial intelligence provided above. The method includes: obtaining target reimbursement data input by a target user; inputting the target reimbursement data into a semantic parsing model to obtain fill-in form intention data output by the semantic parsing model; the semantic parsing model is trained by sample reimbursement data and its corresponding fill-in form intention label results; matching the fill-in form intention data with a form template library to obtain a target reimbursement template that matches the fill-in form intention data; performing a data filling operation on the target reimbursement template based on the fill-in form intention data to obtain a target reimbursement form; auditing the target reimbursement form and performing financial reporting based on the audit confirmation result.
[0107] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.
[0108] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The software product can be stored in a readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0109] As described above, only the preferred specific embodiments of the present invention are provided, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. A financial reimbursement method based on artificial intelligence, characterized in that, It includes the following steps: Obtain target reimbursement data input by a target user; Input the target reimbursement data into a semantic parsing model to obtain fill-in form intention data output by the semantic parsing model; The semantic parsing model is trained by sample reimbursement data and its corresponding fill-in form intention label results; Match the fill-in form intention data with a form template library to obtain a target reimbursement template matched by the fill-in form intention data; Perform data filling operations on the target reimbursement template based on the fill-in form intention data to obtain a target reimbursement form; Audit the target reimbursement form and perform financial reporting based on the audit confirmation result; The matching of the fill-in form intention data with the form template library to obtain a target reimbursement template matched by the fill-in form intention data includes: Extract key features from the fill-in form intention data and the form template library to obtain a first fill-in form feature vector and a form feature vector respectively; For each historical reimbursement record, extract the fill-in form feature vector corresponding to the first fill-in form feature vector, and perform aggregation analysis on the fill-in form feature vectors to obtain clustering categories; each clustering category represents a historical reimbursement characteristic; Obtain historical reimbursement feature vectors based on the first fill-in form feature vector and the form feature vector, and calculate the historical usage frequency of each template in the clustering category in the form template library; the historical usage frequency of the template is obtained by dividing the number of occurrences of each form template in the clustering category by the total number of reimbursement records in the clustering category; Calculate the original similarity between the first fill-in form feature vector and the form feature vector using cosine similarity; Weight the original similarity based on the historical usage frequency of the template to obtain a weighted similarity; Cluster the templates in the form template library based on the features of the templates to form multiple clustering sets; Determine the clustering to be selected in the multiple clustering sets based on the weighted similarity; the clustering to be selected is the first preset number of clusters in the multiple clustering sets; Obtain the overall similarity between the second fill-in form feature vector in the clustering to be selected and each template, and determine the target reimbursement template based on the overall similarity; the second fill-in form feature vector is a vector obtained by dimensionality reduction of the first fill-in form feature vector; the overall similarity includes feature vector similarity, business rule similarity, and text semantic similarity; The calculation formula for the weighted similarity is: ; ; Among them, represents the weighted similarity; represents the raw similarity between represents the first filling form feature vector; represents the form feature vector; represents the dot product operation of represents the vector modulus length of represents the vector modulus length of represents the raw similarity calculated under the j th historical reimbursement characteristic, R represents the number of categories of historical reimbursement characteristics, represents the j th template historical usage frequency under the historical reimbursement characteristic.
2. The financial reimbursement method based on artificial intelligence according to claim 1, wherein The semantic parsing model is obtained by fusing a first target model, a second target model, and a third target model.
3. The financial reimbursement method based on artificial intelligence according to claim 2, wherein The training steps of the first target model include: Obtain first sample data; the first sample data includes structured data in historical financial reimbursement data; Input the first sample data into the first pre-trained model to obtain a first prediction result output by the first pre-trained model; obtain a first loss value according to the first prediction result based on a first loss function of the first pre-trained model; adjust a first optimization function in the first pre-trained model based on the first loss value; predict a prediction result of the first sample data based on the adjusted first pre-trained model until the loss value output by the first loss function is less than or equal to a first preset value, the square of the difference between adjacent loss values is less than 1, and the difference between at least three consecutive adjacent loss values is greater than zero, so as to obtain the first target model.
4. The financial reimbursement method based on artificial intelligence according to claim 2, wherein The training steps of the second target model include: Obtain second sample data; the second sample data includes semi-structured data in historical financial reimbursement data; Input the second sample data into the second pre-trained model to obtain a second prediction result output by the second pre-trained model; obtain a second loss value according to the second prediction result based on a second loss function of the second pre-trained model; adjust a second optimization function in the second pre-trained model based on the second loss value; predict a prediction result of the second sample data based on the adjusted second pre-trained model until the loss value output by the second loss function is less than or equal to a second preset value and the product of adjacent loss values is greater than 0 and less than 1, so as to obtain the second target model.
5. The method for financial reimbursement based on artificial intelligence according to claim 2, wherein The training steps of the third target model include: Obtain third sample data; the third sample data includes unstructured data in historical financial reimbursement data; Input the third sample data into the third pre-trained model to obtain a third prediction result output by the third pre-trained model; obtain a third loss value according to the third prediction result based on a third loss function of the third pre-trained model; adjust a third optimization function in the third pre-trained model based on the third loss value; predict a prediction result of the third sample data based on the adjusted third pre-trained model until the loss value output by the third loss function is less than or equal to a third preset value and the reciprocals of the products of alternate loss values are all greater than a fourth preset value, so as to obtain the third target model.
6. A financial reimbursement device based on artificial intelligence, characterized in that, Applied to the financial reimbursement method based on artificial intelligence according to any one of claims 1 to 5; The financial reimbursement device based on artificial intelligence includes: An acquisition unit, configured to acquire target reimbursement data input by a target user; A model output unit, configured to input the target reimbursement data into a semantic parsing model to obtain fill-in intention data output by the semantic parsing model; the semantic parsing model is trained by sample reimbursement data and its corresponding fill-in intention label results; A matching unit, configured to match the fill-in intention data with a form template library to obtain a target reimbursement template matched by the fill-in intention data; A writing unit, configured to perform data filling operations on the target reimbursement template based on the fill-in intention data to obtain a target reimbursement form; An uploading unit: configured to review the target reimbursement form and perform financial reporting based on the review confirmation result.
7. An electronic device, comprising a memory, a processor, and a program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, the steps of a financial reimbursement method based on artificial intelligence as described in any one of claims 1 to 5 are implemented.
8. A non-transitory readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, the steps of a financial reimbursement method based on artificial intelligence as described in any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Data processing method based on information identification technology and related device
CN110264288A
Examination report editing method, device and equipment
CN114841136A
Multi-modal model-based reimbursement process optimization method and device
CN119515573A