Active reimbursement processing method and system based on semantic graph and electronic equipment
By deploying connectors in the enterprise expense reimbursement system for multimodal parsing and structured extraction, and constructing a semantic graph, the problems of low reimbursement processing efficiency and poor data quality are solved, and an intelligent pre-approval process is realized, improving the efficiency and compliance of the expense reimbursement system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 北京合思信息技术有限公司
- Filing Date
- 2026-02-02
- Publication Date
- 2026-05-15
Smart Images

Figure CN122048548A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer information technology, and more specifically, to a proactive expense reimbursement processing method, system, and electronic device based on semantic graphs. Background Technology
[0002] In modern enterprise operations, expense management, as a crucial component of financial management, involves various expenditure types such as travel, entertainment, transportation, and procurement, characterized by high frequency, wide sources, and diverse vouchers. With increasing informatization, enterprises commonly adopt electronic expense control systems to replace traditional paper-based reimbursement processes, achieving digital processing of reimbursement applications, approval workflows, and financial entries. These systems typically rely on users actively uploading invoices, filling out forms, and submitting them for approval, with the system then performing rule verification and compliance reviews. Although some systems have introduced Optical Character Recognition (OCR) technology to assist in field extraction and combined it with simple logical judgments to achieve preliminary automation, the overall process still revolves around "post-event reporting." The system's intervention mainly occurs after the consumption behavior is completed, making it difficult to effectively capture the business context at the time of the consumption.
[0003] In recent years, the development of artificial intelligence and big data technologies has driven the advancement of smart office applications. Some systems have attempted to improve the accuracy of invoice recognition or optimize approval processes through machine learning models. Simultaneously, some platforms have begun to integrate external data sources, such as hotel booking interfaces or travel service provider APIs, in an effort to enrich the channels for obtaining expense information. However, these improvements largely remain at the level of localized functional enhancements and have not fundamentally changed the reimbursement model, which is still primarily driven by human intervention. Data collection remains fragmented, with a lack of semantic connections between vouchers from different sources. The system's ability to understand complex business scenarios is limited, and it cannot automatically reconstruct the complete chain of consumption events. Furthermore, compliance controls are mostly concentrated in the approval process, leading to delays in problem detection and resulting in repeated rejections and resubmissions, impacting efficiency and user experience. Summary of the Invention
[0004] The purpose of this invention is to provide a proactive reimbursement processing method, system, and electronic device based on semantic graphs, so as to alleviate the technical problems of low reimbursement processing efficiency, poor data quality, and high compliance risks in the prior art.
[0005] In a first aspect, embodiments of the present invention provide a proactive expense reimbursement processing method based on semantic graphs. The method includes: performing multimodal semantic parsing and structured extraction on acquired original vouchers related to expense time to generate standardized voucher metadata; proactively acquiring the original vouchers through connectors deployed in a multi-source system; constructing a semantic graph representing the same business activity based on the temporal, spatial, and business semantic features among the voucher metadata; performing compliance pre-review checks on the associated voucher metadata based on the semantic graph and generating pre-review results; generating a structured expense reimbursement draft based on the pre-review results; submitting the expense reimbursement draft for approval according to preset rules; and synchronizing it to the financial system after approval.
[0006] In some optional implementations, multimodal semantic parsing and structured extraction are performed on the original voucher, including: performing optical character recognition and layout analysis on the original voucher, and using a large language model for semantic understanding and named entity recognition to extract key fields; based on the above key fields, a dual-track path of template matching and model prediction is used to generate field extraction results, and the results of the two outputs are fused based on confidence to generate standardized voucher metadata; the above standardized voucher metadata includes time information, spatial information and business semantic information.
[0007] In some optional implementations, a semantic graph representing the same business activity is constructed, including: determining the association features between the original vouchers corresponding to multiple voucher metadata based on the time information, spatial information and business semantic information of each of the above voucher metadata; the association features include at least one of time proximity, geographical proximity, overlap of participating entities and business keyword matching degree; and clustering the scattered original vouchers through graph matching and community discovery algorithms to generate a semantic graph representing the same business activity.
[0008] In some alternative implementations, the graph matching and community discovery algorithms described above are guided by itinerary plans or travel application information as priors and employ a spatiotemporally constrained community discovery model for node clustering.
[0009] In some optional implementations, compliance pre-audit checks are performed on the associated voucher metadata based on the semantic graph, including: determining the single voucher dimension and the combined scenario dimension composed of multiple associated vouchers based on the semantic graph; performing parallel verification on the single voucher dimension and the combined scenario dimension according to a preset compliance strategy; and using an ensemble learning risk control model to perform anomaly detection on the voucher metadata and generate a pre-audit result containing risk level and rectification suggestions.
[0010] In some optional implementations, the above method further includes: when the above compliance pre-audit detects missing information or compliance anomalies, triggering proactive human-computer interaction to complete or confirm the information; the above proactive human-computer interaction includes: based on a low-interference strategy, aggregating multiple items to be confirmed and proactively reaching the user through a message interface; wherein, the above low-interference strategy includes dynamically selecting the timing of reaching the user based on the user's status, and prioritizing and throttling the interaction requests.
[0011] In some optional implementations, the above method also includes: using user confirmation during the interaction process, approval conclusions of the draft expense report, and audit results from the financial system as feedback signals to dynamically update the model parameters or decision rules used in at least one of the above-mentioned multimodal semantic parsing, semantic graph construction, compliance pre-audit detection, and draft expense report generation.
[0012] Secondly, embodiments of the present invention provide a proactive expense reimbursement processing system based on semantic graphs. This system includes: a semantic parsing and structured extraction module, used to perform multimodal semantic parsing and structured extraction on acquired original vouchers related to expense time, generating standardized voucher metadata; the aforementioned original vouchers are proactively collected through connectors deployed in a multi-source system; a semantic graph construction module, used to construct a semantic graph representing the same business activity based on the temporal, spatial, and business semantic features between the aforementioned voucher metadata; a pre-review detection module, used to perform compliance pre-review detection on the associated voucher metadata based on the aforementioned semantic graph, and generate pre-review results; and an approval module, used to generate a structured draft expense reimbursement form based on the aforementioned pre-review results, submit the draft expense reimbursement form for approval according to preset rules, and synchronize it to the financial system after approval.
[0013] Thirdly, embodiments of the present invention provide an electronic device, including a memory and a processor, wherein the memory stores a computer program that can run on the processor, and the processor executes the computer program to implement the steps of the method described in any of the first aspects above.
[0014] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing computer-executable instructions, which, when invoked and executed by a processor, cause the processor to perform the method described in any of the first aspects above.
[0015] This invention provides a proactive expense reimbursement processing method, system, and electronic device based on semantic graphs. The method first proactively collects original vouchers through connectors deployed in a multi-source system, performs multimodal parsing and structured extraction, and generates standardized voucher metadata. Then, based on the temporal, spatial, and business semantic features between metadata, a semantic graph representing the same business activity is constructed, enabling intelligent association and scenario reconstruction of dispersed consumption data. Before the expense report is generated, compliance pre-review and risk scoring are performed based on this semantic graph to achieve early problem detection. Finally, a structured expense report draft is automatically generated based on the pre-review results. This invention reshapes the traditional passive, fragmented, and post-audit expense reimbursement process into a proactive, associative, and pre-audit intelligent process, effectively solving the technical problems of low expense processing efficiency, poor data quality, and high compliance risks. It significantly reduces the burden of data entry for users and improves financial audit efficiency and risk control. Attached Figure Description
[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments of the present invention will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a flowchart illustrating a proactive reimbursement processing method based on semantic graphs, provided in an embodiment of the present invention. Figure 2 A schematic diagram of the structure of a proactive reimbursement processing system based on semantic graphs provided in an embodiment of the present invention; Figure 3 A flowchart illustrating another proactive reimbursement processing method based on semantic graphs provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] In modern enterprise management, expenses (including travel expenses, business entertainment expenses, transportation expenses, and procurement expenses) are characterized by high frequency and wide distribution. The sources of these expense vouchers are diverse, and their formats vary. Furthermore, there is a common time lag and process disconnect between business activities and reimbursement procedures. Traditional expense control systems primarily handle "post-event reporting and approval," collecting information only when users actively submit it. This passive management model easily leads to scattered original vouchers, incomplete financial information, and difficulty in effectively preventing duplicate reimbursements and compliance risks.
[0020] Based on this, the present invention provides a proactive reimbursement processing method, system and electronic device based on semantic graphs, which solves the technical problems of low reimbursement processing efficiency, poor data quality and high compliance risk in the prior art.
[0021] To facilitate understanding of this embodiment, a proactive reimbursement processing method based on semantic graphs, as disclosed in this embodiment of the invention, will first be described in detail. (See [link to relevant documentation]). Figure 1 The diagram shows a flowchart of a proactive reimbursement processing method based on semantic graphs. This method can be executed by an electronic device and mainly includes the following steps S102 to S108: Step S102: Perform multimodal semantic parsing and structured extraction on the acquired original vouchers related to cost and time to generate standardized voucher metadata; the original vouchers are actively collected through connectors deployed in the multi-source system.
[0022] Among them, the connector of the multi-source system can refer to a lightweight data acquisition agent deployed on different data source sides. Preferably, it can be a lightweight data acquisition agent deployed on different data source sides such as enterprise email system, payment billing system, travel / hotel / car service provider open platform, instant messaging tool, calendar and travel application system.
[0023] Preferably, the connector, under user authorization and corporate compliance constraints, can continuously monitor various data sources and automatically trigger data collection when it detects an event signal related to expenses. This event signal may include email subjects containing keywords such as "flight ticket," "hotel," or "e-invoice," overlapping bill transaction times with travel itinerary time windows, IM messages containing invoice images and contextual descriptions such as "client dinner," and calendar events marked "business trip" or "meeting." This embodiment allows voucher collection to be synchronized with the business process, avoiding voucher omissions and delays caused by reliance on user recollection and manual uploading in traditional models, thus alleviating the problem of low reimbursement processing efficiency.
[0024] Multimodal semantic parsing can include optical character recognition and layout analysis, large language model semantic understanding and named entity recognition, and a dual-track path fusion of template matching and model prediction.
[0025] In one embodiment, multimodal semantic parsing and structured extraction of the original voucher may include: performing optical character recognition and layout analysis on the original voucher, and using a large language model for semantic understanding and named entity recognition to extract key fields; based on the key fields, generating field extraction results using a dual-track path of template matching and model prediction, and fusing the results of the two outputs based on confidence levels to generate standardized voucher metadata; the standardized voucher metadata includes time information, spatial information, and business semantic information. Specifically, the aforementioned optical character recognition and layout analysis can be used to locate text regions, table structures, and QR code / barcode positions in document images or PDFs; the aforementioned large language model can perform semantic understanding and named entity recognition to identify business entities such as merchants, amounts, tax rates, project / cost centers, participants, and travel time and location from unstructured text; the aforementioned template matching is suitable for high-frequency standard format vouchers, and the aforementioned model prediction is suitable for long-tail heterogeneous format vouchers; the outputs of the two are fused by confidence weighting to generate standardized voucher metadata containing time information, spatial information, and business semantic information.
[0026] This embodiment realizes automated, high-precision, context-aware structured conversion of multi-source, multi-modal original vouchers, and solves the problem of inconsistent format standards caused by diverse voucher sources and fragmented data collection, thereby providing standardized input for subsequent semantic graph construction.
[0027] Step S104: Based on the temporal, spatial, and business semantic features among the credential metadata, construct a semantic graph representing the same business activity.
[0028] Specifically, the temporal, spatial, and business semantic features can include temporal proximity, geographical proximity, overlap of participating entities, and matching degree of business keywords; the graph matching and community discovery algorithm can use itinerary plans or travel application information as prior guidance, and adopt a community discovery model based on spatiotemporal constraints for node clustering.
[0029] In one embodiment, constructing a semantic graph representing the same business activity may include: determining the association features between the original vouchers corresponding to multiple voucher metadata based on the temporal information, spatial information, and business semantic information of each voucher metadata; the association features include at least one of temporal proximity, geographical proximity, overlap of participating entities, and business keyword matching degree; and clustering the scattered original vouchers through graph matching and community discovery algorithms to generate a semantic graph representing the same business activity. Preferably, the temporal proximity can be obtained by calculating the timestamp difference between voucher metadata and mapping it to Gaussian kernel function similarity; the geographical proximity can be obtained by calculating the spherical distance between geographical coordinates using the Haversine distance formula and mapping it to exponential decay similarity; the overlap of participating entities can be obtained by comparing the intersection ratio of personnel identifiers, department codes, or project numbers extracted from voucher metadata; and the business keyword matching degree can be obtained by calculating the cosine similarity of the voucher text embedding vector using the BERT model.
[0030] The graph matching and community discovery algorithm described above uses multidimensional similarity as edge weights to construct a directed relationship graph between credential metadata nodes, and runs an improved Louvain algorithm for node clustering. This embodiment supports the stable identification of credential sets belonging to the same business activity even under noise conditions such as time misalignment, geographical offset, or subject ambiguity, thereby achieving intelligent association and scene reconstruction of scattered consumption data.
[0031] In one embodiment, the graph matching and community discovery algorithm described above uses itinerary plans or travel application information as prior guidance and employs a spatiotemporal constraint-based community discovery model for node clustering.
[0032] Preferably, itinerary plans or travel application information can serve as anchor nodes in the semantic graph, with their time windows, destinations, and participant lists constituting strong constraints. The spatiotemporal constraint-based community discovery model can use the time overlap rate, geographical radius coverage, and participant matching degree between credential metadata nodes and anchor nodes as hard filtering thresholds during the clustering process, allowing only nodes that meet the constraints to enter the candidate cluster. This embodiment can effectively suppress erroneous cross-trip associations, improve the business accuracy of the semantic graph, and solve the problem of limited understanding of complex business scenarios in existing technologies.
[0033] The above embodiments can intelligently cluster and associate original vouchers collected from different systems and at different times according to actual business activities, so as to realize the intelligent association and scene restoration of scattered consumption data, thereby solving the problems of lack of semantic association between vouchers from different sources and the inability to automatically restore the complete consumption event chain.
[0034] Step S106: Perform compliance pre-audit checks on the associated credential metadata based on the semantic graph, and generate pre-audit results.
[0035] In one embodiment, performing compliance pre-audit checks on associated voucher metadata based on semantic graphs may include: determining the single voucher dimension and the combined scenario dimension composed of multiple associated vouchers based on the semantic graph; performing parallel verification in the single voucher dimension and the combined scenario dimension according to a preset compliance strategy; and using an ensemble learning risk control model to perform anomaly detection on the voucher metadata and generate a pre-audit result containing risk level and rectification suggestions.
[0036] Verification at the single voucher level can include checking the authenticity of invoices, verifying the consistency of arithmetic amounts, determining the validity of invoice dates, and matching merchants to blacklists and whitelists. Verification at the combined scenario level can include verifying the integrity of travel itineraries (such as the time logic closure of air tickets, hotel tickets, and transportation vouchers), analyzing the rationality of geographical trajectories (such as the matching degree between travel time between cities and modes of transportation), checking the compliance of budget allocation, and verifying the applicable rules for cross-currency exchange rates.
[0037] The aforementioned pre-defined compliance strategies can be defined by the company's configurable expense policies, tax compliance requirements, and historical audit experience. Examples include: spending limits, duplicate expense detection, invoice compliance verification, timeliness assessment, merchant blacklist / whitelist matching, and travel standard verification.
[0038] Ensemble learning risk control models can combine multiple algorithms (such as isolated forest, local anomaly factor, autoencoder, etc.) to generate preliminary review results with risk levels and rectification suggestions.
[0039] The purpose of this step is to complete pre-compliance verification and early risk detection before the expense report is generated, that is, to move compliance review from the approval stage to before the expense report is generated, in order to solve the problem that compliance control in the existing technology is mostly concentrated in the approval stage, resulting in delayed problem detection and high rejection rate. In one embodiment, the above method may further include: when the compliance pre-review detects missing information or compliance anomalies, triggering proactive human-computer interaction to complete or confirm the information; proactive human-computer interaction includes: based on a low-interference strategy, aggregating multiple items to be confirmed and proactively contacting the user through a message interface; wherein, the low-interference strategy includes dynamically selecting the contact timing according to the user's status and prioritizing and throttling the interaction requests.
[0040] Proactive human-computer interaction can include aggregating multiple items to be confirmed based on a low-interference strategy and proactively reaching users through a message interface. The low-interference strategy can refer to dynamically selecting the timing of reaching users based on their status and prioritizing and throttling interaction requests.
[0041] Preferably, the above-mentioned aggregation of multiple pending confirmation items refers to merging the missing items, conflicting fields, and strategy suggestions that need to be confirmed under the same business activity into one interaction card; the above-mentioned message interfaces include enterprise instant messaging tools, email systems, and mobile notifications; the above-mentioned user status includes online / offline, working hours / non-working hours, in a meeting / idle; the above-mentioned priority management is set based on the problem risk level, timeliness, and user's historical response habits; the above-mentioned throttling control limits the number of interaction requests pushed to the same user per unit time.
[0042] This embodiment reduces the frequency of user operations and cognitive load, significantly alleviating the burden of data entry for users.
[0043] Step S108: Generate a structured draft expense report based on the preliminary review results, submit the draft expense report for approval according to preset rules, and synchronize it to the financial system after approval.
[0044] As another example, the aforementioned preliminary review results can also be structured outputs generated based on semantic graphs, including risk levels, references to violation clauses, evidence chain indexes, and rectification suggestions. Furthermore, the draft reimbursement form can be a standardized data object formed by organizing voucher metadata based on the preliminary review results, using subgraphs of the semantic graph as units, and completing expense classification, tax separation, cross-currency exchange rate conversion, apportionment calculation, and bill-invoice alignment. Preferably, the aforementioned preset rules can refer to an approval matrix composed of amount thresholds, expense categories, affiliated departments, applicant ranks, regional policies, and approval node permissions.
[0045] Step S108 above can eliminate the need for manual piecing together of expense reports, enabling the automatic generation of structured expense report drafts based on the pre-approval results, which can significantly reduce the burden of filling out forms for users.
[0046] In one embodiment, the above method may further include: using the user confirmation operation during the interaction process, the approval conclusion of the draft expense report, and the audit result of the financial system as feedback signals to dynamically update the model parameters or decision rules used in at least one of the steps of multimodal semantic parsing, semantic graph construction, compliance pre-audit detection, and generation of the draft expense report.
[0047] As specific examples, user confirmation can refer to a user's one-click confirmation, field fine-tuning, or natural language supplementary input on an interactive card pushed by the system; the approval conclusion of a draft expense report can refer to the approval system's response of approval, rejection, return for modification, or additional comments; the audit results of the financial system can refer to the reasons for order cancellation, reconciliation discrepancies, or image missing prompts fed back by the financial accounting system or ERP system during the accounting process; dynamic updates can refer to writing the above three types of signals into feature storage, triggering model retraining or rule hot loading, with an update cycle typically on an hourly or daily basis.
[0048] This embodiment can establish a data-driven continuous optimization path, enabling each processing step to adapt to the evolution of enterprise rules and changes in business scenarios.
[0049] Preferably, the dynamic update specifically includes: using the field annotation data corrected by user confirmation operation to fine-tune the confidence fusion weight of the named entity recognition head and template matching engine of the large language model in the multimodal semantic parsing module; using the semantic graph subgraph corresponding to the rejected reimbursement draft in the approval conclusion to adjust the similarity calculation coefficient of time proximity and geographical proximity in the community discovery model; using the duplicate reimbursement or invoice occupancy anomalies fed back by the financial system audit to enhance the recognition ability of the integrated learning risk control model in the compliance pre-audit detection module for the multi-path reporting pattern of the same invoice number; this implementation method can gradually improve the accuracy of the system to form a unique reimbursement knowledge body for the enterprise.
[0050] Furthermore, the entire process of collecting and applying the feedback signals is traceable, with each signal associated with the original voucher fingerprint, semantic graph snapshot version, pre-review decision path, and update effective timestamp. Before official deployment, the updated model parameters or decision rules are verified using an A / B testing framework to assess their impact on the initial review pass rate, document integrity rate, and average reimbursement cycle. This mechanism ensures that technical improvements are quantifiable, auditable, and traceable, thereby achieving the technical effects of improving financial audit efficiency and risk control, reducing user reporting burden, and increasing reimbursement processing efficiency. This invention provides a proactive reimbursement processing method based on semantic graphs. First, original vouchers are proactively collected through connectors deployed in a multi-source system, and multimodal parsing and structured extraction are performed to generate standardized voucher metadata. Then, based on the temporal, spatial, and business semantic features between metadata, a semantic graph representing the same business activity is constructed, enabling intelligent association and scenario reconstruction of dispersed consumption data. Before the reimbursement form is generated, compliance pre-review and risk scoring are performed based on this semantic graph to achieve early problem detection. Finally, a structured reimbursement form draft is automatically generated based on the pre-review results. This invention reshapes the traditional passive, fragmented, and post-audit reimbursement process into an intelligent process that is proactive, interconnected, and pre-audit. It effectively solves the technical problems of low reimbursement processing efficiency, poor data quality, and high compliance risks, significantly reduces the burden of data entry for users, and improves the efficiency of financial auditing and risk control.
[0051] Based on the same inventive concept, this invention also provides a proactive reimbursement processing system based on semantic graphs, see [link to relevant documentation]. Figure 2 As shown, the system mainly includes the following parts: The semantic parsing and structured extraction module 210 is used to perform multimodal semantic parsing and structured extraction on the acquired original vouchers related to cost and time, and generate standardized voucher metadata; the original vouchers are actively collected through connectors deployed in the multi-source system; Semantic graph construction module 220 is used to construct a semantic graph representing the same business activity based on the temporal, spatial and business semantic features between credential metadata. The pre-audit detection module 230 is used to perform compliance pre-audit detection on the associated credential metadata based on semantic graphs and generate pre-audit results; The approval module 240 is used to generate a structured draft expense report based on the pre-approval results, submit the draft expense report for approval according to preset rules, and synchronize it to the financial system after approval.
[0052] The semantic graph-based proactive reimbursement processing system provided in this embodiment of the invention can be specific hardware on a device or software or firmware installed on the device. The system provided in this embodiment of the invention has the same implementation principle and technical effects as the aforementioned method embodiments. For the sake of brevity, any parts not mentioned in the system embodiments can be referred to the corresponding content in the aforementioned method embodiments. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can all be referred to the corresponding processes in the aforementioned method embodiments, and will not be repeated here.
[0053] To facilitate understanding, this invention also provides an application example of a proactive expense reimbursement processing method based on semantic graphs. Preferably, this embodiment proposes a fully intelligent expense reimbursement method and system, breaking through the traditional "passive form filling" efficiency tool positioning and realizing proactive, continuous, and closed-loop intelligent services from expense occurrence to reimbursement recording. Its core lies in transforming the role of AI from a "form-filling clerk at the final destination" to a "personal assistant throughout the journey," silently completing voucher collection, parsing, association, pre-approval, form grouping, and synchronized recording in the background. Employees only need to "confirm" key issues with a single click, and finance can conduct efficient and interpretable approval and auditing in combined scenarios. See also... Figure 3 As shown, preferably, the overall process of this method includes the following steps S301 to S308: Step S301: Actively collect credentials from all channels.
[0054] Under user authorization and corporate compliance constraints, the system continuously monitors and captures original vouchers and contextual data related to expenses through connectors such as email, billing, service provider APIs, IM, and calendar / travel systems. These data include itineraries, e-invoices, order receipts, credit card statements, business trip applications, and meeting minutes, forming a continuous record of "before, during, and after" expenses.
[0055] Preferably, under the constraints of user authorization and enterprise compliance, the system deploys multiple types of connectors and data collection agents to form "continuous data collection." The email connector subscribes to highly relevant topics (flight tickets, hotels, ride-hailing, conference registration, e-invoices, etc.) through OAuth 2.0 / enterprise IMAP agents and uses both templates and models to identify sources and content, automatically locating key credentials from sender domains, subject semantics, attachment MIME types, and body structure. The billing connector connects to corporate / personal card (controlled scenarios) clearing and electronic statements, capturing payment events based on transaction time, merchant MCC, amount, currency, and geographic location, and establishing weak connections with associated credentials. Service provider connectors for travel / hotels / car services synchronize order receipts and change notifications through open APIs or enterprise aggregation platforms; IM / collaboration platform data collection receives invoice images and contextual descriptions (e.g., "customer dinner") via an enterprise whitelist robot, performs anonymization and quality checks on the client side, and uploads minimized versions; the calendar and travel application system provides "planned" information (destination, time window, budget, participants). All raw events enter the event bus and are stored in the "raw credential repository," accompanied by information such as source, collection time, authorized scope, integrity assertion, fingerprint, and playback identifier. Deduplication is achieved using a hash fingerprint and source ID combination key, and time-series verification prevents duplicate entries and dirty data. To adapt to different enterprise network boundaries and compliance requirements, the collection agent supports edge deployment, zero-trust access, and granular data logging policies, achieving a controllable, manageable, and traceable multi-source collection closed loop.
[0056] Step S302, semantic parsing and structured extraction.
[0057] Perform OCR and layout analysis on multimodal data (text / image / PDF / email / HTML), and extract fields such as merchant, amount, tax rate, project / cost center, participants, trip time and location, and transaction number by combining Large Language Model (LLM) and Named Entity Recognition (NER). Then perform authenticity verification and consistency verification to generate standardized voucher metadata.
[0058] Preferably, the system employs cascaded parsing for multimodal vouchers: First, it performs layout analysis and OCR on the PDF / image to identify key blocks, table structures, QR codes / barcodes, and other elements on the voucher and constructs a layout structure tree; then, it uses domain-adjusted LLM to perform semantic segmentation, intent recognition, and named entity recognition (NER), and combines layout features to complete field and relation extraction, forming a set of multiple candidate fields with confidence scores.
[0059] Perform arithmetic consistency checks (subtotal / tax / total) and time window and exchange rate window checks on numerical fields such as amount, tax rate, and currency; call the authenticity verification channel to check the header, tax number, invoice date, unique invoice number and the risk of duplicate reimbursement (in conjunction with the invoice pool) for invoice-type vouchers.
[0060] Image quality assessment and tamper detection (such as secondary compression and splicing traces) are introduced for images, and abnormal samples are entered into a manual confirmation queue.
[0061] To achieve a balance between stability and generalization, the system adopts a dual-track approach of "template + model": high-frequency templates are extracted quickly using lightweight templates, while long-tail templates are generalized using models. The two outputs are merged by a conflict resolver based on confidence, consistency, and source credibility, and finally standardized into a unified "expense voucher schema" for the enterprise. The image fingerprint, parsing version number, rule hit trajectory, and evidence chain are archived together to ensure traceability and reproducibility.
[0062] Multimodal recognition architecture: The system constructs a CNN+Transformer hybrid network, using an EfficientNet-B7 image encoder as the backbone for document layout analysis, and combining it with the LayoutLM-v3 model to achieve document structure understanding. Image preprocessing employs adaptive Gaussian filtering (σ=1.2) and Canny edge detection (threshold 50-150). Text regions are located through the EAST network, and end-to-end character recognition is achieved using the CRNN+CTC loss function, with an accuracy exceeding 95%.
[0063] Page layout analysis algorithm: Based on Mask R-CNN for page element segmentation, it identifies key blocks on the ticket (header, amount, date, merchant, tax number, etc.), table structure (row and column boundary detection), QR code / barcode (ZXing decoding), and other elements. TreeLSTM is used to encode spatial relationships when constructing the page layout structure tree, supporting hierarchical representation of complex layouts.
[0064] Semantic understanding engine: A domain-fine-tuned BERT-Base-Chinese model is used for semantic segmentation, and BiLSTM+CRF is employed for named entity recognition (NER), achieving an F1 score of 92%. Entity types include 18 financial entities such as ORG (merchant), MONEY (amount), DATE (date), LOC (address), and TAX (tax number). Relation extraction uses the RE-BERT model to identify semantic relationships such as "invoice issuer-invoice recipient" and "product-price".
[0065] Confidence calculation: A weighted confidence assessment based on model output probability, layout consistency, and historical accuracy. The formula is: Confidence = 0.4×P_model + 0.3×P_layout + 0.3×P_history, where P_model is the model prediction probability, P_layout is the layout feature matching degree, and P_history is the historical validation accuracy.
[0066] Data verification mechanism: The amount field uses the regular expression pattern="\\d{1,3}(,\\d{3})". Matching (\.\\d{2})?", performing arithmetic consistency checks (subtotal + tax = total, error tolerance ±0.01). Time window checks use the sliding window algorithm to detect abnormal time jumps. Exchange rate checks obtain the benchmark exchange rate in real time through the central bank's API, with a tolerance of ±3%.
[0067] Authenticity verification process: Invoices are verified for authenticity through the State Taxation Administration's invoice verification interface, including verification of four elements: invoice code, number, invoice date, and invoice amount. A bitmap index for the invoice pool is established, and duplicate reimbursements are detected in O(1) time complexity.
[0068] Image quality control: Laplacian variance is used to detect blurriness (threshold 100), JPEG compression quality is assessed (threshold 85), and Error Level Analysis is used to detect tampering traces. Abnormal samples (quality score <60) are automatically entered into a manual review queue.
[0069] Conflict resolution algorithm: Employing the Dempster-Shafer fusion framework based on evidence theory, probabilistically fuses the outputs of template extraction and model prediction. The confidence weight calculation formula is as follows: W_template = 0.8×Match_score + 0.2×Template_coverage, W_model = Entropy_score×Generalization_factor.
[0070] Finally, softmax normalization is used to output a uniform confidence distribution.
[0071] Standardized Schema: Defines a unified JSON schema for enterprise expense vouchers, containing 47 standard fields and 18 extended fields. Data types include string, number, date, geopoint, array, etc., supporting standardized representations across multiple currencies, languages, and regions. Versioning is employed during storage, preserving parsing traces, rule hit paths, and evidence chain pointers, supporting audit backtracking and performance tuning.
[0072] Step S303: Semantic association and process / event graph construction.
[0073] Based on features such as time, location, participants, application form, budget / subject, etc., the system performs spatiotemporal and business semantic associations on multi-source vouchers across invoices and channels, constructs a directed graph of "trip-consumption-application-approval", automatically identifies multiple expenses under the same business trip / event, resolves ambiguities and completes the context.
[0074] After the vouchers are structured, the system constructs a directed multi-relationship graph with time, geography and business semantics as the axes: nodes include “itinerary plan”, “actual travel segment”, “consumption voucher”, “application form”, “meeting / visit event”, “participant / department / project”, etc., and edge relationships represent “belonging to the same itinerary”, “triggered by this event”, “paid for by this event”, “approved by this event”, etc.
[0075] By leveraging multimodal features such as spatiotemporal proximity (same day or reasonable time window, same city or geographical radius), overlapping participants, consistent destinations, matching schedule keywords (e.g., client name / project name / exhibition name), application form reference relationships, and similarity between the merchant MCC on the bill and the purpose of the trip, a graph matching and community discovery model is trained. This model automatically assigns scattered vouchers to the corresponding event clusters and provides candidate associations and confidence rankings for ambiguous or conflicting scenarios (e.g., multiple meals in the same city), allowing for confirmation to be delayed until the interaction stage.
[0076] The graph, acting as a "semantic platform," provides a combined perspective for pre-screening (completeness of flight + hotel + vehicle bookings), natural boundaries for group bookings (packaging by event or itinerary), and a visualized reconstruction of "business storylines" for auditing. The system incrementally updates the graph, maintains versions and snapshots to support approval reruns and audit evidence collection, and provides indicator alignment and backtracking capabilities to ensure consistency and transparency in governance.
[0077] Knowledge Graph Architecture Design: A Property Graph model is adopted to construct a multi-level graph structure containing eight main node types: Person, Travel, Event, Expense, Merchant, Location, Project, and Policy. Twenty-three relationship types are defined between nodes, such as BELONGS_TO, HAPPENS_AT, PAYS_FOR, and APPROVED_BY. Each edge carries attributes such as timestamp, confidence level, and evidence chain.
[0078] Entity extraction and linking: A named entity recognition model based on BiLSTM+CRF, combining rule matching and dictionary lookup, achieves an F1 score of 94%. Entity linking employs a fuzzy matching algorithm based on edit distance, combined with TF-IDF similarity calculation, using the following formula: Similarity = α×EditDistance_norm + β×TF-IDF_score + γ×Context_match, where α=0.4, β=0.4, γ=0.2.
[0079] Relation extraction algorithm: A Transformer-based relation classification model is adopted. The input format is [CLS]Entity1[SEP]Entity2[SEP]Context[SEP], and the output is a 23-dimensional relation probability distribution. The model architecture includes a 12-layer Transformer encoder, 768 hidden layer dimensions, and 12 attention heads. After fine-tuning on the domain dataset, the accuracy reaches 89%.
[0080] Spatiotemporal correlation algorithm: Design of a multi-dimensional similarity calculation framework: Time similarity: The Gaussian kernel function Temporal_sim(t1,t2) = exp(-|t1-t2|² / 2σ²) is used, where σ = 24 hours.
[0081] Spatial similarity: Spatial_sim(loc1,loc2) = exp(-distance(loc1,loc2) / threshold) is calculated based on Haversine distance, with the threshold set to 5 kilometers.
[0082] Semantic similarity: Text embeddings were calculated using BERT-Base, with a cosine similarity threshold of 0.75.
[0083] Graph Neural Network Inference: A 3-layer GraphSAGE network is constructed for multi-hop relationship inference, with each layer containing 256 hidden units. The mean aggregator is used to aggregate neighbor node features, and the activation function is ReLU. The loss function combines node classification loss and edge prediction loss: Loss = α×L_node + β×L_edge, where α=0.6, β=0.4.
[0084] Dynamic graph update mechanism: Implements an incremental update algorithm, with a time complexity of O(log n) for inserting a new node and O(1) for edge updates. Versioned snapshot storage is used, with each graph snapshot containing node / edge status, metadata version, and change logs. Timestamp-based snapshot backtracking and concurrent write control are supported.
[0085] Community discovery algorithm: An improved Louvain algorithm is used for event clustering, combined with modularity optimization and spatiotemporal constraints: Q = 1 / 2m × Σ[Aij - kikj / 2m]×δ(ci,cj)×w_temporal×w_spatial, where w_temporal and w_spatial are spatiotemporal weight factors.
[0086] Confidence propagation mechanism: Confidence propagation is based on the Label Propagation algorithm. The iterative formula is: conf(vi)^(t+1) = Σ wij×conf(vj)^(t) / Σ wij, where wij is the edge weight, reflecting the relationship strength and confidence. The convergence condition is that the change in confidence between two consecutive iterations is <0.01.
[0087] Graph query optimization: A Cypher-based query engine is built, supporting path queries, subgraph matching, and aggregation calculations. Graph indexes (B+ trees + adjacency tables) and query plan optimizations are employed, resulting in typical query response times of <100ms. Graph sharding storage is implemented, supporting horizontal scaling and load balancing.
[0088] Business rule integration: Enterprise expense policies are encoded into graph constraint rules, such as "the total cost of transportation and accommodation for the same business trip shall not exceed the standard × number of days". Complex constraints are expressed using the SWRL rule language, and consistency checks and violation detection are performed through the inference engine.
[0089] Step S304, compliance strategy pre-review and risk scoring.
[0090] Based on the company's configurable expense policies, tax compliance, and historical audit experience, combined with strategy engines and risk control models, we conduct preliminary reviews of single invoices and combined scenarios, covering issues such as limit limits, duplicate reimbursements, invoice compliance, timeliness, blacklists and whitelists, and travel standards, and generate explainable risk conclusions and rectification suggestions.
[0091] Specifically, the system's built-in strategy engine and risk control model work together: the strategy engine supports both graphical and DSL configurations, expressing limits, travel standards (star rating, city classification, subsidies), time windows, duplicate reimbursements, necessity of invoices (e.g., tickets must correspond to boarding / itinerary), merchant blacklists and whitelists, policy exceptions and authorization exemptions, etc.; the risk control model learns abnormal patterns (unconventional time / location, abnormal amount distribution, suspicious merchants, frequent order splitting, etc.) based on the company's historical rejections and audit samples.
[0092] The preliminary review operates in parallel at both the individual ticket and combined ticket levels: at the individual ticket level, it verifies authenticity, field consistency, and basic policies; at the combined ticket level, it checks the completeness of the itinerary, the rationality of cost allocation (tax rate / subsidy / currency), budget alignment, and the correctness of cost centers. Each hit provides an explanatory description, a chain of evidence, and actionable rectification suggestions (such as changing the project affiliation or merging it into a specific itinerary), and sets a blocking / warning / approval level. The preliminary review aims to complete automatic repairs or prepare "one-click confirmation" suggestions before employee intervention, reducing repeated communication and rejection cycles during the submission stage and improving the initial review pass rate.
[0093] Strategy engine architecture: Based on a hierarchical decision tree structure, supporting both graphical drag-and-drop configuration and DSL syntax dual modes. The RETE algorithm optimizes rule matching efficiency, with an average matching time of <5ms. Strategies are divided into four levels: Basic Compliance Layer (invoice authenticity, amount consistency), Corporate Policy Layer (travel standards, quota restrictions), Industry Standards Layer (tax compliance, audit requirements), and Personalized Strategy Layer (department-specific rules, project constraints).
[0094] Multidimensional risk feature engineering: Constructing a 168-dimensional risk feature vector, including: Time-related features (24 dimensions): Distribution of consumption periods, weekday / weekend ratio, frequency of late-night consumption, anomaly of time intervals, etc. Amount-related features (32 dimensions): deviation of amount distribution, consecutive integer amounts, abnormally large amounts, frequent small-amount splits, etc. Geographical features (28 dimensions): Dispersion of consumption locations, proportion of consumption in other locations, consistency of GPS trajectory, high-risk areas, etc. Behavioral features (36 dimensions): submission frequency, number of modifications, changes in merchant preferences, distribution of consumption categories, etc. Related Dimension Features (48 dimensions): Multi-account collaboration mode, team consumption correlation, supplier concentration, risk of duplicate invoices, etc.
[0095] Anomaly detection algorithm ensemble: Employing an ensemble learning framework that combines multiple algorithms: Isolation Forest: Detects global outliers, with a contamination parameter of 0.1 and 100 trees. LOF Local Anomaly Factor: Identifies local density anomalies, with neighbor count k=20 and anomaly threshold of 1.5; OCSVM single-class support vector machine: learns the boundary of normal patterns, kernel function RBF, gamma=0.1; Autoencoder anomaly detection: network structure [168-128-64-32-64-128-168], reconstruction error threshold is determined based on 95th percentile.
[0096] Risk scoring algorithm: Design of a multi-level scoring system: 1. Basic anomaly score: S_base = Σ wi×fi, where wi is the feature weight and fi is the feature value; 2. Pattern matching score: Calculation of pattern similarity based on historical cases; 3. Group Deviation Score: The degree of deviation relative to colleagues in the same department / job level; 4. Time-series trend score: The degree of trend anomaly based on time series analysis; 5. Final Risk Score: Risk_Score = α×S_base + β×S_pattern + γ×S_group +δ×S_trend; where α=0.4, β=0.3, γ=0.2, δ=0.1.
[0097] Dynamic threshold adjustment: The risk threshold is dynamically adjusted using sliding window statistics and Z-score standardization. Window size: 30 days of historical data; Threshold update frequency: automatically updated every day at midnight; Adjustment formula: Threshold_new = μ + k×σ, where k is dynamically adjusted according to the false alarm rate (target false alarm rate 5%).
[0098] Interpretability mechanism: Generate risk interpretations based on SHAP (SHapley Additive exPlanations) values. Feature importance ranking: Calculate the contribution of each feature to the final prediction; Decision path visualization: Displays the decision tree path and key nodes; Comparative analysis: Comparison with similar normal cases; Rectification suggestion generation: Automatically generates specific and actionable remediation suggestions based on decision-making rules.
[0099] Strategy Learning and Optimization: Implementing an Online Learning Mechanism Feedback collection: Record feedback signals such as approval results, user confirmations, and audit findings; Feature weight adjustment: Adjust feature weights using gradient descent based on feedback; Rule mining: Using the FP-Growth algorithm to mine new rules from the feedback data; A / B testing framework: Supports canary releases and performance evaluation of new strategies.
[0100] Real-time early warning system: Constructing a three-tiered early warning mechanism: Green Channel: Automatic approval if risk score <30 points; Yellow alert: 30-70 minutes, marked as a reminder but not blocked; Red Block: >70 points, mandatory manual review; The early warning response time is less than 100ms, supporting real-time decision-making and immediate feedback.
[0101] Step S305, Proactive human-computer interaction and missing information completion.
[0102] When missing or inconsistent items are detected, the system proactively reaches users with non-intrusive cards / messages, displaying context and suggested values. Users can complete the task with a single click to confirm or make minor adjustments. The system provides multi-round natural language interaction for complex scenarios and leaves a complete traceable record throughout the process.
[0103] Specifically, when the system detects missing fields, uncertain associations, or policy suggestions requiring confirmation, it proactively reaches users via enterprise IM, email cards, and mobile notifications, adhering to the principle of "non-intrusive and low-friction." Interactive cards aggregate multiple questions, displaying the system's understanding of the context (trip name, participants), suggested values and reasons, and remaining risks and impacts. Users can confirm or edit with a single click; complex questions support multi-round natural language input and are instantly structured and stored in the database.
[0104] The interactive orchestrator features prioritization and throttling controls, dynamically selecting the timing and channel for communication based on user work hours, device online status, and historical response behavior to avoid "message bombardment." It also supports proxy confirmation and role allocation (e.g., administrative / assistant confirmation of itinerary information), and strictly records the confirmer, time, and discrepancies to ensure a clear and auditable chain of responsibility. Through this mechanism, the traditional "end-of-month review and centralized organization" is broken down into "light confirmation of current or recent consumption," significantly reducing the probability of forgetting or omissions.
[0105] Step S306: Intelligent grouping and automatic merging of expense reports.
[0106] Based on the data map and preliminary review results, the system automatically merges relevant vouchers according to business events or travel itineraries, completes expense classification, tax separation, currency conversion, allocation rules and bill-invoice alignment, and generates an auditable and recordable reimbursement form draft.
[0107] Specifically, based on the graph and pre-approval results, the system triggers a group synthesizer to automatically generate a draft expense report. The synthesis logic includes: aggregating vouchers by event / itinerary; classifying them according to the company's accounting and expense account system; performing currency conversion across currencies based on transaction day or company-wide rules; separating tax amounts and allocating tax rates according to tax-inclusive / tax-exclusive rules; automatically merging small-amount transactions by the same merchant on the same day; aligning corporate credit card statements with invoices / receipts; and generating allocation details for transactions that may cross projects / cost centers based on participants, application forms, and allocation rules. The generated draft includes structured fields, compliance status, necessary attachments, image links, explanatory notes, and change history. The user interface follows a simplified "check-confirm" path, while generating a "scene view" (timeline / map) and compliance summary for the approver, reducing back-and-forth communication. Before submission, the system runs a lightweight consistency check again to ensure the final version is consistent with the voucher chain.
[0108] Step S307: Approval linkage and synchronization with the financial system.
[0109] Once the draft is confirmed, it is routed for approval according to the approval matrix; after approval, the review data and images are synchronized to the financial accounting / expense control / image archiving system, and the approval opinions and changes are written back to the map to support auditing and traceability; if rejected, the completion is triggered precisely.
[0110] Specifically, after a draft is submitted, the system automatically routes data according to the enterprise's approval matrix (amount threshold, subject / project, department, region), supporting serial / parallel processing and conditional branching. The approval interface provides a highly interpretable interface: the voucher chain, process context, rule hits, and rectification history for each expense are clearly visible; drill-down to the original image and email source is supported for suspicious points. After approval, the system pushes the reviewed data to the financial accounting / ERP / expense control system and synchronizes the image file and index metadata to the archive system, ensuring traceable association between image and entry; if rejection occurs, it automatically writes back to the graph and triggers an interactive card for precise completion, avoiding back-and-forth emails. The system also provides cross-period consistency verification, invoice holding and release mechanisms, and reproducible processing logs to meet the stringent requirements of taxation and auditing.
[0111] Step S308: Closed-loop learning and dynamic policy optimization.
[0112] By using user confirmation, approval conclusions, and audit results as feedback signals, the analysis model, strategy thresholds, and group order strategies are continuously optimized to form a unique "reimbursement knowledge body" for the enterprise, achieving increasing accuracy over time.
[0113] Specifically, the system writes user confirmations, approval / rejection status, audit conclusions, and discrepancies between financial reconciliation records as feedback signals into the feature storage, driving three types of continuous learning: adaptation of the long-tail format and domain terminology of the parsing model; A / B optimization of strategy thresholds and combination rules without reducing compliance rates to minimize unnecessary outreach and blocking; and business line personalization of grouping strategies to suit different allocation and packaging habits. Through federated learning and differential privacy technology, the system leverages cross-organizational statistics to improve model generalization capabilities while protecting privacy and compliance. The system provides a "Change Impact Panel" to quantify the impact of improvements on reimbursement cycles, rejection rates, document completeness, and reconciliation efficiency, forming a data-driven continuous evolution mechanism.
[0114] The above process is achieved collaboratively by connectors and acquisition agents, parsing and extraction engines, process / event graphs, strategy and risk control engines, proactive interaction orchestrators, group order synthesizers, approval routing and financial adapters, and feedback learning closed-loop modules, ensuring timeliness, stability and traceability through event-driven and streaming processing.
[0115] To implement the methods described in the above embodiments, the system adopts a cloud-native microservice architecture based on Kubernetes and integrates an interpretable AI decision-making framework to ensure high performance, high reliability, scalability, and trustworthy decision-making throughout the entire processing flow.
[0116] (1) Distributed microservice architecture supporting the method flow: This system adopts a microservice architecture based on container and orchestration technologies, mapping the steps of the core methods in the above embodiments to specialized, elastically scalable components hosted by independent services. Each service collaborates through lightweight APIs and event streams to jointly complete an intelligent pipeline from voucher collection to expense report generation.
[0117] Data Acquisition Service Group: This corresponds to the "active acquisition" step in the embodiments of the above method and is further defined. This group includes multiple microservices such as email connector service, API gateway service, and IM acquisition agent service, which are deployed and interface with external data sources (such as corporate email, travel service providers, and instant messaging tools). With user authorization, these services actively listen for and capture original credentials and context data in the form of "connectors," and publish standardized credential acquisition events through a unified event bus (such as one built on Apache Kafka). This can be used in the active acquisition step of the above method through connectors deployed in multi-source systems, and to provide data sources for subsequent real-time processing.
[0118] Voucher Processing and Intelligent Analysis Service Group: This group works collaboratively to implement the three core steps of "parsing and structured extraction", "constructing semantic graphs" and "compliance pre-review" in the embodiments of the above methods.
[0119] Multimodal parsing service: This service receives collection events and invokes a parsing pipeline that integrates an OCR engine, layout analysis model, and domain-adjusted language model (LLM). Specifically, it implements a "dual-track path of template matching and model prediction," performing structured extraction of documents such as images and PDFs, generating standardized document metadata, and attaching parsing confidence and trajectory.
[0120] Semantic Graph Construction Service: This service receives standardized credential metadata and runs a graph computing engine. Based on features such as time, space, and participants, it uses graph matching and community discovery algorithms to associate and cluster discrete credential metadata nodes, dynamically constructing and updating a "trip / event" semantic graph in memory and a graph database (such as Neo4j). This graph is the concrete data embodiment of the "semantic graph representing the same business activity" in the above method embodiments.
[0121] Compliance Pre-screening and Risk Control Service: This service embeds a configurable strategy engine and an ensemble learning-based risk control model. It can perform compliance verification and risk scoring for single tickets and combined scenarios in parallel, based on "semantic graphs." Its decision-making logic (especially the risk scoring part) is supported by an interpretable AI framework, ensuring the interpretability of the output results.
[0122] Business Process Orchestration Service Group: Interaction and Form Arrangement Service: This service monitors the pre-approval results. When a missing or abnormal item is detected, it can initiate proactive interaction with the user by calling the push notification service according to the "low-interference strategy." Upon confirmation or when the pre-approval is passed, it triggers the intelligent form arrangement service. Based on the semantic graph and the pre-approval results, the form arrangement service automatically performs expense merging, tax separation, and apportionment calculation, generating the structured reimbursement form draft described in the embodiments of the above methods.
[0123] Approval Routing and Feedback Learning Service: This service handles draft submissions, dynamically routes approvals according to rules, and synchronizes the results as feedback signals to the feedback learning service upon completion of the process. The feedback learning service uses these signals to continuously optimize the parameters of the analytical model, graph association algorithm, and risk control model, achieving closed-loop evolution of the system.
[0124] (2) Technical architecture for implementing interpretable pre-screening decisions: To specifically support the "generation of interpretable pre-review results" and precise interaction based on the results in the embodiments of the above methods, this system integrates an interpretable AI decision-making system in compliance pre-review and risk control services.
[0125] Multi-layered explanation generation mechanism: The system generates differentiated explanations for different objects. For employees, simple and clear reasons and suggestions are provided in natural language in the interactive cards; for financial auditors, detailed reports including feature contribution analysis, key evidence highlighting, and comparison with similar cases are provided in the approval interface.
[0126] SHAP-based risk attribution: When the risk control model outputs a risk score, the system simultaneously runs the SHAP value calculation engine. This engine can quantify the impact of each input feature (such as "abnormal consumption time", "merchant located in a high-frequency risk area", "deviation from the average consumption of similar trips") on the final score and present it in the form of a visual chart (such as a waterfall chart), so that the conclusion of "high risk" or "recommendation to be rejected" has clear numerical evidence.
[0127] Encapsulation of the Decision Evidence Chain: Each preliminary audit result (whether approved, warned, or blocked) is not an isolated conclusion. The system automatically encapsulates a complete evidence chain that can be traced back to: the original document image that triggered the decision, the parsed structured fields, the specific compliance policy provisions that were hit, and the key features and weights in the risk control model inference. This directly supports the technical effects achieved by the embodiments of the above method and provides technical feasibility for audit traceability.
[0128] Based on this, the technical effects brought about by the above-mentioned architectural collaboration are as follows: The aforementioned distributed architecture and explainable AI architecture do not exist in isolation, but rather work in deep collaboration: the event bus and stream processing ensure the real-time and continuous nature of the process from "active collection" to "preliminary review" and then to "grouping," enabling "preliminary review" to be implemented in the event.
[0129] Microservices enable core capabilities such as "parsing", "graph construction" and "risk control" to be iterated and scaled elastically to meet the massive and ever-changing voucher processing needs of different enterprises.
[0130] The interpretable AI framework is embedded in the risk control service, so that "compliance pre-review based on semantic graph" is no longer a "black box". Its output "interpretable pre-review results" become the key input to drive "proactive human-computer interaction" and "credible approval", thus forming a closed loop from intelligent decision-making to human-computer collaboration.
[0131] In summary, the distributed architecture and explainable AI architecture described in this specific embodiment are preferred technical solutions for implementing the intelligent expense reimbursement processing method defined in the above embodiments of the present invention. This architecture ensures that the method can be implemented and operated in an efficient, reliable, and trustworthy manner, ultimately achieving the inventive objective of improving expense reimbursement processing efficiency, accuracy, and risk control level.
[0132] Based on the same inventive concept, embodiments of the present invention also provide an electronic device, specifically, the electronic device includes a processor and a storage device; the storage device stores a computer program, and the computer program, when run by the processor, executes the method described in any of the above embodiments.
[0133] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. The electronic device 400 includes: a processor 410, a memory 420, a communication interface 430, and a bus 440. The memory 420 stores machine-readable instructions that can be executed by the processor 410. When the electronic device is running, the processor 410 communicates with the memory 420 through the bus 440. The processor 410 executes the machine-readable instructions to perform the steps of the method described above.
[0134] Specifically, the memory 420 and processor 410 can be general-purpose memory and processor, without any specific limitations. When the processor 410 runs the computer program stored in the memory 420, it can execute the above method.
[0135] Processor 410 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 410 or by instructions in software form. The processor 410 may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this invention can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software module can reside in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 420, and processor 410 reads the information from memory 420 and, in conjunction with its hardware, completes the steps of the above method.
[0136] Corresponding to the above method, this embodiment of the invention also provides a computer-readable storage medium storing machine-executable instructions. When the computer-executable instructions are called and run by a processor, the computer-executable instructions cause the processor to perform the steps of the above method.
[0137] In the embodiments provided by this invention, it should be understood that the disclosed apparatus and method can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.
[0138] Furthermore, the units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0139] Furthermore, the functional modules in the various embodiments of the present invention can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0140] It should be noted that if the functionality is implemented as a software module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0141] In this document, relational terms such as first and second are used only to distinguish one entity or operation from another entity or operation, without necessarily requiring or implying any such actual relationship or order between these entities or operations.
[0142] The above description is merely an embodiment of the present invention and is not intended to limit the scope of protection of the present invention. For those skilled in the art, the present invention can have various modifications and variations. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A proactive expense reimbursement processing method based on semantic graphs, characterized in that, The method includes: The acquired original vouchers related to cost and time are subjected to multimodal semantic parsing and structured extraction to generate standardized voucher metadata; the original vouchers are actively collected through connectors deployed in a multi-source system. Based on the temporal, spatial, and business semantic features of the aforementioned credential metadata, a semantic graph representing the same business activity is constructed. Based on the semantic graph, a compliance pre-audit check is performed on the associated credential metadata, and a pre-audit result is generated; A structured draft expense report is generated based on the preliminary review results. The draft expense report is then submitted for approval according to preset rules and synchronized to the financial system after approval.
2. The method according to claim 1, characterized in that, Multimodal semantic parsing and structured extraction of original vouchers, including: Optical character recognition and layout analysis are performed on the original vouchers, and semantic understanding and named entity recognition are performed using a large language model to extract key fields; Based on the key fields, field extraction results are generated using a dual-track approach of template matching and model prediction. The results from the two outputs are then fused based on confidence levels to generate standardized voucher metadata. The standardized voucher metadata includes time information, spatial information, and business semantic information.
3. The method according to claim 2, characterized in that, Constructing a semantic graph representing the same business activity includes: Based on the temporal information, spatial information, and business semantic information of each of the aforementioned voucher metadata, the association characteristics between the original vouchers corresponding to the multiple voucher metadata are determined; the association characteristics include at least one of temporal proximity, geographical proximity, overlap of participating entities, and business keyword matching degree. By using graph matching and community discovery algorithms, the scattered original vouchers are clustered to generate a semantic graph representing the same business activity.
4. The method according to claim 3, characterized in that, The graph matching and community discovery algorithm uses itinerary plans or travel application information as prior guidance and employs a spatiotemporal constraint-based community discovery model for node clustering.
5. The method according to claim 1, characterized in that, Based on the semantic graph, a compliance pre-audit check is performed on the associated credential metadata, including: Based on the semantic graph, the dimensions of a single voucher and the dimensions of a combined scenario consisting of multiple related vouchers are determined; According to the preset compliance strategy, parallel verification is performed at both the single voucher dimension and the combined scenario dimension. An anomaly detection is performed on the voucher metadata using an ensemble learning-based risk control model, generating a preliminary review result that includes risk level and rectification suggestions.
6. The method according to claim 5, characterized in that, The method further includes: When the compliance pre-audit detects missing information or compliance anomalies, it triggers proactive human-computer interaction to complete or confirm the information; The proactive human-computer interaction includes: based on a low-interference strategy, aggregating multiple items to be confirmed and proactively reaching the user through a message interface; wherein, the low-interference strategy includes dynamically selecting the timing of reaching the user based on the user's status, and prioritizing and throttling the interaction requests.
7. The method according to claim 1, characterized in that, The method further includes: The user confirmation during the interaction process, the approval conclusion of the draft expense report, and the audit results of the financial system are used as feedback signals to dynamically update the model parameters or decision rules used in at least one of the following steps: multimodal semantic parsing, semantic graph construction, compliance pre-audit detection, and generation of the draft expense report.
8. A proactive expense reimbursement processing system based on semantic graphs, characterized in that, The system includes: The semantic parsing and structured extraction module is used to perform multimodal semantic parsing and structured extraction on the acquired original vouchers related to cost and time, and generate standardized voucher metadata; the original vouchers are actively collected through connectors deployed in a multi-source system; The semantic graph construction module is used to construct a semantic graph representing the same business activity based on the temporal, spatial, and business semantic features among the credential metadata. The pre-audit detection module is used to perform compliance pre-audit detection on the associated credential metadata based on the semantic graph and generate pre-audit results; The approval module is used to generate a structured draft expense report based on the pre-approval results, submit the draft expense report for approval according to preset rules, and synchronize it to the financial system after approval.
9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions that, when invoked and executed by a processor, cause the processor to perform the method according to any one of claims 1 to 7.