AI tax auditing method based on multi-modal data and attention mechanism
Through the AI tax audit method of multimodal data and attention mechanism, the problem of opaque decision-making process of the AI tax audit model is solved, visual evidence links and blind spot warnings are generated, audit efficiency and interpretability of conclusions are improved, and auditors' trust and risk identification capabilities are enhanced.
Patent Information
- Application Number
- CN202510639810.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-08-19
AI Technical Summary
The decision-making process of the existing AI tax audit model is opaque, audit conclusions are difficult to explain, and the potential cognitive blind spots of the model are difficult to be effectively identified and circumvented.
The AI tax audit method based on multimodal data and attention mechanism is adopted. By pre-processing and unified semantic representation of structured data, text data and image data, the artificial intelligence audit model of the attention mechanism is used to identify potential tax risk points, and a visual attention audit evidence link is generated, which supports interactive traceability operations, identify attention distribution abnormalities and generate audit blind spot warning information.
It enhances auditors' understanding and trust in AI audit conclusions, reduces the time and energy of manually searching and sorting out evidence, and improves the intelligence level and risk resistance of tax audits.
Smart Images

Figure CN120509974A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of tax auditing, and in particular to an AI tax auditing method based on multimodal data and an attention mechanism. Background Art
[0002] With the development of big data and artificial intelligence technologies, the application of AI in the field of tax auditing has attracted increasing attention. AI models, especially deep learning models, have shown great potential in processing massive tax data, identifying potential tax risks and abnormal patterns, and can assist auditors in quickly screening high-risk targets from complex data.
[0003] However, the current application of AI in tax audits still faces several key challenges. Many high-performance AI models have complex and opaque internal decision-making logic. Auditors often only receive a risk score or classification, but struggle to understand the specific evidence and logic used to reach the model's conclusion. This impacts auditors' trust in and acceptance of AI-generated results and makes it difficult to use AI-generated audit conclusions as direct and fully explainable audit evidence.
[0004] In complex tax cases, key audit evidence is often scattered across multiple data sources of various types. Even if the AI model can point out that an individual is at risk, it is still necessary to manually sort out the core chain of evidence supporting the risk judgment from massive amounts of data in a complete and accurate manner.
[0005] Auditors struggle to determine whether the high-risk conclusions drawn by AI models in specific cases are based on a deep understanding of key case information and accurate logical reasoning, or simply stem from superficial statistical features or coincidences learned by the model. This leads to the risk of misjudgment or omission of AI judgments.
[0006] Moreover, any AI model may exhibit "cognitive limitations" or "understanding biases" due to biases in training data, inherent limitations of the model structure, or when faced with certain types of data or complex scenarios. The existing audit process lacks an effective mechanism to proactively identify and warn of these weaknesses that may exist in AI models in specific audit tasks. Summary of the Invention
[0007] The present invention provides an AI tax audit method based on multimodal data and attention mechanism to solve the technical problems in existing AI tax audits, such as the opaque decision-making process of deep learning models, the difficulty in interpreting and trusting audit conclusions, and the difficulty in effectively identifying and avoiding potential cognitive blind spots of the models.
[0008] The present invention provides an AI tax audit method based on multimodal data and attention mechanism, the method comprising:
[0009] Preprocessing multimodal tax audit raw data including structured data, text data, and image data, and performing unified semantic representation on the preprocessed data;
[0010] Using an artificial intelligence audit model based on an attention mechanism to analyze the data after the unified semantic representation to identify potential tax risk points, and capturing the attention weights generated within the artificial intelligence audit model during the analysis process;
[0011] Mapping the captured attention weights back to corresponding data units of the multimodal tax audit raw data, and generating a visual attention audit evidence chain that represents the strength of association between different audit evidence and the focus of the artificial intelligence audit model, wherein the visual attention audit evidence chain supports auditors in performing interactive tracing operations;
[0012] Analyze the attention weights generated internally during the analysis process of the artificial intelligence audit model to identify attention distribution anomalies when the artificial intelligence audit model processes specific data areas or evidence chain links, and generate audit blind spot warning information based on the identified attention distribution anomalies.
[0013] Preferably, the multimodal tax audit raw data is preprocessed and semantically represented in a unified manner, including:
[0014] Standardized cleansing and engineering of audit-related features for structured data selected from corporate financial statements and tax returns;
[0015] Performing natural language processing on text data selected from audit contracts, business agreements, and tax laws and policies, wherein the natural language processing includes at least named entity recognition, semantic relationship extraction, and generation of text semantic embedding representations;
[0016] Applying optical character recognition technology to image data selected from the invoice image and the original voucher scan to extract key text information, and associating layout information of the key text information in the image data;
[0017] Cross-modal alignment technology is used to associate and uniformly represent the structured data processing results, text data processing results, and image data processing results obtained in the above processing steps at the semantic level.
[0018] Preferably, the artificial intelligence audit model based on the attention mechanism is used for analysis and capture of attention weights, including: using a deep learning model architecture that includes a self-attention mechanism or a cross-attention mechanism or a combination of the two to perform tax risk analysis on data after unified semantic representation; in the reasoning process of risk identification by the artificial intelligence audit model, the attention weight matrix generated by its internal attention layer related to the final decision, which reflects the degree of mutual attention between different input data units, is captured and stored in real time.
[0019] Preferably, the generating of a visual attention audit evidence chain representing the strength of association between different audit evidence and the focus of the artificial intelligence audit model and supporting interactive tracing operations includes:
[0020] Aggregate the multi-head attention weights and multi-layer attention weights captured by the AI audit model to form a comprehensive attention view;
[0021] Mapping the attention weights in the comprehensive attention view to corresponding data units of the multimodal tax audit raw data, where the corresponding data units include words or sentences in text data, accounting items or amounts in structured data, and specific text regions or image regions in image data;
[0022] Use at least one visual representation method selected from dynamic heat maps, highlighted text tags, and variable-strength association lines to display on the user interface the focus of the AI audit model's attention when making risk assessments, as well as the process of establishing evidence associations and allocating attention across multiple documents of different types;
[0023] In response to the auditor's interactive click operation on the visual elements presented in the user interface and marked as high attention values, the original audit evidence fragments or complete documents directly related to the highlighted elements are instantly located and displayed, and the auditor is allowed to further drill down to view deeper detailed information.
[0024] Preferably, the attention weights generated internally during the analysis process of the artificial intelligence audit model are analyzed to identify the attention distribution anomaly, including: based on pre-set heuristic rules derived from audit field knowledge, and attention anomaly patterns learned and summarized from historical audit data and expert experience through machine learning methods, jointly defining multiple types of attention distribution anomaly patterns that may characterize the artificial intelligence audit model's insufficient understanding or insufficient evidence.
[0025] Preferably, the types of abnormal attention distribution patterns include at least the following four: a diffuse attention pattern in which the model's attention weight distribution in key decision areas is too evenly distributed without a clear focus; an attention misplacement pattern in which the model's attention is highly concentrated on secondary information or noise data that is irrelevant to the risk judgment logic; an attention loss pattern in which the model fails to allocate sufficient attention to key information areas of evidence that should be judged as key based on audit experience or domain knowledge; and an attention pattern inconsistency in which the attention pattern of the current audit case deviates significantly from the typical attention pattern of historically confirmed cases with similar risk characteristics.
[0026] Preferably, the audit blind spot warning information is generated based on the identified attention distribution anomalies, including: when one or more attention distribution anomaly patterns are identified and the quantitative indicators of the anomaly patterns meet the preset significance conditions, the data areas or evidence chain links corresponding to the anomaly patterns are specially visually marked in the user interface of the visual attention audit evidence chain; and clear warning information is issued to the auditors, the warning information includes a brief explanation of the cause of the attention distribution anomaly and recommended audit directions to guide the auditors to conduct targeted manual review.
[0027] Preferably, the method further includes the following steps: integrating the potential tax risk points identified by the artificial intelligence audit model, the presented visual attention audit evidence chain, the generated audit blind spot warning information, and the manual review opinions, corrections and supplementary annotations of the auditors on the aforementioned information during the interaction process; and automatically or semi-automatically generating a structured audit working paper or audit report based on all the integrated information, which includes a risk description, key evidence links, potential blind spot prompts of the artificial intelligence audit model and manual review.
[0028] The technical solution provided by this application has at least the following technical effects or advantages: by converting the attention weights within the AI model into an intuitive visual attention audit evidence chain and supporting interactive traceability, it enhances auditors' understanding, trust, and willingness to adopt AI audit conclusions. The interactive visual evidence chain can help auditors quickly locate, connect, and review core evidence related to specific risk points from massive amounts of multimodal data, significantly reducing the time and effort required to manually search, screen, and organize evidence.
[0029] Even relatively inexperienced auditors can use this technology to intuitively observe and learn how AI models analyze complex tax cases, identify key risk points, and correlate core evidence. The AI's focus reflects, to a certain extent, condensed "expert experience," while its "blind spot warnings" help cultivate critical thinking and risk awareness.
[0030] To address the tedious, time-consuming and error-prone problems of manually concatenating core evidence chains in multi-source heterogeneous data, this invention provides clear insights and action guidance through "visual explanation" and "blind spot warning".
[0031] The AI "audit blind spot" warning data recorded during operation can, through long-term accumulation and statistical analysis, form an "AI audit cognitive difficulty map" or "audit risk pattern knowledge base." This "metadata" can not only guide subsequent iterations and optimizations of the AI audit model, but also improve auditor training content and skills development priorities. It can even provide valuable data reference for clarifying the revision of tax regulations and improving audit standards.
[0032] This invention's "Audit Blind Spot Warning" mechanism enables AI to proactively "expose" potential misunderstandings or uncertainties in its analysis of specific cases, thereby guiding human auditors to conduct more targeted and efficient reviews. The human auditors' professional judgment, review conclusions, and supplementary findings serve as high-quality feedback, used to revise the final conclusion of the current audit case and continuously learn and improve the AI model in the future. This comprehensively enhances the intelligence level and risk mitigation capabilities of tax audit work. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 This is a flow chart of the AI tax audit method based on multimodal data and attention mechanism of the present invention. DETAILED DESCRIPTION
[0034] The present invention relates to an AI tax audit method based on multimodal data and an attention mechanism to address technical issues in existing AI tax audits, such as the opaque decision-making process of deep learning models, the difficulty in interpreting and trusting audit conclusions, and the difficulty in effectively identifying and circumventing potential cognitive blind spots of the models.
[0035] The above technical solution will be described in detail below in conjunction with the accompanying drawings and specific implementation methods of the specification to better understand the above technical solution. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments of the present invention. It should be understood that the present invention is not limited to the example embodiments used only to explain the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. In addition, it should be noted that, for the convenience of description, only the parts related to the present invention, rather than all, are shown in the drawings. Example 1
[0036] like Figure 1 The flowchart of the AI tax audit method based on multimodal data and attention mechanism is shown in the figure. The method includes the following steps:
[0037] Preprocessing multimodal tax audit raw data including structured data, text data, and image data, and performing unified semantic representation on the preprocessed data;
[0038] Using an artificial intelligence audit model based on an attention mechanism to analyze the data after the unified semantic representation to identify potential tax risk points, and capturing the attention weights generated within the artificial intelligence audit model during the analysis process;
[0039] Mapping the captured attention weights back to corresponding data units of the multimodal tax audit raw data, and generating a visual attention audit evidence chain that represents the strength of association between different audit evidence and the focus of the artificial intelligence audit model, wherein the visual attention audit evidence chain supports auditors in performing interactive tracing operations;
[0040] Analyze the attention weights generated internally during the analysis process of the artificial intelligence audit model to identify attention distribution anomalies when the artificial intelligence audit model processes specific data areas or evidence chain links, and generate audit blind spot warning information based on the identified attention distribution anomalies.
[0041] The specific principle of this method is as follows:
[0042] Receive electronic financial statements submitted by businesses, such as income statements, balance sheets, cash flow statements, etc. that comply with specific accounting standards, tax returns (such as VAT returns and corporate income tax returns), and transaction data and account balance sheets exported from enterprise resource planning (ERP) systems and financial software. First, identify the type and format of this data.
[0043] Unify the format of structured data from different sources or versions. For example, unify dates to the "year-month-day" format, unify numeric data to a specific number of decimal places, and map or convert accounting account codes according to the preset unified account system to ensure consistency in subsequent processing.
[0044] Identify and address obvious logical errors (e.g., negative revenue in the income statement without a reasonable explanation, an unbalanced balance sheet, etc.). This can be detected through pre-set validation rules or comparison with historical data. These errors can be flagged, corrected (e.g., based on contextual inference), or deleted when there is clear evidence.
[0045] Detect missing values in key fields. Depending on the importance of the field and the characteristics of the data, you can use mean / median / mode filling, regression-based prediction filling, or mark them as special values for subsequent model processing.
[0046] Identify and remove exact duplicate rows to avoid biasing statistical analysis and model training.
[0047] Based on tax audit experience and risk analysis needs, new, more valuable features for risk identification are derived from cleansed structured data. For example, key financial ratios such as debt-to-asset ratio, quick ratio, gross profit margin, net profit margin, and accounts receivable turnover can be calculated. Year-on-year or month-on-month growth rates for specific accounting items can be calculated to identify unusual fluctuations.
[0048] Construct binary indicator variables, such as "whether there are large overseas payments," "whether there are unusual transaction amounts with related parties," and "whether specific expense items exceed normal proportions." These newly constructed features will serve as input to the AI model along with the raw data.
[0049] Semi-structured / unstructured data processing flow:
[0050] Text data processing (based on natural language processing (NLP) technology): including audit engagement letters, various business contracts (purchase, sales, services, loans, etc.), board of directors or shareholders' meeting resolutions, important correspondence emails, text descriptions of related-party relationships and transactions, compilations of tax laws and policies, and corporate responses to tax inquiries.
[0051] Split continuous text paragraphs into independent word units (Chinese word segmentation) or sentence units.
[0052] Automatically identify and annotate predefined entity categories in text, such as organization names, personal names, place names, dates, amounts, contract numbers, tax names, regulatory clause numbers, and specific business terms (such as "technology transfer" and "equity incentives").
[0053] Further analysis and extraction of possible semantic relationships between identified entities. For example, from a contract text, entities such as "Party A," "Party B," "Contract Amount," "Signing Date," and "Subject Matter" and their constituent relationships such as "Sign," "Pay," and "Provide" can be extracted.
[0054] The processed text units (words, sentences, or entire document fragments) are converted into dense numerical vectors (i.e., word embeddings or sentence embeddings) that capture their deep semantic information. This is typically achieved using deep learning language models (such as BERT and RoBERTa) pre-trained on large-scale text corpora, or fine-tuned on text data from a specific tax audit domain to obtain vector representations that better align with domain semantics.
[0055] Image data processing: mainly scans or photos of various invoices (such as special VAT invoices, ordinary invoices), bank receipts, warehouse receipts, warehouse delivery receipts and other original documents.
[0056] Apply OCR technology to image data to convert printed or handwritten text information into machine-readable text format.
[0057] Based on the OCR-recognized text, combined with the image's layout information (if available, for example, through a layout analysis model like LayoutLM), predefined key fields are automatically extracted. For example, for invoices, information such as the invoice code, invoice number, invoice date, purchaser's name and taxpayer identification number, seller's name and taxpayer identification number, name of goods or taxable services, specification model, unit, quantity, unit price, amount, tax rate, tax amount, and total price and tax are automatically extracted. This step can utilize regular expression matching, rule-based extraction logic, or training of specialized information extraction models.
[0058] The extracted key text information will be associated with its location information in the original image (e.g., coordinate box) and can be further linked to other text data processed by NLP (e.g., contract content).
[0059] Voice / video data processing: such as recordings of audit interviews with management personnel or key personnel of the audited unit, and video materials recorded during on-site inspections of inventory, fixed assets, etc.
[0060] Automatic speech recognition technology was applied to the interview recordings to convert them into text transcripts.
[0061] For video data, key video frames or video clips containing important audit clues can be extracted through techniques such as scene change detection, object recognition, or behavioral analysis. These extracted text, images, or video clips can then be further processed according to the above-mentioned processing flow for text or image data.
[0062] Utilize the time information contained in various types of data (such as financial reporting period, invoice date, contract signing date, and bank transaction date) to preliminarily align and sort data from different sources along the time dimension.
[0063] Link different pieces of data related to the same audit object or transaction based on common entities identified in various types of data (such as company name, individual name, contract number, invoice number).
[0064] Cross-modal learning techniques are employed, for example, to construct a shared semantic vector space. The goal is to ensure that data from different modalities describing the same concept or fact (e.g., the textual description of "payment terms" in a contract, the "payment" record of the corresponding amount in a bank statement, and the "payee" information displayed on an image of an invoice) have similar or close vector representations in this shared space. This can be achieved by designing a specific multimodal neural network architecture and training it using self-supervised or supervised learning methods such as contrastive learning and translation tasks.
[0065] After the above alignment and fusion processing, each audit-related evidence or data point (regardless of whether its original form is structured, textual or image) will be converted into a unified numerical representation containing rich context and cross-modal correlation information (for example, an enhanced embedding vector that integrates multiple information sources), which lays the foundation for the unified processing and in-depth analysis of subsequent AI audit models.
[0066] In-depth analysis is performed on the pre-processed and fused multimodal tax data to identify potential tax risk points, and in the process, attention weights are generated that can be used for interpretation and tracing.
[0067] Transformer-based deep learning model architectures are preferred. These models are favored for their powerful sequence processing capabilities and ability to capture long-range dependencies. For example, BERT and its variants (such as RoBERTa and ALBERT) can be used as foundational modules for text understanding; LayoutLM and its variants can be used to process document images containing both text and layout information (such as invoices and scanned contracts); or a custom multimodal Transformer model can be designed that directly accepts and processes unified semantic representations from different data modalities.
[0068] Application of the Attention Mechanism: Self-Attention: The model's internal self-attention layer allows it to dynamically calculate the mutual importance or correlation between individual elements (words, accounts) in an input sequence (for example, all the words in a contract or all the account lines in a financial statement). The model generates a weighted contextual representation for each element in the sequence, where the weight is derived from the "attention" paid to it by other elements. This enables the model to understand the meaning of a word in a specific context or the relative importance of a financial figure within the overall structure of a report.
[0069] Cross-Attention: This mechanism is particularly important when processing multimodal data or when comparing or correlating different information sources. For example, when a model needs to determine the authenticity of an invoice, it can leverage the cross-attention mechanism to allow invoice information (such as the amount or product name) to "focus" on relevant clauses in the contract, or vice versa, allowing the contract clauses to "focus" on the invoice information, thereby determining whether the two match. Attention weights indicate which invoice fields are most closely and critically associated with which contract sections.
[0070] Multi-Head Attention: Transformer models typically employ a multi-head attention mechanism. This means the model doesn't have a single "attention perspective," but rather independently calculates attention weights from multiple different "subspaces" or "perspectives" simultaneously, then combines the attention results from these different "heads." This allows the model to simultaneously capture multiple different types or levels of dependencies and association patterns in the data.
[0071] Input representation combination: The output of unified semantic representations of various types of aligned and fused data (e.g., word embedding sequences for text, feature vectors for image regions, numerical features of structured data, etc.) is combined in a predetermined manner as input to the AI model. A simple combination method can involve directly concatenating feature vectors from different sources; a more complex method can involve designing a specific input structure. For example, different input segments are set for data of different modalities, and special segmentation markers (such as the "[SEP]" marker) are used to distinguish them. The model then learns how to integrate this cross-segment information.
[0072] Task design for model training: The main risk identification task can be designed as follows:
[0073] Classification tasks: For example, output a predefined risk category (such as "high risk - false invoicing", "medium risk - unfair pricing of related-party transactions", "low risk", etc.) for each audit case (or enterprise, specific transaction set).
[0074] Risk level assessment task: Output a continuous risk score (e.g., 0 to 100), where higher scores indicate greater risk.
[0075] Anomaly detection tasks involve identifying specific transaction patterns or data combinations that significantly deviate from normal business or financial behavior. Training data is typically a large number of historically audited cases labeled with risk types and levels.
[0076] Auxiliary learning tasks: To enhance the model's understanding of tax audit logic and the effectiveness of representation learning, auxiliary learning objectives can be introduced during training. For example, the evidence importance ranking task requires the model to not only determine risk but also identify which input pieces of evidence (such as specific contract terms or invoices) contribute most to the final risk judgment. Key risk factor extraction tasks require the model to identify and output the specific risk factors or feature combinations that lead to a specific risk judgment. By simultaneously learning the main task and these auxiliary tasks, the model can be encouraged to learn more interpretable internal representations, potentially indirectly improving the performance of the main task.
[0077] Attention weight capture and storage mechanism: During the process of the AI model performing forward propagation and reasoning on a new audit case (i.e., performing risk judgment), the calculated attention weight matrix is extracted in real time from each attention layer within the model (especially those key layers that are closely connected to the final risk output layer, or the attention layers that perform cross-modal information interaction).
[0078] These weight matrices (for each attention head) typically represent the attention score or probability distribution assigned to one position in the input sequence (the "query") over all other positions in the sequence (the "keys"). A higher score indicates that the "query" has gained more information from the corresponding "key" when generating its own output representation.
[0079] Storage content: The original attention weight data, which may contain multiple heads and multiple levels, needs to be stored and associated with the case ID of the current analysis, the input data segment identifier (such as the document ID, the start and end positions in the text, the financial statement account line number, the invoice area coordinates, etc.), and the risk judgment results made by the model for subsequent visualization interpretation and blind spot warning.
[0080] Convert the complex attention weight data within the AI model into a visual form that auditors can intuitively understand and interactively explore, and allow auditors to trace back to the original audit evidence through the visual interface.
[0081] Attention weight aggregation and mapping backtracking process: Since models such as Transformer usually contain multiple attention layers, each layer contains multiple attention heads, it would be very complicated to directly analyze all these raw weights. Therefore, it is necessary to aggregate these weights to obtain a more macro and easier to interpret attention view. Aggregation methods can include:
[0082] Averaging method: averaging the weight matrices of all attention heads in the same layer element by element, or averaging the weights of all (or a few selected key) attention layers in the model.
[0083] Weighted average method: According to the importance of different attention heads or different layers in a specific task (which may be pre-evaluated through some interpretable analysis methods), different weights are assigned to them for weighted averaging.
[0084] Specific head / layer selection: Based on domain knowledge or model analysis, specific attention heads or weights of specific layers that are considered to best reflect the key reasoning process are selected as representatives.
[0085] Accurate mapping of weights to data units: The aggregated attention weight values are accurately associated and mapped back to the original data units that constitute the model input. This means:
[0086] For text data, attention weights need to be assigned to specific words, phrases, sentences, or even paragraphs.
[0087] For structured data (such as financial statements), weights need to be mapped to specific accounting account names, account amounts, report rows, or cells.
[0088] For image data (such as invoice images), weights need to be assigned to text regions identified by OCR, or to specific pixel regions on the image that the model focuses on (if the model has the ability to process image pixels directly).
[0089] For cross-document or cross-modal attention, it is necessary to clearly know which part of one document generates and how strong the attention connection is with which part of another (or the same) document. This precise mapping is the basis for effective visualization.
[0090] Generation and presentation of the “Audit Evidence Chain Heat Map”:
[0091] Visualization: Using a variety of intuitive visualization methods, on a unified user interface or through multiple synchronously linked view windows, to show which data and evidence the AI model's "focus of attention" is mainly distributed when making a specific risk judgment. Common visualization techniques include:
[0092] Text highlighting / heatmap: For text-based evidence, words or sentences are highlighted using different shades of background color (e.g., darker red indicates higher attention) based on their attention weight.
[0093] Node-link graphs (network graphs): Different audit evidence (e.g., different documents, different financial statements, different transaction records) are represented as nodes in the graph. Nodes are connected by lines with direction and thickness / color coding to indicate the strength of the connections between them revealed by the attention mechanism. The size or color of the node itself can also be used to indicate the importance of the evidence itself (e.g., based on the total amount of attention it has accumulated).
[0094] Table / matrix coloring: For tabular data (such as financial reports), you can directly use color depth (heat map) on the cells to indicate the degree of attention the model pays to the content of each cell.
[0095] Image region overlay: For image evidence such as invoices, a semi-transparent color heat map can be overlaid on the original image to indicate which areas of the image the model has paid more attention to.
[0096] The key to visualizing cross-document and cross-modal associations is not just to show the distribution of attention within a single document. More importantly, it clearly and intuitively demonstrates how the AI model establishes connections across multiple documents of different types and modalities and allocates attention accordingly. For example, a brightly colored, thick line connecting Clause X of Contract A to the amount column of Invoice B could be used on the interface to indicate that the model believes there is a strong association between the two, and that this association has a significant impact on the final risk assessment. In this way, auditors can clearly see how the AI "connects together" scattered evidence to form a judgment.
[0097] Interactive traceability and drill-down functionality: Auditors can select elements marked as having high attention values by clicking the mouse in the visual interface. These elements can include the darkest areas in the heat map, highlighted words in the text, large or colored nodes in the node-link diagram, or edges with high connection strength.
[0098] When users click on these highlighted elements, the system will respond immediately, automatically locate and display the original audit evidence content directly related to the highlighted elements. For example:
[0099] Click on the name of a financial statement account to display the account's detailed account records, related general ledger voucher information, and even link to scanned copies of related original documents in a new window or sidebar.
[0100] Clicking a highlighted clause in a contract will display the full text of the contract and automatically scroll the view to the clause.
[0101] Clicking a highlighted area on an invoice image (such as the product name or amount) will display the full image of the invoice and provide information about all OCR-recognized fields on the invoice.
[0102] Starting from the original evidence currently presented, further "drill-down" operations should be supported, that is, to view other evidence directly related to it. For example, a single accounting voucher can be linked to all invoice images attached to it; and from a single invoice, other related contracts or bank statements can be found based on counterparty information. In this way, auditors can roam and explore complex evidence networks by following the AI's "attention path" or clues of their own interest.
[0103] At the same time, in order to facilitate auditors to focus on specific risk points or types of evidence, the visual interface should also provide auxiliary interactive functions such as filtering (for example, only displaying evidence and attention flows related to "related transactions"), sorting (for example, sorting evidence by attention intensity), and keyword filtering (for example, searching for specific words in highlighted evidence).
[0104] By conducting in-depth analysis of the attention weight distribution patterns generated by the AI model, it aims to automatically identify situations where the model may have insufficient understanding, insufficient evidence, or abnormal reasoning logic when processing specific audit cases or specific information areas, and to issue timely warnings to auditors.
[0105] Definition and identification method of abnormal attention distribution patterns, combining a hybrid identification strategy with preset rules and machine learning:
[0106] Heuristic rule definition: Based on the experience and audit domain knowledge of senior tax audit experts, a series of heuristic rules that can characterize "irrational attention distribution" or "possible understanding bias of the model" are pre-defined.
[0107] Machine learning methods can be used to supplement and optimize these heuristic rules, and even discover some complex abnormal patterns that are difficult for humans to detect.
[0108] Examples of abnormal attention distribution patterns:
[0109] Dispersed attention: This manifests as the AI model's attention weighting being too evenly distributed and dispersed in areas that, according to audit logic, should contain key information (e.g., core transaction terms in contracts, items with large, unusual fluctuations in financial statements), lacking a clear, concentrated focus. This may indicate that the model is failing to effectively filter and capture the truly decisive key points from a wealth of information, raising questions about the reliability of its judgment.
[0110] Excessive focus on non-critical information: This occurs when the AI model's attention is drawn to secondary information, even noise data, that is clearly irrelevant or less important to the current audit risk assessment logic, assigning excessive attention weight to these information. This may indicate that the model has learned some spurious statistical correlations or has a biased understanding of the problem.
[0111] Lack of attention to key information areas: This manifests as the AI model paying significantly less attention to, or even completely ignoring, evidence areas or data items that, based on audit experience, domain knowledge, or relevant regulations, are crucial for assessing a specific tax risk point. This directly indicates that the AI model may have a "blind spot" and is failing to fully utilize key evidence.
[0112] Attention patterns inconsistent with historical cases with similar risks): A historical case library can be maintained, storing historically confirmed cases with similar risk characteristics (for example, the risk of "false VAT invoicing") and their corresponding "typical" or "healthy" attention distribution patterns. When processing a new audit case, if the resulting attention pattern deviates significantly and unexplainably from the typical pattern of similar risk cases in the historical library, this may indicate that the AI's understanding or judgment logic of the current case is unconventional, and caution is required.
[0113] Weak cross-modal attention connections: For risk points that require the comprehensive use of multiple types or sources of evidence (for example, the authenticity of complex transactions that require mutual verification of contract text, bank statements, and invoice images), if the AI model's attention connection strength between these different modal evidence that should be strongly correlated is very weak, or key cross-modal connections are missing, it indicates that the model may have failed to effectively integrate cross-modal information and its comprehensive judgment ability may be flawed.
[0114] Marking and early warning mechanism for “audit blind spots” or “areas of insufficient evidence”:
[0115] Visual marking: When one or more abnormal patterns of attention distribution are identified in the current audit case through the above methods, and the degree of such abnormality (e.g., measured by quantitative indicators) exceeds the preset trigger threshold, the specific data areas, evidence fragments, or links in the chain of evidence related to the abnormal attention will be specially and prominently marked in the provided visual interface. For example, these areas can be highlighted in a different color (such as warning yellow or red) than normal attention, or a warning icon or text prompt box can be attached next to them.
[0116] Clarify the generation and delivery of early warning information: At the same time, clear early warning information will be issued to auditors. The early warning information should succinctly indicate the nature and possible location of the problem. For example: "Warning: The AI model may be distracted by the 'Cost of Main Business' item in the '2023 Income Statement of XX Company'. It is recommended to manually review the rationality of its cost structure and the adequacy of related documents." Or "Reminder: Regarding the currently determined 'related-party transaction pricing risk', the AI model's attention to Clause Y regarding the pricing basis in the 'Service Contract between Company A and Company B' is significantly lower than expected. There may be a risk of insufficient understanding of this key evidence. Please pay special attention."
[0117] Warning Levels and Explanations: Warnings can be categorized into different levels (e.g., "General Concern," "Major Concern," "High Risk") based on the type, number, and severity of the detected attention patterns, as well as their potential impact on potential audit risks. Each warning should, whenever possible, include a brief explanation of its cause (e.g., "distracted attention," "missing key information," or "inconsistent with historical patterns") to help auditors better understand the implications of the warning and determine subsequent audit responses.
[0118] Integrate the analysis results of the AI model, insights from the visual evidence chain, prompts for audit blind spot warnings, and the manual review opinions and supplementary findings made by auditors during the interactive process to ultimately generate structured and well-informed audit working papers and audit reports.
[0119] The integration process of AI analysis results and manual audit opinions: automatically summarizes all risk judgments made by the AI model on the current audit case (such as risk type, risk level score), the visual attention evidence chain constructed for each judgment (which may include snapshots or links of key evidence fragments), and all audit blind spot warning information issued for each judgment link.
[0120] Auditors can manually review each AI judgment, each chain of evidence, and each blind spot warning on the provided interface (usually linked to a visual interface). They can:
[0121] Confirm: Agree with the AI's judgment.
[0122] Revision: If the AI’s risk level assessment is deemed inaccurate, adjustments can be made; if the evidence the AI focuses on is deemed incomplete or incorrect, supplements or corrections can be made.
[0123] Supplementary comments: You can add your own analysis, reasoning, questions, or instructions for further investigation for any findings or warnings made by the AI.
[0124] Handling Alerts: After manually reviewing blind spot alerts issued by AI, auditors should document the review process and conclusions (e.g., confirming the AI's concerns or eliminating the blind spot). These auditors' actions and input should be recorded and clearly distinguished from the AI's original analysis results.
[0125] Automatic generation support for structured audit working papers or reports: Templates for audit working papers and audit reports can be built-in or allowed to be customized by users.
[0126] When generating a report, the following information will be automatically filled into the corresponding positions according to the template: basic information of the audit project.
[0127] Identify the main tax risk points and their descriptions. For each risk point, link to or embed the most illustrative AI-generated visual screenshot of the attention evidence chain.
[0128] List the "audit blind spot" warnings and their contents that AI has issued during the risk point judgment process.
[0129] Clearly display the auditor's manual review of the AI judgment and blind spot warnings, any corrections, and the final confirmed audit conclusions. Attach key snippets or indexes of the original audit evidence related to the risk point.
[0130] The contents of the generated reports or working papers (especially those related to evidence) should, as far as possible, maintain links to the original data and visual analysis results to facilitate subsequent review and tracing.
[0131] Through the above process description, this solution can not only help AI models identify risks, but also help auditors understand the basis for AI's judgment and proactively discover AI's potential limitations, thereby promoting more efficient, reliable, and intelligent human-machine collaborative tax audits.
[0132] It should be understood that the embodiments disclosed in the present invention and the above description can enable those skilled in the art to use the present invention to implement the present invention. At the same time, the present invention is not limited to the embodiments mentioned above. It should be understood that those skilled in the art can still modify the technical solutions described in the above embodiments or replace some of the technical features therein with equivalents; and such modifications or replacements do not deviate from the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention and are all included in the scope of protection of the present invention.
Claims
1. An AI tax audit method based on multimodal data and attention mechanism, characterized by: The method comprises: Preprocessing multimodal tax audit raw data including structured data, text data, and image data, and performing unified semantic representation on the preprocessed data; Using an artificial intelligence audit model based on an attention mechanism to analyze the data after the unified semantic representation to identify potential tax risk points, and capturing the attention weights generated within the artificial intelligence audit model during the analysis process; Mapping the captured attention weights back to corresponding data units of the multimodal tax audit raw data, and generating a visual attention audit evidence chain that represents the strength of association between different audit evidence and the focus of the artificial intelligence audit model, wherein the visual attention audit evidence chain supports auditors in performing interactive tracing operations; Analyze the attention weights generated internally during the analysis process of the artificial intelligence audit model to identify attention distribution anomalies when the artificial intelligence audit model processes specific data areas or evidence chain links, and generate audit blind spot warning information based on the identified attention distribution anomalies.
2. The AI tax audit method based on multimodal data and attention mechanism according to claim 1 is characterized in that: The multimodal tax audit raw data is preprocessed and semantically represented in a unified manner, including: Standardized cleansing and engineering of audit-related features for structured data selected from corporate financial statements and tax returns; Performing natural language processing on text data selected from audit contracts, business agreements, and tax laws and policies, wherein the natural language processing includes at least named entity recognition, semantic relationship extraction, and generation of text semantic embedding representations; Applying optical character recognition technology to image data selected from the invoice image and the original voucher scan to extract key text information, and associating layout information of the key text information in the image data; Cross-modal alignment technology is used to associate and uniformly represent the structured data processing results, text data processing results, and image data processing results obtained in the above processing steps at the semantic level.
3. The AI tax audit method based on multimodal data and attention mechanism according to claim 1 is characterized in that: The use of an artificial intelligence audit model based on an attention mechanism for analysis and capturing attention weights includes: using a deep learning model architecture that includes a self-attention mechanism or a cross-attention mechanism or a combination of the two to perform tax risk analysis on data after unified semantic representation; in the reasoning process of risk identification by the artificial intelligence audit model, real-time capture and storage of the attention weight matrix generated by its internal attention layer related to the final decision, which reflects the degree of mutual attention between different input data units.
4. The AI tax audit method based on multimodal data and attention mechanism according to claim 1 is characterized in that: The generation of a visual attention audit evidence chain representing the strength of association between different audit evidence and the focus of the artificial intelligence audit model and supporting interactive tracing operations includes: Aggregate the multi-head attention weights and multi-layer attention weights captured by the AI audit model to form a comprehensive attention view; Mapping the attention weights in the comprehensive attention view to corresponding data units of the multimodal tax audit raw data, where the corresponding data units include words or sentences in text data, accounting items or amounts in structured data, and specific text regions or image regions in image data; Use at least one visual representation method selected from dynamic heat maps, highlighted text tags, and variable-strength association lines to display on the user interface the focus of the AI audit model's attention when making risk assessments, as well as the process of establishing evidence associations and allocating attention across multiple documents of different types; In response to the auditor's interactive click operation on the visual elements presented in the user interface and marked as high attention values, the original audit evidence fragments or complete documents directly related to the highlighted elements are instantly located and displayed, and the auditor is allowed to further drill down to view deeper detailed information.
5. The AI tax audit method based on multimodal data and attention mechanism according to claim 1, characterized in that: The attention weights generated internally during the analysis process of the artificial intelligence audit model are used to identify the attention distribution anomaly, including: based on pre-set heuristic rules derived from audit field knowledge, and attention anomaly patterns learned and summarized from historical audit data and expert experience through machine learning methods, jointly defining multiple types of attention distribution anomaly patterns that may represent insufficient understanding or insufficient evidence of the artificial intelligence audit model.
6. The AI tax audit method based on multimodal data and attention mechanism according to claim 5 is characterized in that: The abnormal attention distribution patterns include at least the following four types: a diffuse attention pattern in which the model's attention weight distribution in key decision areas is too evenly distributed without a clear focus; The model's attention is highly focused on secondary information or noise data that is irrelevant to the risk judgment logic; The model's attention deficit pattern is characterized by a failure to allocate sufficient attention to key information areas that should be considered key evidence areas based on audit experience or domain knowledge; and an inconsistency in attention patterns is characterized by a significant deviation between the attention pattern of the current audit case and the typical attention pattern of historically confirmed cases with similar risk characteristics.
7. The AI tax audit method based on multimodal data and attention mechanism according to claim 1, characterized in that: The method of generating audit blind spot warning information based on the identified attention distribution anomalies includes: when one or more attention distribution anomaly patterns are identified and the quantitative indicators of the anomaly patterns meet the preset significance conditions, in the user interface of the visual attention audit evidence chain, the data area or evidence chain link corresponding to the anomaly pattern is specially visually marked; and clear warning information is issued to the auditor, and the warning information includes a brief explanation of the cause of the attention distribution anomaly and recommended audit directions to guide the auditor to conduct targeted manual review.
8. The AI tax audit method based on multimodal data and attention mechanism according to claim 1, characterized in that: The method also includes the following steps: integrating the potential tax risk points identified by the artificial intelligence audit model, the presented visual attention audit evidence chain, the generated audit blind spot warning information, and the manual review opinions, corrections and supplementary annotations of the auditors on the aforementioned information during the interaction process; and automatically or semi-automatically generating a structured audit working paper or audit report containing risk descriptions, key evidence links, potential blind spot prompts of the artificial intelligence audit model and manual review status based on all the integrated information.
Citation Information
Cited By
Purchase file verification method and device based on multi-mode and rule optimization
CN120806825A
File auditing method, device and system and medium
CN121723993A
Document review methods, devices, systems and media
CN121723993B
Audit model construction method and system suitable for multi-modal data
CN121883195A