National library centralized payment anomaly detection system and method based on deep learning
By using deep learning technology, combined with multi-source data perception and intelligent analysis, semantic conflicts and cross-domain risks in centralized treasury payments are identified, solving the problems of missed and false judgments in traditional detection methods, and achieving efficient risk identification and cross-domain sharing.
Patent Information
- Application Number
- CN202511689690.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-18
- Publication Date
- 2026-02-10
AI Technical Summary
Traditional manual review and fixed rule detection are difficult to adapt to the complex multi-source heterogeneous data and dynamic risk patterns in centralized treasury payments. They cannot identify semantic conflicts, hidden related transactions and cross-domain risks, leading to missed judgments and misjudgments.
The system employs a deep learning-based centralized treasury payment anomaly detection system. Through multi-source data perception, intelligent analysis, decision application, and federated collaboration technology, combined with multimodal deep learning, fiscal BERT, T-GNN, and explainable AI, it identifies semantic conflicts and cross-domain risks, enabling real-time risk warnings and human-machine collaborative review.
It improves the accuracy of anomaly detection, can accurately identify hidden risks, provide natural language explanations and visual reports, protect data privacy while enabling cross-domain risk sharing and joint prevention and control, and adapt to business changes and new risks.
Smart Images

Figure CN121504470A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of fiscal informatization and artificial intelligence technology, specifically to a deep learning-based system and method for detecting anomalies in centralized treasury payments. Background Technology
[0002] With the deepening of fiscal informatization, centralized treasury payment, as a core link in fiscal fund management, has seen its business scale continuously expand, and payment data has become increasingly multi-sourced, heterogeneous, and dynamic. On the one hand, data sources cover payment application data (amount, payee, budget item, etc.), internal data (budget indicators, historical payment records, project database information), external data (business registration information of payment recipients, records of dishonesty, government procurement filing data), and unstructured data (scanned copies of contracts, text of application reasons). On the other hand, the risk patterns faced by fiscal fund supervision are becoming increasingly complex, such as related enterprises circumventing approval thresholds by splitting payments across domains and fraudulently obtaining funds by exploiting semantic conflicts between payment reasons and budget declarations. Traditional manual review and fixed rule detection are no longer sufficient to meet the demands of massive data processing and the challenges of identifying hidden risks.
[0003] Traditional risk identification accuracy detection methods rely on fixed rules, which cannot adapt to dynamically changing risk patterns and are difficult to identify complex risks such as semantic conflicts and hidden related transactions (such as kinship between personnel and payment recipients). They also lack cross-domain risk detection mechanisms and have weak ability to identify cross-provincial and cross-level payment risks (such as cross-provincial split payments), which are prone to missed or misjudgments. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides a deep learning-based system and method for detecting anomalies in centralized treasury payments, which solves the problems of difficult identification of complex payment data and hidden risks in the fiscal field, such as the transfer of benefits.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a deep learning-based method for detecting anomalies in centralized treasury payments, comprising the following steps: S1 Multi-Source Data Perception: Collects payment application data, internal data, external data, and unstructured data, cleans, standardizes, and integrates them to build a fiscal risk control data lake and generate multimodal features; S2 Intelligent Analysis: It integrates multimodal deep learning, pre-trained models with enhanced knowledge in the financial domain, dynamic temporal knowledge graphs and graph neural networks, federated learning and explainable AI technologies to learn payment behavior representations, identify semantic conflicts, correlations and cross-domain risk patterns, and output anomaly scores and explanations. S3 Decision Application: Real-time risk warning, human-machine collaborative review, risk source analysis and global situation visualization based on intelligent analysis results; S4 Federation Collaboration: Treasury nodes at all levels exchange encrypted model parameters under privacy protection, collaboratively train a global risk control model, and achieve cross-domain joint prevention and control; S5 Model Optimization: Combining feedback from manual review, the model parameters and knowledge graph are continuously updated through online incremental learning to improve detection accuracy.
[0006] As a further aspect of the present invention: the multi-source data perception in step 1 includes: capturing payment application data in real time through CDC technology, periodically extracting or querying internal data via API, collecting external data through API calls or web crawlers, extracting unstructured data text using OCR, and then generating a sample set through format standardization, value range cleaning, entity association, and feature engineering.
[0007] As a further aspect of the present invention: in step 2, the multimodal deep learning adopts a fusion model of CNN, LSTM and attention mechanism to extract spatial local patterns and temporal dependencies respectively, weighted fuse multimodal features, generate a deep representation of payment behavior and calculate anomaly scores.
[0008] As a further aspect of the present invention: the pre-trained model for enhancing fiscal domain knowledge in step 2 is a fiscal BERT, which achieves semantic understanding of payment reasons, compliance judgment and semantic conflict detection through pre-training on fiscal corpus and multi-task fine-tuning.
[0009] As a further aspect of the present invention: in step 2, the dynamic temporal knowledge graph constructs a five-dimensional network of units, personnel, projects, payment recipients, and time, and uses T-GNN to learn the dynamic representation of nodes to mine risk patterns such as related transactions and transfer of benefits.
[0010] As a further aspect of the present invention: in step 2, the AI can be explained to analyze the feature contribution degree through SHAP, generate counterfactual samples and compare risk changes, and output a report containing natural language explanations, visualization charts and optimization suggestions.
[0011] As a further aspect of the present invention: in step 4, the federal collaboration involves each level of treasury node training the model locally and then uploading encrypted parameters. The central server securely aggregates and generates a global model and distributes and updates it.
[0012] A deep learning-based anomaly detection system for centralized treasury payments includes a multi-source data perception layer, an intelligent analysis engine layer, a decision application layer, and a federated collaboration platform. The multi-source data perception layer is used to collect and process multi-source heterogeneous data to generate multimodal features; The intelligent analysis engine layer deploys modules such as multimodal deep learning and pre-trained models in the financial field to perform anomaly detection and analysis; The decision application layer implements functions such as risk warning and collaborative review. The federal collaboration platform connects treasury nodes at all levels and supports collaborative risk control under privacy protection.
[0013] As a further aspect of the present invention: the intelligent analysis engine layer includes a multimodal deep learning module, a fiscal BERT module, a T-GNN module, a federated learning module, and an interpretable AI module, with each module working together to perform anomaly detection and analysis.
[0014] As a further aspect of the present invention: the decision application layer sets up an early warning level classification mechanism, outputs four levels of early warning (blue, yellow, orange, and red) based on risk scores, and simultaneously generates evidence chains and handling suggestions to support the tracking of the review process and risk tracing.
[0015] This invention provides a deep learning-based system and method for detecting anomalies in centralized treasury payments. Compared with existing technologies, it has the following advantages: (1) This invention uses multi-source data perception technology to integrate structured, unstructured and other types of data to build a fiscal risk control data lake, providing comprehensive data support for risk identification. At the same time, by combining multimodal deep learning, fiscal BERT, and T-GNN, it can accurately identify hidden risks such as semantic conflicts and related transactions, significantly improving the accuracy of anomaly detection and effectively solving the problem that traditional fixed rules cannot adapt to dynamic risk patterns. (2) This invention introduces interpretable AI technology, analyzes feature contribution through SHAP, generates counterfactual samples, outputs a report with natural language explanations and visualization charts, clearly presents the basis for anomaly judgment, effectively assists manual review and decision-making, realizes efficient risk control through human-machine collaboration, and based on the federated learning federated collaboration mechanism, the national treasury nodes at all levels only exchange encrypted model parameters, without sharing the original data. While ensuring data privacy, it builds a global risk control model, realizes cross-domain risk pattern sharing and joint prevention and control, and improves the overall risk control scope and efficiency. (3) This invention combines manual review feedback and continuously updates model parameters and knowledge graphs through online incremental learning to adapt to business changes and new risks in a timely manner, ensuring long-term stability of model detection accuracy and realizing continuous optimization and iteration of the risk control system. Attached Figure Description
[0016] Figure 1 This is a flowchart of the method of the present invention; Figure 2 This is a flowchart of the system of the present invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] Example 1 Please see Figure 1 This application provides a deep learning-based method for detecting anomalies in centralized treasury payments, including: S1: Multi-source data perception 1.1 Data Acquisition: First, data from multiple sources is collected by connecting to the integrated finance system and external databases via a data bus, as follows: Payment application data: CDC (Change Data Capture) technology is used to capture database change logs in real time, obtaining structured fields such as amount (A), payee (P), budget item (S), and application time (T), with latency controlled within 100ms to ensure data timeliness.
[0019] Internal data: Budget indicator data (such as annual budget amount B) and fund usage plan data are extracted in batches through scheduled tasks (2:00 AM every day). At the same time, historical payment record database (data of the past 3 years), project database (including project type C and budget application content D) and budget unit information are queried in real time through API interface.
[0020] External data: Obtain the equity structure (such as shareholding ratio R) and registered capital (Cpt) of the payment recipient by calling third-party business registration APIs; collect negative information (such as credit record F) of the payment recipient by using public opinion APIs; and obtain procurement contract filing data by connecting to the government procurement cloud platform.
[0021] Unstructured data: Directly extract the application reason text (Text), and use OCR technology (recognition accuracy ≥98%) to parse attachments (such as scanned copies of contracts) and extract key information (such as contract amount A_ctr, acceptance status St).
[0022] 1.2 Data Processing The above data is processed in four steps: cleaning, standardization, correlation fusion, and feature engineering. Cleaning: Filter outliers based on business rules, such as non-negative amount verification (A≥0) and time validity verification (T≤current system time). For missing fields (such as missing budget items), use mean filling + manual review to handle them, and keep the missing rate below 0.5%.
[0023] Standardization: Unify data format, convert time to ISO8601 format (e.g., 2025-10-17T09), convert amount to Decimal type (keeping 2 decimal places), and unify encoding to UTF-8; map value range based on fiscal data dictionary, such as mapping bank transfer to 01 and wire transfer to 02 in the payment method field.
[0024] Key integration: Payment applications are linked to project database data via primary and foreign keys (e.g., project ID_proj); fuzzy matching algorithms (e.g., Jaccard similarity ≥ 0.8) are used to align payee names with company names in business registration information to complete payee information; a globally unique Application_ID is generated for each payment application to build a wide table of data centered on payment applications.
[0025] Feature engineering: generating multimodal features, including: Numerical characteristics: Amount as a percentage of budget (R_A = A / B, e.g., A = 1.8 million yuan, B = 2 million yuan, then R_A = 0.9), historical payment frequency (F_p = number of payments made by the unit in the last 30 days / 30, e.g., 6 payments made in the last 30 days, then F_p = 0.2 times / day). Text features: Extract keywords from the text of the event (e.g., project payment for equipment purchase) using the TF-IDF algorithm; Graph characteristics: Initially construct a relationship graph of unit-project-payee, and record the number of node connections (e.g., the number of projects associated with unit U1 N_proj=5).
[0026] 1.3 Sample Construction Multimodal features are combined into a feature vector X = [R_A, F_p, TF-IDF (Text), N_proj]. Sample labels are annotated based on the historical review structure (normal = 0, abnormal = 1), such as a payment application in 2024 with "the reason does not match the budget application" being labeled as "1". At the same time, features are aligned with the application status at the time of application according to the timestamp to avoid future information leakage. Finally, a training sample set (sample size ≥ 100,000) is generated, stored in the feature store, and version management is supported.
[0027] S2: Intelligent Analysis 2.1 Multimodal Deep Learning Modeling A fusion model combining CNN, LSTM, and attention mechanisms is used to learn payment behavior representations. Step 1: Feature Embedding Numerical feature standardization: The standardization formula is as follows: Where μ is the characteristic mean (e.g., the average payment amount of a certain unit in the past 3 years μ = 1.2 million yuan). Standard deviation (e.g.) =300,000 yuan), if a certain payment amount X = 1,800,000 yuan, then X norm = (180-120) / 30 = 2, which means that the amount is 2 standard deviations above the mean and needs to be closely monitored.
[0028] Categorical feature embedding: Discrete features such as budget item (S) and project type (C) are mapped to a 128-dimensional dense vector through an embedding layer. For example, S = education expenditure is mapped to [0.12, 0.35].
[0029] Text feature embedding: Input the fiscal BERT model (see 2.2) to obtain a 768-dimensional semantic vector of the cause text.
[0030] Step 2: Feature Extraction CNN branch: 1D-CNN (kernel size = 3, number = 64) is used to extract local combination patterns of numerical and categorical features (such as the combination feature of education expenditure + amount ratio of 0.9), and key spatial features are preserved through max pooling layers.
[0031] LSTM branch: This involves processing a sequence of historical payment transactions (e.g., X of the last 10 payments). norm Input an LSTM network (hidden layer dimension = 128, number of layers = 2) to learn temporal dependencies, such as capturing the abnormal trend of increasing payment amounts of the last 10 transactions of a certain unit.
[0032] Step 3: Feature Fusion The weights of each feature branch are calculated using a multi-head attention mechanism, as shown in the formula: Feature fusion result = (weights obtained by standardizing the similarity between the query feature and the identifiers of each branch feature) multiplied by the actual content of each branch feature and summed; Among them, the query features are: target features that need to be analyzed in detail (such as abnormal fluctuations in amount), which are used to actively query the degree of correlation between other features and them; Branch feature identifier: The identity label of each feature branch to be fused (such as time series trend, semantic conflict), used to calculate the correlation degree of the query feature; The actual content of the branch features: The specific values of each feature branch (such as the quantitative value of the increasing amount of the last 3 payments) are the original data that ultimately participates in the fusion; Similarity: Measures the degree of association between the identifiers of the query feature and the branch feature (the higher the value, the more important the branch feature is to the query feature). Standardized weights: Convert similarity into a ratio between 0 and 1 (the sum of the weights of all branches is 1), clarifying the contribution percentage of each feature branch.
[0033] For example: Branch 1 (Time Series Characteristics); Branch 2 (semantic features); The similarity between the query feature and the branch 1 identifier is 0.3; The similarity between the query feature and the branch 2 identifier is 0.7; The total similarity is 0.3 + 0.7 = 1, therefore the weight of branch 1 is 0.3 and the weight of branch 2 is 0.7. The fused structure = (branch 1 weight × branch 1 content) + (branch 2 weight × branch 2 content) = (0.3 × 0.6) + (0.7 × 0.8) = 0.18 + 0.56 = 0.74, which means that the fused feature value is 0.74. The weight of the semantic feature (branch 2) (0.7) is significantly higher than the weight of the temporal feature (branch 1) (0.3), indicating that in this application, semantic conflict has a greater impact on the abnormal judgment of over-budget payment.
[0034] Step 4: The fused deep representation vector is input into the fully connected layer, and the anomaly score (0-1) is output through the Sigmoid function. Wherein, Score is the anomaly score, ranging from [0,1]. The closer to 1, the higher the risk of the payment application; the closer to 0, the lower the risk. e is the natural constant (approximately 2.718), which is the base of the exponential function. W is the weight matrix, used to measure the importance of each input feature (such as amount ratio, semantic conflict, correlation, etc.) to anomaly judgment (the larger the weight, the greater the impact of the feature on the result). b is the bias term. WX+b is the linear combination structure of the model with respect to the input features, which can be understood as the "raw risk score". Its value range is (-∞, +∞), and it is mapped to the [0,1] interval through the Sigmoid function for easy and intuitive judgment of risk level.
[0035] For example Suppose the feature vector X and model parameters of a payment application are as follows: Feature vector X: [Amount percentage = 0.9, Semantic conflict = 0.8, Association = 0.7] (all are standardized feature values); Weight matrix W: [0.3, 0.5, 0.4] (indicating that semantic conflict features have the highest weight and the greatest impact on anomaly detection); Bias term b: -0.2; Step 1: Calculate the linear combination result WX+b=(0.3×0.9)+(0.5×0.8)+(0.4×0.7)-0.2=0.27+0.4+0.28-0.2=0.75 Step 2: Score = 1 / (1+e) -0.75If the value is approximately 0.68, it means that the abnormal value of this payment application is 0.68.
[0036] 2.2 Pre-trained model for enhancing knowledge in the fiscal domain Step 1: Corpus Construction: Collect financial laws and regulations, budget and final accounts reports, and texts of payment reasons from the past 5 years (a total of 1 million entries), and clean them to form a corpus in the field of finance.
[0037] Step 2: Model pre-training: Based on the general BERT model, continue pre-training using fiscal corpus. Use the MLM (masked language model) task to mask fiscal terms (such as centralized treasury payment) with a 15% probability, allowing the model to learn the semantics of the terms; at the same time, learn the paragraph logic of fiscal documents through the NSP (next sentence prediction) task.
[0038] Step 3: Task Fine-tuning: Fine-tuning the three tasks: cause classification, compliance judgment, and semantic conflict detection. Semantic conflict detection: Input: Reason = XX project payment and budget application content = equipment purchase; Model output conflict score (0-1), formula is: Conflict score = 1 - semantic similarity (absolute value) between the payment reason text and the budget declaration content; Semantic similarity is used to measure the degree of similarity in meaning between two texts (the value ranges from 0 to 1, with the closer to 1 indicating greater similarity in meaning and the closer to 0 indicating greater difference in meaning). The conflict score ranges from 0 to 1: the closer the score is to 1, the more obvious the semantic conflict between the two texts; the closer the score is to 0, the more consistent the semantics.
[0039] For example Take a payment application from a budget unit as an example: 1. Payment Reason Text: Payment for equipment purchase for Project XX; 2. Budget application details (D): Project XX is a house renovation project, and the budget is for the purchase of building materials; Step 1: Convert the two texts into semantic vectors using the fiscal BERT model and calculate their semantic similarity. Since the meanings of "equipment procurement" and "house repair + building materials procurement" are quite different, the semantic similarity calculation result is 0.2.
[0040] Step 2: Calculate the conflict score using the formula expressed in the text: Conflict score = 1 - 0.2 = 0.8.
[0041] A conflict score of 0.8 (close to 1) indicates that there is a significant semantic conflict between the reason for the payment application and the content of the budget declaration. The model marks it as a high-risk feature and includes it in the subsequent anomaly scoring calculation (such as in the anomaly scoring in step 2.1, this feature may receive a higher weight, pushing up the overall anomaly score).
[0042] 2.3 Dynamic Temporal Knowledge Graphs and Graph Neural Networks Step 1: Graph Construction: Construct a five-dimensional graph of Unit (U) - Personnel (H) - Project (Proj) - Payment Recipient (P) - Time (T), such as U1 Unit - H1 Personnel (Supervisor) - Proj1 Project - P1 Company (Supervisor is a relative of H1) - T20251017, and record the relationship attributes (such as the kinship weight between H1 and P1 = 0.9).
[0043] Step 2: Use a temporal graph neural network (T-GNN) to learn the dynamic representation of nodes. The formula is: the feature vector of the node at the current time step = activation function (weight 1 × the feature vector of the node at the previous time step + weight 2 × the average of the feature vectors of the neighboring nodes at the previous time step + bias term). Among them, the feature vector of a node is used to describe the attributes and status of a node (such as the payment object P1, the person H1) (such as P1's registered capital, credit record, H1's approval authority, etc.), and is the core data for model learning.
[0044] Previous time: refers to the time point before the current time in the time series (e.g., if the current time T is 20251017, the previous time T-1 is 20251016).
[0045] Neighboring nodes: These refer to other nodes that are directly associated with the current node (e.g., P1's neighbors may include the Proj1 project and H1 personnel). Weight 1 and Weight 2 are parameters obtained from model training, which respectively measure the influence of the node's own historical features and the historical features of neighboring nodes on the current node's state (the larger the weight, the more significant the influence).
[0046] Activation function: used to introduce non-linear relationships (commonly ReLU function), allowing the model to learn complex association patterns (such as the combined risk of kinship + high-frequency cooperation).
[0047] Bias term: Adjusts the baseline value of the model output to ensure that the calculation results are more in line with actual business rules.
[0048] For example: The feature vector of node P1 at the previous time step (T-1 = 20251016) includes: registered capital = 10 million yuan, number of historical cooperations = 3 times, and record of dishonesty = 0 records, which are quantized as [0.6, 0.4, 0.1]. P1's neighboring nodes: Project Proj1 and H1 personnel (supervisor), a total of two neighbors.
[0049] The feature vector of Proj1 at the previous time step: [0.8, 0.3] (budget execution progress = 80%, project risk level = low); H1's feature vector at the previous time step: [0.9, 0.2] (Approval authority = high, kinship record = 1 time); Model parameters: weight 1 = 0.2, weight 2 = 0.3, bias term (b) = 0.1, activation function is ReLU (values less than 0 are set to 0).
[0050] Calculation steps: Step 1: Calculate the weighted values of the node's own features at the previous time step: Weight 1 × eigenvector of node P1 at the previous time step = 0.5 × [0.6, 0.4, 0.1] = [0.3, 0.2, 0.05]; Step 2: Calculate the average and weighted values of the neighboring node features from the previous time step: The average neighbor feature = (Proj1 feature + H1 feature) ÷ 2 = ([0.8, 0.3] + [0.9, 0.2]) ÷ 2 = [0.85, 0.25]; Weight 2 × average value = 0.3 × [0.85, 0.25] = [0.255, 0.075]; Step 3: Combine the above results and add a bias term: The sum = [0.3, 0.2, 0.05] + [0.255, 0.075] + 0.1 = [0.655, 0.375, 0.15] (when the vector dimensions are inconsistent, they are unified through internal model mapping); Step 4: Output the feature vector at the current time step using the ReLU activation function: Since all values are greater than 0, the final feature vector of the node at the current time step is [0.655, 0.375, 0.15].
[0051] Step 3: Risk Mining: Identify related transaction circles (such as a closed-loop network formed by U1-Proj1-P1-P2) using the community discovery algorithm (Louvain algorithm), and calculate the circle risk score (if there are 3 semantically conflicting payments within the circle, the risk score = 0.85).
[0052] 2.4 Explainable AI Analysis: Step 1: Feature Contribution Analysis: The contribution of each feature to the outlier score is calculated using the SHAP value. The formula is: Contribution of a feature (SHAP value) = Difference between the model output of all feature subsets that do not contain the feature and feature subsets that contain the feature, multiplied by the weights of the corresponding subsets and then summed. Among them, the feature subset is a set of features selected from all features (such as selecting the proportion of amount, semantic conflict, and association relationship from the proportion of amount and semantic conflict). Subsets that do not contain this feature: All possible subsets that do not contain the current analysis feature (e.g., when analyzing association relationships, the subsets are only amount percentage, only semantic conflict, and amount percentage + semantic conflict). Model output difference: The difference in outlier scores output by the model when the same subset includes the feature and when it does not (the larger the difference, the more significant the impact of the feature on the result). Weight: Calculated based on the number of features contained in the subset (the closer the number of features is to half of the total number of features, the higher the weight), used to balance the influence of different subsets.
[0053] For example, let's take the semantic conflict feature (denoted as B) as an example and calculate its SHAP value: 1. Given conditions: Overall features: Amount percentage (A), semantic conflict (B), and association (C), totaling 3 features; Feature subsets that do not contain B: S1 (empty set), S2 ({A}), S3 ({C}), S4 ({A,C}); Weights of each subset (calculated based on the number of features): S1=1 / 6, S2=1 / 6, S3=1 / 6, S4=1 / 2; Model output differences (including B vs. not including B): S1: 0.2 (including B) - 0.1 (excluding B) = 0.1; S2: 0.6 (including B) - 0.3 (excluding B) = 0.3; S3: 0.5 (including B) - 0.2 (excluding B) = 0.3; S4: 0.9 (including B) - 0.5 (excluding B) = 0.4; Calculation process: SHAP value (B) = (0.1×1 / 6) + (0.3×1 / 6) + (0.3×1 / 6) + (0.4×1 / 2) = 0.017 + 0.05 + 0.05 + 0.2 = 0.317. The SHAP value of the semantic conflict feature is 0.317, which means that the contribution of this feature to the anomaly score is 0.317 (if the total anomaly score is 0.9, the contribution ratio is about 35%).
[0054] Step 2: Generate counterfactual samples. For example, change the payee P1 (H1 relative) to P2 (unrelated), recalculate the anomaly score. If the original score is 0.92 and the modified score is 0.28, then the risk reduction is approximately (0.92-0.28) / 0.92 ≈ 70%. Generate an explanation: If the payee is changed to the unrelated company P2, the risk value will decrease by 70%.
[0055] Step 3: Report Generation: Output a natural language report containing a bar chart of SHAP values and a counterfactual comparison table, clearly identifying the anomaly type (such as suspected fraudulent projects to obtain funds) and the basis for judgment.
[0056] 2.5 Federated Learning Collaborative Analysis Step 1: Task initialization: The central coordinating node defines the federated task (anomaly detection model training), distributes the initial model (same as the CNN-LSTM model in 2.1) to 31 provincial treasury nodes, agrees to use FedAvg as the aggregation algorithm, and adopts homomorphic encryption as the security protocol.
[0057] Step 2: Local Training: Each provincial node trains the model using local data (e.g., the Jiangsu provincial node uses payment data from the past two years, sample size = 50,000 records). The SGD optimizer (learning rate = 0.001, batch size = 32) is used for 10 iterations to obtain the local model parameters W. i (i is the node number).
[0058] Step 3: Parameter Encryption Upload: Each node encrypts W using the central node's public key. i Upload to the central node, and simultaneously upload local data of size n. i (e.g., Jiangsu Province n) i =50,000 entries).
[0059] Step 4: Secure aggregation: The central node aggregates global parameters by weighting the data volume. The formula is: Global model parameter = (sum of local data volume of each node × local model parameter of that node) ÷ total local data volume of all nodes; The local data volume is the number of samples used by each node (such as the provincial treasury) to train the model (such as the number of payment application records in a certain province). Local model parameters: Model parameters (such as related transaction feature weights, semantic conflict feature weights, etc.) are trained on local data for each node. For example: Beijing node: Local data volume = 30,000 records, related transaction feature weight in the local model = 0.7; Jiangsu node: Local data volume = 50,000 records, related transaction feature weight in the local model = 0.6; Guangdong node: Local data volume = 40,000 records, related transaction feature weight in the local model = 0.5; Step 1: Calculate the sum of data volume of each node × local parameters: (30,000 × 0.7) + (50,000 × 0.6) + (40,000 × 0.5) = 21,000 + 30,000 + 20,000 = 71,000; Step 2: Calculate the total local data volume of all nodes: 30,000 + 50,000 + 40,000 = 120,000; Step 3: Calculate global model parameters: The global related transaction feature weight = 7.1 ÷ 12 ≈ 0.592. The related transaction feature weight in the global model is 0.592. This value is closer to the local parameter (0.6) of the Jiangsu node (50,000 data entries) with the largest data volume. At the same time, it integrates the experience of Beijing (high weight) and Guangdong (low weight), which not only reflects the business characteristics of large data volume nodes, but also takes into account the risk identification rules of different regions, enabling the global model to more accurately identify related transaction risks across regions. Step 5: Model Distribution: The central node encrypts the global model parameters and distributes them to each node. Each node updates its local model to achieve cross-domain risk pattern sharing (such as identifying abnormal patterns of cross-provincial related enterprises misappropriating funds).
[0060] S3: Decision Application 3.1 Real-time risk warning Triggering mechanism: After a payment application is submitted, the anomaly detection model in S2 is invoked in real time (response time ≤ 1 second), and the warning level is determined based on the anomaly score. Blue alert Score∈[0,0.3]: Normal application, automatically approved; Yellow alert Score∈[0.3,0.6]: Low risk, indicating that the reviewers should pay attention to characteristics such as a high proportion of the amount involved; Orange alert Score∈[0.6,0.8]: Medium risk, automatic approval suspended, key evidence (such as scanned copies of contracts) requires manual review. Red alert Score∈[0.8,1]: High risk, automatic blocking, and generation of a suspension payment suggestion.
[0061] 3.2 Human-Machine Collaborative Audit Review interface: Displays the warning level, anomaly score, and interpretable report to reviewers. For example, in red warning cases, it highlights key evidence such as semantic conflict (Conflict=0.8) and correlation (SHAP=0.4), and provides the function of viewing the graph path and retrieving historical cases.
[0062] Handling procedures: Supports reviewers to approve, return for modification, or reject the application. The results are recorded in real time (e.g., on October 17, 2025 at 10:30, reviewer Zhang XX rejected the application because: no project acceptance report was provided).
[0063] 3.3 Risk Origin Tracing and Situation Visualization Source tracing analysis: For confirmed abnormal cases (such as fraudulent projects to obtain funds), the risk transmission path is traced through knowledge graph (such as U1 unit → H1 personnel → P1 company → Proj1 project), and related anomalies are discovered (such as P1 company being involved in 3 other similar projects).
[0064] Visual dashboards: Build a global risk situation dashboard to display: Real-time metrics: Current number of alerts (e.g., 5 red alerts, 23 orange alerts), approval rate (e.g., 92%); Historical trends: Changes in the percentage of each warning level over the past 30 days (e.g., the percentage of red warnings increased from 1% to 3%, the reasons need to be analyzed); Regional distribution: Ranking of anomaly rates in each province (e.g., a province with an anomaly rate of 5%, higher than the national average of 2%).
[0065] S4: Federal Collaboration 4.1 Node Registration: Treasury nodes at all levels (central, provincial, municipal, and county) register with the federal collaboration platform, submit node information (such as name and IP address), and authenticate their identity through digital certificates. Only nodes that pass authentication can participate in federal tasks.
[0066] Network topology: A three-tier topology structure of central, provincial, and municipal nodes is adopted. The central node is responsible for task coordination and parameter aggregation, while the provincial nodes connect to the municipal nodes to ensure clear data transmission hierarchy (e.g., municipal nodes only upload parameters to the provincial nodes within their own province and do not interact directly with the central node).
[0067] 4.2 Collaborative Risk Control Risk pattern sharing: Under the premise of privacy protection, the central node shares the cross-domain risk pattern library with each node (such as the cross-provincial split payment pattern: a company splits 3 payments of 500,000 yuan each in province A and 2 payments of 500,000 yuan each in province B to circumvent the single-province approval threshold of 1 million yuan), without sharing the original data.
[0068] Cross-regional early warning and linkage: If a provincial node detects that a company is involved in cross-provincial risk patterns, it will send an early warning notification to related provincial nodes through the federal platform (e.g., Province A sends an early warning of abnormal payment by Company XX to Province B). The related provincial nodes will give priority to reviewing the company's application to achieve joint prevention and control.
[0069] S5: Model Optimization 5.1 Feedback Data Collection Collect feedback from manual reviewers: Add the corrective results of the reviewers' judgments on the model (such as the model being marked as low risk when it is actually abnormal) to the sample library as new labels.
[0070] Collect system operation data: record model accuracy (e.g., 90% of red alert cases are confirmed as abnormal, accuracy = 90%) and false positive rate (e.g., 5% of normal applications are marked as orange alert, false positive rate = 5%).
[0071] 5.2 Incremental Model Learning An online incremental learning algorithm (such as FTRL) is used to update the model parameters periodically (every Sunday morning) with newly added samples (e.g., 10,000 new labeled samples). The formula is as follows: New parameter = Old parameter - 0.0005 × g Where 0.0005 is the learning rate and g is the gradient of the new sample.
[0072] 5.3 Knowledge Graph Update Real-time updates of graph entities and relationships: If a payment recipient has a new record of dishonesty (F=1), its graph attributes are updated; if a new person-payment recipient relationship is discovered (such as H2 person having a shareholding relationship with P2 company), graph edges are added to ensure that the graph dynamically reflects the actual business relationships.
[0073] 5.4 Effect Evaluation and Iteration Monthly evaluation of model performance: calculate accuracy (target ≥ 95%), recall (target ≥ 90%), and F1 score (target ≥ 92%). If a certain indicator fails to meet the target (e.g., recall = 88%), analyze the reasons (e.g., not covering new split payment models), adjust the model structure (e.g., increase the LSTM time window length) or supplement training data.
[0074] Example 2 Please see Figure 1 This illustrates another embodiment of the present invention, which is largely the same as the technical solution of Embodiment 1, so only the differences are described.
[0075] A deep learning-based method for detecting anomalies in centralized treasury payments, comprising: S1: Multi-source data perception: When the treasury payment anomaly detection process is initiated, the system automatically triggers the data collection process, using diverse technologies to cover all relevant data. CDC (Change Data Capture) technology is used to capture structured data (such as applicant unit, amount, payment recipient, etc.) from the payment application system in real time to ensure that the data is synchronized with the business. Obtain relevant internal and external data (internal such as budget indicators and historical payment records, external such as corporate credit information and government procurement information) through API interfaces or scheduled crawlers to supplement the risk assessment dimensions. OCR technology is used to extract text information from unstructured documents such as contracts and invoices, transforming paper information into analyzable data.
[0076] The collected data is cleaned (outliers removed, missing items filled), standardized (unified format and encoding), and correlated and fused (multi-source information is linked through core keys such as payment application number and unit code). Finally, multimodal features containing numerical, text, time series, and image types are generated, and a fiscal risk control data lake is constructed to provide high-quality raw materials for subsequent analysis.
[0077] S2 Intelligent Analysis: Based on the multimodal features generated by S1, the system initiates multi-module parallel analysis to uncover risks from different dimensions: Multimodal deep learning: CNN is used to extract spatial features of invoice images (such as the authenticity of the seal), LSTM is used to capture the temporal patterns of historical payments (such as sudden changes in amount), and then multi-dimensional features are weighted and fused through an attention mechanism to calculate the anomaly score of the current payment behavior (the higher the score, the higher the risk). Fiscal BERT: Through pre-training with fiscal regulations and historical cases and fine-tuning with business scenarios, it accurately understands the semantics of payment reasons (such as the expression characteristics of fictitious expenditures), judges compliance (such as whether the budget is exceeded), and detects text conflicts (such as contradictions between the reason and the contract content). GNN: Construct a five-dimensional knowledge graph of units, personnel, projects, payment recipients, and time, and use time-aware graph neural networks to mine hidden connections (such as the same approver making continuous payments to related companies) to identify hidden risks such as the transfer of benefits. Explainable AI: Analyzes the contribution of each feature to the anomaly score through SHAP values (e.g., 80% over budget is a major risk point), generates counterfactual samples (e.g., if the amount is reduced to within budget, the risk score drops by 50 points), and finally outputs a risk report with natural language explanations and visual charts.
[0078] After the results from multiple modules are integrated, a final anomaly score and risk explanation report are generated, providing a clear conclusion for decision-making regarding what the risk is and why it exists.
[0079] S3 Decision Applications: Based on the anomaly scores and reports output by S2, the system proceeds to the decision execution phase: The system uses a four-level warning system based on scores: blue (low risk, automatic approval), yellow (low to medium risk, random checks), orange (high to medium risk, key review), and red (high risk, payment suspension), which triggers corresponding handling procedures. Provide evidence chains (such as specific evidence of budget overruns and relationship diagrams of related companies) and handling suggestions (such as verifying the authenticity of contracts) for human-machine collaborative audits, and support the tracking of the entire audit process; The risk source map and overall situation dashboard provide an intuitive display of risk distribution (such as high-frequency anomaly types in a certain region), making it easier for managers to trace from a single anomaly to the overall risk pattern.
[0080] The system ultimately outputs early warning notifications, review results, and source tracing reports, achieving a closed loop from risk identification to handling.
[0081] S4 Federation Collaboration: After S3 is completed, the system initiates the federated coordination mechanism: Each provincial, municipal, and county treasury node trains its model based on local data and only uploads the encrypted model parameters to the central server. The central server then uses a secure aggregation algorithm (such as federated averaging) to fuse the parameters, generate a nationwide global risk control model, and distributes it to nodes at all levels to update their local models.
[0082] This approach enables collaborative learning of cross-regional risk characteristics (such as common patterns of a new type of fraud in multiple locations) without sharing the original data, thereby enhancing the system's ability to identify cross-regional risks.
[0083] S5 model optimization: The system initiates online incremental learning by combining feedback from manual review (such as misjudged normal payments and missed abnormal cases) with the global model of S4; Update model parameters using first-case studies to optimize the outlier score calculation logic: Supplement the knowledge graph with entities and relationships (such as adding enterprise-related information) to improve the accuracy of related risk mining.
[0084] The optimized model is fed back into S2 to continuously improve detection accuracy, forming a complete closed loop of data, analysis, decision-making, feedback, and optimization, ensuring that the system adapts to business changes such as policy adjustments and new risks.
[0085] Example 3 Please see Figure 2 This illustrates another embodiment of the present invention, which is largely the same as the technical solutions of Embodiment 1 and Embodiment 2, so only the differences are described.
[0086] A deep learning-based anomaly detection system for centralized treasury payments includes: Multi-source data perception layer: This layer is responsible for data collection and preprocessing. It collects full data through tools such as CDC, API, web crawler, and OCR. After cleaning, standardization, correlation and fusion, it generates multimodal features and transmits them to the upper intelligent analysis engine layer. Its core function is to transform scattered and heterogeneous raw data into usable and reliable analytical materials.
[0087] Intelligent Analysis Engine Layer: This layer is the core processing unit, deploying five major modules: multimodal deep learning, fiscal BERT, T-GNN, federated learning, and interpretable AI, corresponding to the analysis logic of S2; The system receives multimodal features from the perception layer and calculates anomaly scores and risk interpretations through collaborative calculations across various modules. At the same time, it outputs local model parameters to the federal collaboration platform and receives global model updates.
[0088] Decision application layer: This layer transforms the analysis results into actionable functions: Based on abnormal scores, a four-level early warning system is displayed to intuitively indicate the risk level. Provides human-machine collaborative review tools, displays the chain of evidence and handling suggestions, and supports process tracking; The overall risk situation is presented through a visual dashboard, which supports risk tracing and querying.
[0089] Meanwhile, the feedback from manual review is sent back to the intelligent analysis engine layer as a basis for model optimization, thus realizing a closed loop of human-machine collaboration.
[0090] Federal Collaboration Platform: A technical platform connecting treasury nodes at all levels, responsible for encrypted parameter transmission, secure aggregation, and model distribution; Receive encrypted model parameters from each node to prevent leakage of raw data; Generate a global model and distribute it to allow nodes at all levels to share cross-domain risk characteristics.
[0091] Some of the data in the above formulas are numerical calculations with dimensions removed, and the contents not described in detail in this specification are all prior art known to those skilled in the art.
[0092] The above embodiments are only used to illustrate the technical methods of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical methods of the present invention without departing from the spirit and scope of the technical methods of the present invention.
Claims
1. A method for detecting anomalies in centralized treasury payments based on deep learning, characterized in that, Includes the following steps: S1 Multi-Source Data Perception: Collects payment application data, internal data, external data, and unstructured data, cleans, standardizes, and integrates them to build a fiscal risk control data lake and generate multimodal features; S2 Intelligent Analysis: It integrates multimodal deep learning, pre-trained models with enhanced knowledge in the financial domain, dynamic temporal knowledge graphs and graph neural networks, federated learning and explainable AI technologies to learn payment behavior representations, identify semantic conflicts, correlations and cross-domain risk patterns, and output anomaly scores and explanations. S3 Decision Application: Real-time risk warning, human-machine collaborative review, risk source analysis and global situation visualization based on intelligent analysis results; S4 Federation Collaboration: Treasury nodes at all levels exchange encrypted model parameters under privacy protection, collaboratively train a global risk control model, and achieve cross-domain joint prevention and control; S5 Model Optimization: Combining feedback from manual review, the model parameters and knowledge graph are continuously updated through online incremental learning to improve detection accuracy.
2. The method for detecting anomalies in centralized treasury payments based on deep learning according to claim 1, characterized in that, The multi-source data perception in step 1 includes: capturing payment application data in real time through CDC technology, obtaining structured fields such as amount, payee, budget item, and application time, periodically extracting or querying internal data via API, collecting external data via API calls or web crawlers, extracting unstructured data text using OCR, and then generating a sample set through format standardization, value range cleaning, entity association, and feature engineering.
3. The method for detecting anomalies in centralized treasury payments based on deep learning according to claim 1, characterized in that, In step 2, the multimodal deep learning adopts a fusion model of CNN, LSTM and attention mechanism to extract spatial local patterns and temporal dependencies, respectively, and weightedly fuse multimodal features to generate a deep representation of payment behavior and calculate anomaly scores.
4. The method for detecting anomalies in centralized treasury payments based on deep learning according to claim 1, characterized in that, In step 2, the pre-trained model for enhancing knowledge in the fiscal domain is the Fiscal BERT. Through pre-training on fiscal corpora and fine-tuning with multiple tasks, it achieves semantic understanding of payment reasons, compliance judgment, and semantic conflict detection.
5. The method for detecting anomalies in centralized treasury payments based on deep learning according to claim 1, characterized in that, In step 2, the dynamic temporal knowledge graph constructs a five-dimensional network of units, personnel, projects, payment recipients, and time, and uses T-GNN to learn the dynamic representation of nodes to explore risk patterns such as related transactions and transfer of benefits.
6. The method for detecting anomalies in centralized treasury payments based on deep learning according to claim 1, characterized in that, In step 2, the AI can be explained to analyze the contribution of features through SHAP, generate counterfactual samples and compare risk changes, and output a report containing natural language explanations, visualization charts and optimization suggestions.
7. The method for detecting anomalies in centralized treasury payments based on deep learning according to claim 1, characterized in that, In step 4, the federal collaboration involves each level of treasury node training the model locally and uploading encrypted parameters. The central server then securely aggregates and generates a global model, which is then distributed and updated.
8. A deep learning-based treasury centralized payment anomaly detection system, capable of executing the deep learning-based treasury centralized payment anomaly detection method according to any one of claims 1-7, characterized in that, It includes a multi-source data perception layer, an intelligent analysis engine layer, a decision application layer, and a federated collaboration platform; The multi-source data perception layer is used to collect and process multi-source heterogeneous data to generate multimodal features; The intelligent analysis engine layer deploys modules such as multimodal deep learning and pre-trained models in the financial field to perform anomaly detection and analysis; The decision application layer implements functions such as risk warning and collaborative review. The federal collaboration platform connects treasury nodes at all levels and supports collaborative risk control under privacy protection.
9. A deep learning-based anomaly detection system for centralized treasury payments according to claim 8, characterized in that, The intelligent analysis engine layer includes: The system includes a multimodal deep learning module, a fiscal BERT module, a T-GNN module, a federated learning module, and an interpretable AI module. These modules work together to perform anomaly detection and analysis.
10. A deep learning-based anomaly detection system for centralized treasury payments according to claim 8, characterized in that, The decision application layer is equipped with an early warning level classification mechanism, which outputs four levels of early warning: blue, yellow, orange, and red, based on risk scores. It also generates evidence chains and handling suggestions simultaneously, supporting the tracking of the review process and risk tracing.