An invoice risk control management method and system based on multi-source data linkage

By constructing a hypergraph of associations using a multi-source data fusion gateway and graph neural networks, combined with a heterogeneous temporal Transformer model and a federated collaborative processing mechanism, the problems of data isolation and low risk identification efficiency in traditional invoice management are solved, enabling accurate identification and timely processing of false invoices and duplicate financing.

CN121120289BActive Publication Date: 2026-01-13GUANGZHOU LESHUI INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511649154.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-12
Publication Date
2026-01-13
Estimated Expiration
2045-11-12

AI Technical Summary

Technical Problem

Traditional invoice management relies on a single data source, making it difficult to identify risks such as fraudulent invoices and duplicate financing. Data processing efficiency is low, and information is isolated between departments, making it impossible to share information in a timely manner and coordinate risk management.

Method used

Data from the entire invoice business process is collected through a multi-source data fusion gateway. A hypergraph of relationships between invoices, enterprises, personnel, and assets is constructed. Graph neural networks and heterogeneous time-series Transformer models are used for minute-level analysis and multi-scale anomaly analysis. A federated collaborative handling mechanism is adopted for secure sharing and coordinated handling.

Benefits of technology

It has enabled the accurate identification and timely handling of fraudulent invoice issuance groups and risks of repeated financing, improved data processing efficiency, and ensured the stability of national tax revenue and financial markets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121120289B_ABST
    Figure CN121120289B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of invoice risk control management, and discloses an invoice risk control management method and system based on multi-source data linkage. The method realizes real-time collection of multi-source data of an invoice business whole process through a multi-source data fusion gateway. The multi-source data is subjected to flow ETL processing and encryption integration to generate an encrypted unified data view. Based on the view, a correlation hypergraph is constructed by using a dynamic graph correlation technology, and entity relationships are analyzed by means of a graph neural network to identify virtual invoice gang and repeated financing risks. Risk quantification and traceability technologies are adopted to calculate a dynamic risk score, and a risk level, a responsibility node identifier and a fund flow direction heat map are output. When the risk score exceeds a threshold value, risk gradient information is safely shared among a group headquarters, a branch company and a tax and bank institution to trigger linkage disposal actions such as invoice freezing, red letter notification or credit adjustment. According to an artificial review result, a risk model is fine-tuned and a rule library is hot-updated by using a knowledge self-updating technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of invoice risk control management technology, specifically to an invoice risk control management method and system based on multi-source data linkage. Background Technology

[0002] Invoices, as crucial financial and tax documents in economic activities, hold a paramount position. From the perspective of internal financial management, they are the legal documents for financial receipts and expenditures. Every income and expenditure is recorded based on an invoice, providing a clear picture of the company's financial situation. When purchasing raw materials, the invoices issued by suppliers detail the types, quantities, and prices of the purchased goods. Company finance personnel use this as a basis for accounting procedures and accurate cost calculation. Simultaneously, invoices are the foundation of accounting, providing indispensable data support for preparing financial statements, helping management understand operating results and make informed decisions.

[0003] In the field of tax collection and administration, invoices are a crucial basis for tax authorities to control tax sources and collect taxes. By monitoring the issuance and circulation of invoices, tax authorities ensure that enterprises pay taxes in accordance with the law and prevent tax evasion. For general VAT taxpayers, the input and output VAT amounts stated on special VAT invoices directly relate to the amount of VAT payable. Tax authorities use this information as a clue for tax audits to ensure tax fairness. Invoices are also an important means for the state to supervise economic activities, maintain economic order, and protect national property security, playing a key role in promoting the healthy and stable development of the market economy.

[0004] Traditional invoice management often relies on a single data source, which severely limits the effective identification of invoice risks. For example, while relying solely on tax bureau records can provide basic invoice issuance information, it is difficult to detect fraudulent invoicing activities concealed by complex cash flows and contractual relationships. A company may sign false contracts with multiple related parties, circulating funds among them to create the illusion of fictitious transactions while issuing corresponding invoices. Due to the lack of comprehensive analysis of multi-source data such as bank statements and logistics GPS signals, it is difficult to detect these anomalies solely from tax bureau records, allowing fraudulent invoicing to remain hidden for extended periods, seriously harming national tax revenue and market economic order. Similarly, when identifying the risk of double financing, relying on a single data source cannot provide a comprehensive understanding of a company's asset and liability situation and cash flow, making it difficult to accurately determine whether a company is using the same assets for double financing.

[0005] Traditional invoice management methods suffer from significant shortcomings in data processing efficiency, failing to achieve real-time monitoring of the entire invoice business process. In the data collection stage, reliance on manual entry or periodic batch imports is not only time-consuming and labor-intensive but also prone to data errors. After invoices are issued, it may take several days or even longer to input the relevant data into the tax system or the company's internal financial system. During the data processing stage, traditional processing technologies and processes struggle to meet the demands for rapid analysis of massive amounts of invoice data, leading to risks being detected only after they have occurred. When a company engages in fraudulent invoicing, due to data processing delays, tax authorities may only discover the problem during routine inspections months later. By this time, the company may have already incurred substantial tax losses, and those responsible may have already transferred assets, increasing the difficulty of subsequent recovery and punishment.

[0006] Information silos exist among various departments and institutions in invoice management, hindering the formation of a collaborative risk management mechanism. Enterprises, tax authorities, and banks operate independently, failing to share risk information in a timely manner. When enterprises issue invoices, tax authorities cannot obtain real-time information on bank fund flows and logistics, making it difficult to determine the authenticity and reasonableness of the invoices. When an enterprise is suspected of issuing false invoices, the lack of an effective information-sharing mechanism with banks prevents tax authorities from quickly obtaining data on the enterprise's fund flows, hindering the tracking and freezing of funds and impacting the efficiency of case investigation. Furthermore, poor communication exists between different departments within enterprises. The finance department focuses only on the accounting treatment of invoices, while the business departments are responsible for the authenticity of the business transactions behind the invoices. The lack of effective collaboration and information sharing between the two makes it difficult to promptly detect and resolve invoice risks within the enterprise. Summary of the Invention

[0007] The purpose of this invention is to provide an invoice risk control management method and system based on multi-source data linkage to solve the problems mentioned in the background art.

[0008] To achieve the above objectives, the present invention provides an invoice risk control management method based on multi-source data linkage, the method comprising:

[0009] The system collects multi-source data from the entire invoice business process in real time through a multi-source data fusion gateway, including tax bureau ledgers, bank statements, logistics GPS signals, contract blockchain evidence, industrial and commercial judicial records, and ERP image data. The system then performs streaming ETL processing and encryption integration on the multi-source data to generate an encrypted unified data view.

[0010] Based on the encrypted unified data view, a hypergraph of relationships between invoices, enterprises, personnel, and assets is constructed using dynamic graph association technology. Entity relationships are then analyzed in minutes using graph neural networks to identify fraudulent invoice groups and risks of repeated financing.

[0011] Using risk quantification and tracing techniques, a heterogeneous time-series Transformer model is used to perform multi-scale anomaly analysis on the edge weights of the associated hypergraph, calculate dynamic risk scores, and output risk levels, responsibility node identifiers, and heatmaps of fund flows.

[0012] When the dynamic risk score exceeds the predefined risk threshold, the risk gradient information is securely shared among the group headquarters, subsidiaries, and tax and banking institutions through a federal collaborative handling mechanism based on differential privacy and zero-knowledge proof, triggering coordinated handling actions such as invoice freezing, red-letter notification, or credit adjustment.

[0013] Based on the results of manual review of the coordinated response actions, the review feedback is transformed into reinforcement learning rewards through knowledge self-updating technology, and the risk model is fine-tuned online and the rule base is updated hot.

[0014] Preferably, the real-time collection of multi-source data throughout the entire invoice business process via the multi-source data fusion gateway includes:

[0015] Configure multi-source data access rules, specifying the real-time push interface for tax bureau ledger data, the timed retrieval cycle for bank transaction data, the streaming reception frequency for logistics GPS signals, the smart contract event monitoring for contract blockchain notarization, the batch update trigger conditions for industrial and commercial judicial records, and the asynchronous upload channel for ERP image data.

[0016] Perform format standardization, field mapping, and data validation on the multi-source data to eliminate semantic conflicts between heterogeneous data sources;

[0017] Homomorphic encryption is performed on the standardized multi-source data using an encryption algorithm to generate the encrypted unified data view, which is then stored in a distributed database.

[0018] Preferably, the construction of the hypergraph relating invoices, enterprises, personnel, and assets includes:

[0019] Extract invoice codes, enterprise unified social credit codes, personnel ID numbers, and asset serial numbers from the encrypted unified data view as entity nodes, and extract invoice issuance time, transaction amount, logistics trajectory, and contract terms as edge attributes;

[0020] The embedding representation of entity nodes is learned using the neighborhood aggregation function of graph neural network, and the similarity matrix between nodes is calculated.

[0021] The hypergraph structure is dynamically updated based on the similarity matrix, virtual edges are added to capture potential associations across entity types, and subgraph patterns of risky groups are identified through community detection algorithms.

[0022] Preferably, the step of parsing entity relationships using a graph neural network at the minute level includes:

[0023] The scheduling graph neural network model loads pre-trained weights, performs multiple rounds of message passing iterations on the associated hypergraph, and updates node embeddings.

[0024] An attention mechanism is used to calculate the importance score of edge weights, and high-weight edges are selected as key paths for risk association.

[0025] Entities on the critical path are clustered in real time to generate topological fingerprints of risky groups and matched with a historical risk pattern library.

[0026] Preferably, the multi-scale anomaly analysis of the edge weights of the associated hypergraph includes:

[0027] The edge weight data is converted into a time series and input into the encoder layer of the heterogeneous temporal Transformer model to capture long-term dependencies.

[0028] Anomalies at different time scales, including short-term fluctuations, medium-term trends, and long-term cycles, are calculated using a multi-head self-attention mechanism.

[0029] By fusing multi-scale anomaly features, a dynamic risk score is output using a fully connected layer, and a responsibility node identifier is generated through gradient backpropagation.

[0030] Preferably, the triggering of coordinated response actions through the federal collaborative response mechanism includes:

[0031] Under the federated learning framework, each participant's local model is initialized, and the encrypted risk gradient information is uploaded to the coordination server.

[0032] Noise is added to gradient information based on differential privacy to meet privacy budget constraints, and the authenticity of gradient data is verified by zero-knowledge proof.

[0033] After the coordinating server aggregates gradients, it issues a global model update command, triggering the local system's action executor to complete the modification of invoice status or adjustment of credit parameters.

[0034] Preferably, the security sharing risk gradient information includes:

[0035] A temporary session key is generated, the risk gradient information is symmetrically encrypted, and then transmitted to the participants through a secure channel.

[0036] After the participants decrypt the gradient information using their private keys, they perform model inference in the sandbox environment and output disposal suggestions.

[0037] Record decision logs of handling recommendations and generate audit trail records for subsequent tracing.

[0038] Preferably, the step of converting review feedback into reinforcement learning rewards through knowledge self-updating technology, and fine-tuning the risk model online and hot-updating the rule base includes:

[0039] The results of manual review are converted into reward signals, with positive rewards corresponding to correct risk identification and negative rewards corresponding to false alarms or omissions.

[0040] The policy gradient algorithm of reinforcement learning is used to adjust the parameters of the risk model to maximize the long-term reward accumulation.

[0041] The confidence threshold and entity parsing rules of the rule base are dynamically updated through an online learning framework and deployed to the production environment in real time.

[0042] Preferably, the step of converting review feedback into reinforcement learning rewards includes:

[0043] Define a reward function, where risk identification accuracy, response time, and resource consumption are multi-dimensional reward indicators;

[0044] Using a convolutional neural network to perform sentiment analysis on the review feedback text and extract implicit reward weights;

[0045] By employing reward pruning techniques, the sparse reward problem can be balanced, ensuring the model's convergence stability.

[0046] Preferably, the present invention also includes an invoice risk control management system based on multi-source data linkage, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of the above-mentioned invoice risk control management method based on multi-source data linkage.

[0047] Compared with the prior art, the beneficial effects of the present invention are:

[0048] This invention utilizes a multi-source data fusion gateway to collect multi-source data in real time across the entire invoice business process, encompassing tax bureau ledgers, bank statements, logistics GPS signals, blockchain-based contract notarization, business and judicial records, and ERP image data. These data sources are extensive, reflecting relevant information about the invoice business from different perspectives. Tax bureau ledgers record basic invoice issuance information, bank statements reflect fund flows, logistics GPS signals show the transportation trajectory of goods, blockchain-based contract notarization ensures the authenticity and immutability of contracts, business and judicial records reflect the legality of business operations and credit status, and ERP image data contains relevant image data of internal business processes. By performing streaming ETL processing and encrypted integration on this multi-source data, an encrypted unified data view is generated. This process organically combines previously scattered and isolated data, providing a comprehensive and accurate data foundation for subsequent risk analysis. It enables relevant personnel to grasp the overall picture of the invoice business, understand each stage and related factors, and thus more effectively identify potential risk points.

[0049] Based on an encrypted unified data view, a hypergraph is constructed using dynamic graph association technology to connect invoices, enterprises, personnel, and assets. This hypergraph clearly displays the complex relationships between entities, closely linking invoices with the issuing enterprises, involved personnel, and related assets. Minute-level parsing of entity relationships using graph neural networks allows for in-depth mining of potential information within these relationships. Analyzing the relationship between invoices and enterprises reveals abnormal invoice issuance and receipt patterns between companies; when researchers associate invoices with individuals, it identifies whether individuals are frequently involved in abnormal invoice transactions. This precise analysis accurately identifies fraudulent invoice issuance groups and the risk of duplicate financing. For fraudulent invoice issuance groups, analyzing invoice flows, fund transfers, and personnel connections allows for rapid identification of group members and their modus operandi, enabling timely action to combat them and prevent tax revenue losses. Regarding the risk of duplicate financing, correlation analysis of enterprise asset and financing data accurately determines whether an enterprise is using the same asset for duplicate financing, providing accurate risk warnings to financial institutions, reducing credit risk, and ensuring the stability of the financial market.

[0050] This study employs risk quantification and attribution techniques, using a heterogeneous temporal Transformer model to perform multi-scale anomaly analysis on the edge weights of a connected hypergraph. This process enables in-depth risk analysis from multiple dimensions, considering risk changes across different time scales and the strength and trends of relationships between entities. Through this analysis, a dynamic risk score is calculated, which intuitively reflects the severity of the risk. Simultaneously, risk levels, responsibility node identifiers, and a fund flow heatmap are output. The risk level classification allows managers to quickly understand the severity of the risk and take appropriate measures accordingly. Responsibility node identifiers clearly identify the source of the risk and the relevant responsible parties, enabling managers to quickly locate the problem and investigate and address the responsible parties. The fund flow heatmap visually displays the flow path and key areas of funds, helping managers clearly understand where funds are going, track their flow throughout the business process, and better grasp the propagation path and scope of risk, providing strong support for developing effective risk response strategies.

[0051] When the dynamic risk score exceeds a predefined risk threshold, a federal collaborative handling mechanism securely shares risk gradient information among the group headquarters, subsidiaries, and tax and banking institutions based on differential privacy and zero-knowledge proofs. Differential privacy technology protects sensitive information while sharing data, ensuring data security; zero-knowledge proofs allow recipients to verify the authenticity and validity of information without obtaining the specific data content. This secure information sharing mechanism enables all relevant parties to understand the risk situation in a timely manner, breaking down information barriers. Based on this, coordinated actions such as invoice freezing, red-letter notifications, or credit adjustments are triggered. When invoices are found to be risky, they are frozen promptly to prevent further circulation and use, avoiding the expansion of risk; red-letter notifications are issued to correct erroneous or problematic invoices, ensuring the accuracy of financial data; for enterprises involved in financing risks, their credit lines are adjusted to reduce the risk exposure of financial institutions. These coordinated actions can quickly and effectively address risks, reduce the scope and losses of risks to enterprises and financial institutions, and protect the legitimate rights and interests of all parties. Attached Figure Description

[0052] Figure 1 This is a schematic diagram illustrating the working principle of the invoice risk control management method based on multi-source data linkage described in this invention.

[0053] Figure 2 A flowchart for multi-source data acquisition and encrypted integration;

[0054] Figure 3 A flowchart for constructing a hypergraph relating invoices, companies, personnel, and assets (quadruple associations);

[0055] Figure 4 This is a graph for temporal edge weight multi-scale anomaly detection and risk assessment.

[0056] Figure 5 This is a performance monitoring chart for the federal collaborative response mechanism. Detailed Implementation

[0057] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0058] Please see Figure 1 This invention provides an invoice risk control management method and system based on multi-source data linkage. The method includes: collecting multi-source data from the entire invoice business process through a multi-source data fusion gateway. The multi-source data covers tax bureau ledgers, bank statements, logistics GPS signals, blockchain-based contract notarization, industrial and commercial judicial records, and ERP image data. The multi-source data fusion gateway performs streaming ETL processing and encrypted integration on the multi-source data. Streaming ETL processing includes data parsing, format conversion, field mapping, and validity verification. Encryption integration uses a homomorphic encryption algorithm to encrypt the processed data, generating an encrypted unified data view and persisting it in a distributed database. Based on the encrypted unified data view, a hypergraph of relationships between invoices, enterprises, personnel, and assets is constructed using dynamic graph association technology. A graph neural network model performs minute-level parsing of the hypergraph, learning entity relationships through message passing and node embedding to identify fraudulent invoice groups and repeated financing risk patterns. Using risk quantification and tracing technology, a heterogeneous time-series Transformer model performs multi-scale anomaly analysis on the edge weight data of the hypergraph, calculates dynamic risk scores, and outputs risk levels, responsibility node identifiers, and a heatmap of fund flows. When the dynamic risk score exceeds a predefined risk threshold, a federal collaborative handling mechanism is activated. Based on differential privacy and zero-knowledge proof technology, risk gradient information is securely shared among the group headquarters, subsidiaries, and tax and banking institutions. Secure sharing triggers coordinated actions such as invoice freezing, red-letter notifications, or credit adjustments. These coordinated actions generate manual review results. Knowledge self-updating technology transforms the review feedback into reinforcement learning reward signals, fine-tuning the risk model parameters online through a policy gradient algorithm, and hot-updating the confidence threshold and parsing rules in the rule base.

[0059] Example 1: See Figure 2The configuration module of the multi-source data fusion gateway predefines multi-source data access rules, which are stored and managed in XML format configuration files. The real-time push interface configuration for tax bureau ledger data includes URL endpoints, authentication tokens, data format protocols, and timeout retry mechanisms. The scheduled retrieval cycle for bank transaction data depends on the batch processing window of the banking system, typically set to perform a full retrieval during the low-load period in the early morning, combined with incremental retrieval during specific hours after peak transaction times. The streaming reception frequency threshold for logistics GPS signals is dynamically adjusted based on the speed of logistics vehicles and business density, and is used in high-speed trunk transportation. High-frequency reporting is used, while lower frequency is used during the delivery phase within the city to balance data accuracy and network load. The smart contract event monitoring for contract blockchain notarization is configured with event signature and contract address filters. The monitoring node immediately captures on-chain events related to invoice business after the blockchain network produces a block. The batch update trigger conditions for industrial and commercial judicial records are related to the update announcements of external government data release platforms. After the platform releases the update notification, it triggers the batch data download and comparison process. The asynchronous upload channel for ERP image data has opened an independent bandwidth resource pool. The channel is configured with a breakpoint resume mechanism and a file integrity verification hash function to prevent the transmission of large-capacity image files from being interrupted or damaged. The multi-source data fusion gateway's core data processing performs streaming ETL operations. The data parsing engine initially unpacks the incoming raw byte stream, identifies the data packet encoding format and compression algorithm, and in the format standardization stage, the application mode mapping table converts heterogeneous data into an internal standard data model. Date and time information is uniformly converted to the ISO8601 standard format, currency amounts are uniformly converted to RMB units and retained to two decimal places, and geographic coordinates are uniformly converted to the WGS84 coordinate system. The field mapping process is based on a central metadata repository, which maintains a field mapping dictionary for each data source. The dictionary clearly records the correspondence between the source system field names and the standard model field names. The data verification stage implements multi-layer rule checks. Integrity checks verify whether required fields have null values, logical consistency checks verify business rules, such as invoice amounts must match contract amounts, bank transaction directions must match transaction types, and business rule compliance judgments refer to the latest tax regulations and financial systems. Invalid data is marked and transferred to a dead-letter queue, which is equipped with monitoring and alarm functions to notify data administrators for manual intervention and repair.

[0060] The encryption integration phase employs homomorphic encryption algorithms to protect the processed clean data. The Paillier encryption scheme is selected for encrypting numerical data, allowing addition operations in ciphertext to meet the needs of subsequent risk analysis for numerical aggregation such as transaction amounts. Text and identifier data are symmetrically encrypted using the AES algorithm, with the encryption process performed in memory to avoid the risk of information leakage caused by plaintext data being written to disk. The generated encrypted unified data view is a logical data layer that provides a unified SQL-like query interface to upper-layer applications. The query interface supports equality queries and range queries in ciphertext. The distributed database uses a columnar storage structure to store the encrypted unified data view, which improves the performance of analytical queries. The data partitioning strategy performs horizontal sharding according to time range and business entities. The multi-replica mechanism ensures data consistency through the Raft consensus protocol, with each data replica stored in a different availability zone. The data access interface implements role-based access control, with access permissions granular down to the row and column levels. Audit logs record all query operations on the encrypted unified data view. The operation of the multi-source data fusion gateway relies on a highly available microservice architecture. The core components of the gateway are deployed on a container orchestration platform, which enables automatic scaling and failover of services. Each data processing module runs as an independent microservice, and the microservices communicate with each other through a lightweight RPC framework. Service mesh technology provides load balancing, service discovery, and circuit breaking mechanisms. The monitoring system collects the performance metrics of the multi-source data fusion gateway, including data throughput, processing latency, and error rate. The log aggregation center collects the running logs of all microservices for troubleshooting and performance optimization. Configuration information is stored in a distributed configuration center, and configuration changes can be pushed to each microservice instance in real time without restarting the service. The security module is integrated into the traffic entry point of the multi-source data fusion gateway and is responsible for network firewall, DDoS attack protection, and API access rate limiting.

[0061] The data quality governance process runs through the entire lifecycle of the multi-source data fusion gateway. The data lineage tracing tool records the complete transformation path from the data source to the encrypted unified data view. The data quality rule engine performs data quality assessments regularly, with assessment dimensions including accuracy, timeliness, and completeness. The data quality dashboard visually displays the compliance status of various quality indicators. Data quality anomalies trigger alarms to notify the data governance team. The data governance team analyzes the root causes and optimizes data access rules or streaming ETL processing logic. The data standards management committee is responsible for maintaining and updating the standard data model and field mapping relationships in the central metadata repository to adapt to business changes and the needs of new data sources. The integration of the multi-source data fusion gateway with the downstream risk analysis system is achieved through an event-driven architecture. Once a batch of new multi-source data has been processed by streaming ETL and encrypted and successfully persisted to an encrypted unified data view, the multi-source data fusion gateway will publish a data ready event to a highly reliable message queue. The event message includes the data batch identifier, data time range, and data source system identifier. The downstream risk analysis system subscribes to this message topic and triggers the subsequent dynamic graph association construction process after receiving the event. This loosely coupled integration method ensures the asynchronicity and scalability of the data processing process and avoids bottlenecks caused by strong dependencies between systems.

[0062] Example 2: See Figure 3The process of extracting entity identifiers from the encrypted unified data view is executed by a dedicated entity parsing engine. This engine connects to a distributed database and uses pre-compiled query statements to batch extract invoice codes, enterprise unified social credit codes, personnel ID numbers, and asset serial numbers. These identifiers, after decryption and formatting, serve as the basic entity nodes of the associative hypergraph. Edge attribute extraction is a parallel data processing flow. Invoice issuance timestamps are extracted from tax bureau ledger data and converted to Unix timestamp format. Transaction amounts are obtained after cross-validation of bank statements and invoice amount fields. Logistics trajectory coordinate sequences are reconstructed from GPS signal streams to reconstruct the complete transportation path. Contract term hash values ​​are directly read from blockchain storage to ensure immutability. Entity nodes and edge attributes are assembled in an in-memory graph computation engine. The initial associative hypergraph is represented as an attribute graph containing multiple node and edge types. The graph structure is maintained in memory using adjacency lists and multi-dimensional indexes to enable fast traversal and querying. Learning node embedding representations using the neighborhood aggregation function of a graph neural network is a computationally intensive process. The graph attention network mechanism is instantiated as a differentiable message-passing function. For each entity node in the graph, the graph attention network calculates its attention coefficients with all its first-order neighbors. These attention coefficients are calculated by a small feedforward neural network that takes the feature vectors of the source and target nodes as input. In the neighborhood aggregation stage, each node weights and sums the feature vectors of its neighbors according to the attention coefficients. The aggregated neighbor features are then concatenated with or averaged with the node's own features, and a new node embedding vector is generated through a nonlinear transformation layer. The node embedding vector is a low-dimensional, dense floating-point vector. The geometric distance in the vector space reflects the semantic similarity between nodes. The similarity matrix is ​​constructed by calculating the cosine similarity of all nodes with respect to their embedding vectors. The cosine similarity value ranges from -1 to 1; the closer the value is to 1, the more similar the node features are. Dynamically updating the hypergraph structure based on the similarity matrix is ​​an iterative optimization step. The system sets a dynamically adjusted similarity threshold. For node pairs belonging to different entity types and whose similarity exceeds the threshold, a virtual edge is automatically added to the hypergraph. The virtual edge represents a potential, indirect association relationship. For example, if an individual acts as the legal representative of multiple newly registered companies within a short period, virtual edges will be established between these companies. The community detection algorithm uses the Louvain method on the updated hypergraph. The Louvain method is a hierarchical clustering algorithm based on modularity optimization. The algorithm continuously moves nodes to neighboring communities to maximize the overall modularity, thereby identifying node clusters with tight internal connections and sparse external connections. These clusters correspond to potential fraudulent invoicing groups or units with repeated financing risks.

[0063] Minute-level parsing of entity relationships using graph neural networks requires efficient model scheduling and inference pipelines. The scheduler monitors the frequency of changes in the association hypergraph. When new data injection causes the graph structure to update to a certain scale, the scheduler triggers the graph neural network model to load pre-trained weights. These pre-trained weights are stored in a model repository, which supports version management and fast loading. The graph neural network model architecture includes multiple graph convolutional layers and attention layers. After loading, the model resides in GPU memory to accelerate inference. Multiple rounds of message passing iterations unfold on the complete association hypergraph. In each iteration, each node receives messages from its neighboring nodes. The message content is the current embedding representation of the neighboring node. Nodes use a learnable aggregation function to integrate all received messages and update their own state. Typically, after 2 to 3 iterations, the node embeddings tend to stabilize. The importance score calculation using the attention mechanism is specifically manifested in calculating the attention weight attached to each edge in the associated hypergraph. This attention weight is calculated by an independent attention network, which takes the embedding vectors of the two connected nodes and the attribute features of the edge as input, and outputs a scalar score. Edges with high scores are considered critical paths; for example, the path traversed by an invoice frequently flowing between multiple companies will receive a high attention score. The selected high-weight edges and their associated nodes are extracted to form a subgraph, which is considered the critical path of risk associations. The critical path reveals the core links of abnormal patterns such as fund circulation, centralized invoice issuance, and goods idleness. Real-time clustering of entities on the critical path uses the density-based DBSCAN algorithm. The DBSCAN algorithm can discover clusters of arbitrary shapes and effectively identify noise points. The clustering results divide the entities on the critical path into several groups. The topological fingerprint of the risk group is generated based on the graph structure features of each group. The topological fingerprint features include the average clustering coefficient, average path length, and node degree distribution entropy of the group, which are encoded into a fixed-length feature vector. The real-time matching engine calculates the similarity between the topological fingerprint feature vector and the feature vector in the historical risk pattern library. The historical risk pattern library stores the topological fingerprints of previously confirmed fraudulent invoicing gangs and tax evasion gangs. The matching process uses an approximate nearest neighbor search algorithm to balance accuracy and speed. The successfully matched groups are marked as high-risk gangs and an early warning event is generated.

[0064] The entire graph computation process relies on a distributed graph computation framework. This framework partitions the associative hypergraph and distributes the partitions across multiple computing nodes. Each node is responsible for storing and computing a subgraph, and message passing between nodes is conducted via a high-speed network. A monitoring system tracks the inference latency, memory usage, and convergence of the community detection algorithm in real time. The resource manager dynamically allocates computing resources based on load, such as increasing the parallelism of graph computation tasks during peak business periods. Every graph structure update and risk identification result is recorded in an audit log. The audit log includes the graph version identifier, model inference parameters, a list of identified risk groups, and their topological fingerprints, used to meet compliance requirements and for subsequent model performance evaluation and optimization. Persistent storage of the associative hypergraph uses a snapshot plus incremental log approach. Full graph data snapshots are periodically saved to the distributed file system, and incremental changes during this period are recorded in log form.

[0065] Example 3: The heterogeneous temporal Transformer model performs multi-scale anomaly analysis on the edge weight data of the associated hypergraph. The edge weight data comes from the connection weights between entities in the associated hypergraph. These weights change continuously over time, forming a data stream with temporal characteristics. The edge weight data is converted into a time series with equal intervals. The conversion process involves dividing the time window. The window size is dynamically adjusted according to the business scenario. The short-term window is set to several hours to capture immediate fluctuations, the medium-term window is set to several days to several weeks to observe trends, and the long-term window is set to several months to identify periodic patterns. Each data point in the time series contains the edge weight value, timestamp, and associated entity identifier. Missing values ​​or outliers are filled and smoothed using linear interpolation or moving average methods to ensure the continuity and integrity of the sequence. The encoder layer of the heterogeneous temporal Transformer model receives a preprocessed temporal sequence as input. The encoder consists of multiple identical layers stacked together. Each layer contains a multi-head self-attention mechanism and a feedforward neural network. The self-attention mechanism calculates the correlation strength between each position in the sequence and all other positions, capturing long-term dependencies, such as the transaction patterns of an invoice with multiple companies over a year. Positional encoding information is added to the input sequence to inject temporal order information. Sine and cosine functions are used to generate positional codes. The encoder output is a high-dimensional feature representation of each time point, which contains deep semantic information about the changes in edge weights.

[0066] Anomaly features at different time scales are calculated using a multi-head self-attention mechanism. This mechanism projects the input features into multiple subspaces, each focusing on a different feature dimension. Short-term anomalies are extracted using a small attention window covering the most recent time steps, focusing on transient changes or spikes in edge weights. Mid-term anomalies are extracted using a moving average attention weight, smoothing short-term noise and revealing trends over several weeks. Long-term anomalies utilize a global attention mechanism to analyze the overall shape and periodicity of the sequence. The feature vectors output by each attention head are concatenated and linearly transformed to form a multi-scale feature tensor, the dimension of which is related to the number of time steps and the feature dimension. The multi-scale anomaly features are fused using a weighted summation method, with the following formula:

[0067] ;

[0068] in: Represents a dynamic risk score. Representing the Fusion weights for each time scale Representing the Abnormal feature values ​​at each time scale, Value , , These correspond to short-term, medium-term, and long-term scales, respectively.

[0069] The gradient backpropagation process is used not only for parameter optimization during model training but also for generating responsibility node identifiers during inference. It calculates the gradient of the dynamic risk score with respect to the input edge weights. The gradient value reflects the contribution of each edge weight to the risk score. The entity nodes corresponding to edge weights with larger gradient magnitudes are marked as responsibility nodes. The responsibility node identifiers indicate the key links in the risk transmission path. The gradient calculation is based on the chain rule, propagating back from the output layer to the input layer. The gradient values ​​of intermediate layers are obtained through automatic differentiation. The responsibility node identifiers are output in list form, containing node identifiers and corresponding gradient contribution scores. The fund flow heatmap is rendered based on the edge weight gradient values. The color depth in the heatmap is proportional to the gradient magnitude, intuitively displaying areas of concentrated risk. The entire analysis process is deployed on a distributed computing framework that supports model parallelism and data parallelism. It automatically scales computing resources when processing massive amounts of edge-weighted data. Model inference results are written to a risk database in real time. Database index optimization supports fast query and aggregation operations. The monitoring module tracks model performance metrics, including inference latency, accuracy, and resource utilization. When metrics are abnormal, alarms are triggered to notify the operations team. The system regularly backs up model parameters and configuration information to ensure rapid reconstruction of the analysis environment after fault recovery. The training of the heterogeneous time-series Transformer model uses historical edge-weighted data as samples. Sample labels are based on confirmed risk events. The training objective is to minimize the cross-entropy loss between predicted scores and true labels. The optimizer employs... The algorithm adaptively adjusts the learning rate, employs early stopping during training to prevent overfitting, and determines the final model version based on performance on the validation set. Model updates are synchronized with business data update frequencies. New models are rolled over to the production environment after validation in the testing environment. A version control system manages model iteration history and supports rapid rollback to a stable version. The results of multi-scale anomaly analysis are integrated with downstream risk management systems. When dynamic risk scores exceed thresholds, early warning processes are automatically triggered. Responsibility node identification and heatmap visualization assist human decision-making. Analysis logs record the complete data processing path, including input data hashes, model version, output scores, and timestamps, meeting audit and compliance requirements. Log analysis tools regularly generate analysis reports summarizing changes in risk patterns and trends in model performance.

[0070] See Figure 4The chart illustrates the changes in connection weights between entities over a complete time period, and the dynamic risk score calculated based on these changes. The blue curve represents the time-series changes in edge weight values, reflecting the dynamic fluctuations in connection strength between entities. Three colored areas identify different types of anomaly patterns: the red area shows short-term anomaly patterns, characterized by instantaneous abrupt changes and spikes in edge weight values, which typically correspond to immediate risk events in the business; the yellow area shows medium-term trend anomalies, characterized by persistent changes in edge weight values ​​over a longer period, reflecting abnormal deviations in business trends; and the green area presents long-term periodic anomalies, showing the disruption of the periodic change pattern of edge weight values ​​on a longer time scale. The red dashed line represents the dynamic risk score curve, calculated by fusing anomaly features from different time scales using a multi-head self-attention mechanism. The risk score directly reflects the severity of business risks at the current point in time, providing a quantitative basis for risk warning and response decisions. The chart clearly demonstrates the correlation between edge weight changes and the risk score; when edge weights experience abnormal fluctuations, the risk score rises accordingly, validating the effectiveness of multi-scale anomaly analysis.

[0071] Example 4: The federated collaborative handling mechanism achieves secure risk information sharing and coordinated control in a distributed environment. Under the federated learning framework, the mechanism initializes the local risk identification models of each participant. Participants include the group headquarters risk control system, subsidiary business systems, tax bureau systems, and the bank's core system. Each participant deploys an identical model skeleton locally, containing feature extraction and classification layers, but with initially randomly generated weights. Each participant uses its own encrypted data to calculate risk gradient information, which is the partial derivative of the model parameters with respect to the loss function, reflecting the model's optimization direction under the current data. The process of adding noise to the gradient information based on differential privacy follows strict privacy budget constraints. Noise addition uses a Gaussian mechanism, and the standard deviation of the noise is related to the gradient sensitivity and privacy budget parameters. The privacy budget is controlled by the epsilon-differential privacy parameter, with the epsilon value set at a low level to limit the risk of information leakage from individual data points. Noise addition is completed before gradient aggregation. The gradient vector generated locally by each participant is added to the Gaussian noise vector, preserving the statistical properties of the perturbed gradient information while protecting individual privacy. The zero-knowledge proof protocol verifies the authenticity and compliance of gradient data. Participants need to prove to the coordination server that the gradient information they submit is calculated from compliant local data without tampering with the original data or the calculation process. The proof process uses non-interactive zero-knowledge proof. Participants generate a proof string about the correctness of the gradient calculation. The coordination server verifies the validity of the string without knowing the specific local data. The verified encrypted gradient information is uploaded to the federated learning coordination server through a TLS-based secure channel.

[0072] After aggregating gradients, the coordination server issues a global model update command. Gradient aggregation uses the FedAvg algorithm to perform a weighted average of encrypted gradients from multiple participants. The weights are allocated according to the amount of data from each participant; participants with larger data volumes contribute more to the global model update. The aggregated global gradients are used to update a central global model. The coordination server encodes the global model update command into a message with a specific format, containing the updated parameter tensor, version number, and timestamp. The update command is broadcast to all participants via a message queue. The global model update command triggers the local system's action executor, which is a predefined rule engine. The rule engine parses the received update command and executes specific risk handling actions based on the command code and local business context. For example, modifying invoice status calls the invoice management system's voiding interface, adjusting credit parameters connects to the credit system's limit management module, and red-letter notifications are sent to the tax bureau system via a direct tax connection. The secure sharing of risk gradient information involves cryptographic operations and secure communication. The generation of temporary session keys employs the Elliptic Curve Diffie-Hellman key exchange protocol. Each participant and the coordinating server generates a pair of temporary public and private keys. By exchanging public key materials, they independently calculate an identical shared secret as the session key, which is valid only within a single communication session. The risk gradient information is symmetrically encrypted using the AES-256 algorithm, with GCM mode selected to provide both confidentiality and integrity protection. The encrypted ciphertext, along with the authentication tag, is transmitted to the target participant through a secure channel. Each participant uses its private key to decrypt the session key, and then uses the session key to decrypt the gradient information. This decryption operation is performed within a hardware security module to protect the key materials. The decrypted gradient information is then used for model inference in a memory-isolated sandbox environment. This sandbox environment restricts network access and file write permissions to prevent the leakage of sensitive gradient data. The model inference outputs disposal suggestions, which are returned in structured JSON format. The decision log for the proposed actions is recorded in detail. The decision log includes timestamps, participant identifiers, action types, and risk scoring criteria. Audit trail records are generated based on the decision logs and are used to meet regulatory requirements and support post-event traceability analysis.

[0073] The operation of the federal collaborative processing mechanism relies on a series of configuration parameters that determine the mechanism's privacy protection strength, communication efficiency, and model performance. See Table 1 for a list of key configuration parameters and their typical values.

[0074]

[0075] Table 1: Key Configuration Parameters of the Federal Coordination Mechanism

[0076] The coordination server's architecture supports horizontal scaling and adopts a microservice architecture. Core components include a gradient aggregator, a model updater, a communication coordinator, and a security manager. The gradient aggregator receives and aggregates encrypted gradients; the model updater maintains the global model version and calculates parameter updates; the communication coordinator manages participant registration, heartbeat detection, and message routing; and the security manager implements cryptographic operations and access control logic. Communication between participants and the coordination server is asynchronous. Participants upload gradients immediately after local training without waiting for other nodes. The coordination server triggers aggregation operations after collecting a sufficient number of gradients. Asynchronous communication improves system throughput and avoids the overhead of synchronous waiting. Fault tolerance mechanisms handle node failures and network partitioning. The coordination server monitors the online status of participants; gradients from failed nodes are excluded from aggregation. Model update commands support retransmission mechanisms to ensure eventual consistency, and checkpoint mechanisms save intermediate states for easy fault recovery. The performance monitoring of the federated collaborative handling mechanism is achieved through a distributed tracing system. This system injects a tracing identifier into each processing step, recording end-to-end latency for gradient calculation, encryption, transmission, aggregation, and model updates. Monitoring metrics include privacy budget consumption, model convergence curves, and communication overhead statistics. These metrics are visualized on a monitoring dashboard, and anomaly detection rules trigger alarms to notify system administrators. The mechanism's security is formally verified; the strength of differential privacy protection is guaranteed by mathematical proof; the security of the zero-knowledge proof protocol is based on the assumption of computational complexity; and the security of the communication channel relies on the TLS protocol standard. Regular security audits ensure consistency between the code and design specifications. Integration of the federated collaborative handling mechanism with existing risk control systems is achieved through an adapter layer. This layer converts data formats and API interfaces, ensuring seamless integration of the new mechanism into the enterprise's existing IT architecture and gradually replacing the traditional centralized risk processing model.

[0077] See Figure 5 This section showcases the performance of risk information sharing and coordinated control. The blue curve represents the trend of model recognition accuracy with training epochs, reflecting the gradual improvement of global model performance during federated learning. The convergence curve of model accuracy demonstrates the effect of model optimization through gradient aggregation while protecting data privacy. The red dashed line represents the consumption of the privacy budget, reflecting the strength of the differential privacy protection mechanism in protecting individual privacy information. As the training epochs increase, the privacy budget is gradually consumed, ensuring that the risk of information leakage of individual data points is effectively controlled. The green dotted line shows the changes in communication resource overhead, indicating the data transmission cost between participants and the coordination server during federated learning. Fluctuations in communication overhead reflect changes in the system's resource requirements at different training stages. The background bar chart shows the number of active participants in each training epoch, reflecting the collaborative work of participating nodes in the distributed system.

[0078] Example 5: Manual review results are obtained by risk control experts who conduct secondary confirmation of early warning cases triggered by the federal collaborative handling mechanism. Experts review the related hypergraphs, risk scores, fund flow heatmaps, and original transaction voucher images provided by the system through the review workbench. For correctly identified risk cases, such as a genuine fraudulent invoicing group being accurately marked by the system, the expert clicks the confirmation button to generate a positive reward signal. For false alarm cases, such as a normal trading company being misjudged as risky due to its special business model, the expert clicks the false alarm button to generate a negative reward signal. Missed cases are discovered during post-audit and manually entered into the system, generating a stronger negative reward signal. The defined reward function includes multi-dimensional reward indicators. Risk identification accuracy is the core indicator, calculated by comparing system early warnings and expert confirmation results. Handling response time measures the time interval from risk identification to the completion of the handling action. Resource consumption considers the computing and storage resources used by the system during data analysis and model inference. The multi-dimensional reward indicators are combined into a comprehensive reward value through a weighted summation method, with weight coefficients set by domain experts based on business priorities.

[0079] Convolutional neural networks (CNNs) are used to perform sentiment analysis on the review feedback text to extract implicit reward weights. Risk control experts often leave text annotations in the system when making review decisions. These annotations include the reasons for the decision and subjective evaluations, such as "Although the company has concentrated transactions, it has a real logistics background" or "The obvious traces of fund repatriation support the high-risk judgment." The CNN model takes the word vector sequence of these annotation texts as input. The word vectors are obtained through a pre-trained language model. The convolutional kernel performs a sliding convolution operation on the word vector sequence to extract local semantic features. Max pooling layers select the most significant features, and fully connected layers map the features to sentiment polarity scores. The sentiment polarity score acts as an implicit reward weight multiplier, adjusting the base reward value. Positive comments such as "accurate judgment" generate a multiplier greater than 1 to amplify the reward, while negative comments such as "insufficient evidence" generate a multiplier less than 1 to decay the reward. To balance the sparse reward problem, a reward pruning technique is used. This technique sets an upper and lower reward threshold; rewards exceeding the upper threshold are pruned to the upper limit, preventing excessive impact of extreme cases on the model. Reward pruning also ensures stable policy gradient update steps, preventing drastic fluctuations during training. A reinforcement learning-based policy gradient algorithm is used to adjust the parameters of the risk model. This algorithm directly optimizes the policy function of the risk model, which takes the system state as input. State features include topological indicators of the hypergraph, statistical characteristics of edge weight time series, and entity historical behavior profiles. The policy function outputs the probability distribution of different risk judgment actions for the target entity given the state. The action space includes marking as high-risk, marking as low-risk, and suggesting manual review. The optimization objective is to maximize the long-term cumulative reward. The long-term reward considers the sum of reward discounts over multiple future time steps, with a discount factor set to less than 1 to reduce the present value contribution of future rewards. The policy gradient algorithm calculates the gradient of the objective function with respect to the model parameters and updates the parameters along the gradient direction to improve the expected reward. Gradient estimation uses the Monte Carlo method, calculating the average gradient based on a complete set of decision trajectories.

[0080] The rule base is dynamically updated through an online learning framework, which supports hot updates of model parameters and business rules without interrupting risk control services. Updates include adjusting the confidence threshold for risk assessment. This threshold, the minimum score boundary for risk determination, is dynamically adjusted based on the model's accuracy and recall on recent data. The threshold is raised to reduce false positives when accuracy declines and lowered to capture more risks when recall is insufficient. The logical conditions of entity parsing rules are optimized based on changes in entity relationship patterns. For example, a new holding relationship pattern is added to identify related companies under the same ultimate controlling shareholder. Risk pattern feature templates are expanded based on newly emerging risk methods, such as adding a feature vector to identify fraudulent invoicing through cross-border e-commerce platforms. Updated rules and model parameters are deployed to the production environment in real-time after integrity verification. Integrity verification includes syntax checking, logical conflict detection, and performance benchmarking. A blue-green deployment strategy is used, maintaining two identical production environments: one running the old version and the other deploying the new version. Seamless upgrades are achieved through traffic switching, and a version rollback mechanism automatically switches back to the old version when performance metrics are abnormal.

[0081] A specific example illustrates the operation of the knowledge self-updating technology. The system generates a high-risk warning for a company called "ABC Trading Co., Ltd.", with a dynamic risk score of 0.92, triggering an invoice freezing action. Risk control expert Zhang Gong retrieves comprehensive data on ABC Trading Co., Ltd. on the review workbench. The hypergraph shows that the company has frequent financial transactions with three newly registered companies, and bank statements show a clear rapid inflow and outflow characteristic. However, the contract blockchain evidence is complete, and logistics GPS signals show actual goods movement. Based on experience, Zhang Gong judges that this is a typical financing trade based on real transactions rather than fictitious invoicing. Therefore, he clicks the "false alarm" button in the system and enters the text in the comment box: "Abnormal cash flow but real goods flow, complete contract, recommended for key monitoring but not frozen for now." The system records the review result and generates a negative reward signal. The convolutional neural network performs sentiment analysis on the comment text "Abnormal cash flow but real goods flow, complete contract, recommended for key monitoring but not frozen for now." The analysis results identify positive words such as "real" and "complete" and cautious words such as "abnormal" and "monitoring," ultimately outputting a neutral to negative sentiment polarity strength score of 0.8. The reward function integrates accuracy, response time, and sentiment weights to calculate a negative overall reward. The policy gradient algorithm uses this reward to update the parameters of the risk model, adjusting the weighting of patterns like "abnormal cash flow but genuine goods flow," reducing the contribution of risk scoring solely based on cash flow characteristics. Simultaneously, based on this feedback, the rule base update module adds a new rule to the entity parsing rules: for enterprises with genuine logistics and complete contracts but abnormal cash flow, the risk level is adjusted to "monitoring" instead of "high risk," and the confidence threshold is adjusted from 0.9 to 0.95. This update will take effect in the next model release cycle, and the system will handle similar cases more accurately in the future.

[0082] The infrastructure for knowledge self-updating technology includes a versioned feedback knowledge base that stores all historical review records, corresponding reward signals, model parameter versions, and rule snapshots. Data traceability supports analysis of the root causes of model performance changes. An A / B testing framework conducts small-scale experiments before rule updates, redirecting some traffic to the new rule version and comparing business metrics between the old and new versions. Experimental data assists in deciding the timing of large-scale deployment. A monitoring system tracks key indicators of knowledge self-updating technology, including model stability, rule effectiveness rate, and feedback loop latency. Anomaly monitoring ensures the self-learning process remains under control. Knowledge self-updating technology enables the risk control system to continuously evolve from real-world experience, constantly absorbing expert experience to optimize judgment logic and adapt to increasingly complex financial violation methods.

[0083] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0084] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A multi-source data linkage-based invoice risk control management method, characterized in that, The method comprises: Real-time acquisition of multi-source data of the whole process of invoice business through a multi-source data fusion gateway, including tax bureau bottom account, bank flow, logistics GPS signal, contract blockchain storage, business judicial record and ERP image data, and performing stream ETL processing and encryption integration on the multi-source data to generate an encrypted unified data view; Based on the encrypted unified data view, a dynamic graph correlation technology is used to construct a correlation hypergraph between the four tuples of invoices, enterprises, personnel and assets, and a graph neural network is used to analyze the entity relationship at a minute level to identify fake invoice gangs and repeated financing risks; Risk quantification and traceability technology is used to perform multi-scale anomaly analysis on the edge weight of the correlation hypergraph using a heterogeneous time series Transformer model, calculate a dynamic risk score, and output a risk level, a responsibility node identifier and a fund flow heat map; When the dynamic risk score exceeds a predefined risk threshold, a federal collaborative disposal mechanism is used to securely share risk gradient information between the group headquarters, subsidiary companies and tax and banking institutions based on differential privacy and zero-knowledge proof, triggering the linkage disposal actions of invoice freezing, red notice or credit adjustment; According to the artificial review results of the linkage disposal actions, knowledge self-updating technology is used to convert the review feedback into reinforcement learning rewards, online fine-tune the risk model and hot update the rule library; The construction of the correlation hypergraph between the four tuples of invoices, enterprises, personnel and assets comprises: Extracting invoice codes, enterprise unified social credit codes, personnel ID numbers and asset serial numbers as entity nodes from the encrypted unified data view, and extracting invoice issuing time, transaction amount, logistics track and contract terms as edge attributes; Using the neighborhood aggregation function of the graph neural network to perform embedded representation learning on the entity nodes, and calculating the similarity matrix between the nodes; Based on the similarity matrix, dynamically updating the hypergraph structure, adding virtual edges to capture potential associations across entity types, and identifying subgraph patterns of risk gangs through community detection algorithms; The minute-level analysis of entity relationships by the graph neural network comprises: Scheduling the graph neural network model to load pre-trained weights, performing multiple rounds of message passing iterations on the correlation hypergraph, and updating node embeddings; Using an attention mechanism to calculate the importance score of the edge weight, and selecting high-weight edges as the key path of risk correlation; Real-time clustering of entities on the key path to generate the topological fingerprint of the risk gang, and matching with the historical risk pattern library; The multi-scale anomaly analysis of the edge weight of the correlation hypergraph comprises: Converting the edge weight data into a time series sequence and inputting it into the encoder layer of the heterogeneous time series Transformer model to capture long-term dependencies; Calculating abnormal features at different time scales through a multi-head self-attention mechanism, including short-term fluctuations, medium-term trends and long-term cycles; Fusing multi-scale anomaly features, using a fully connected layer to output a dynamic risk score, and generating a responsibility node identifier through gradient backpropagation.

2. The invoice risk control management method based on multi-source data linkage of claim 1, wherein, The real-time acquisition of multi-source data of the whole process of invoice business through a multi-source data fusion gateway comprises: Configure multi-source data access rules, specifying the real-time push interface for tax bureau ledger data, the timed retrieval cycle for bank transaction data, the streaming reception frequency for logistics GPS signals, the smart contract event monitoring for contract blockchain notarization, the batch update trigger conditions for industrial and commercial judicial records, and the asynchronous upload channel for ERP image data. Perform format standardization, field mapping, and data validation on the multi-source data to eliminate semantic conflicts between heterogeneous data sources; Homomorphic encryption is performed on the standardized multi-source data using an encryption algorithm to generate the encrypted unified data view, which is then stored in a distributed database. 3.The invoice risk control management method based on multi-source data linkage of claim 1, wherein, The triggering of coordinated response actions through the federal collaborative response mechanism includes: Under the federated learning framework, each participant's local model is initialized, and the encrypted risk gradient information is uploaded to the coordination server. Noise is added to gradient information based on differential privacy to meet privacy budget constraints, and the authenticity of gradient data is verified by zero-knowledge proof. After the coordinating server aggregates gradients, it issues a global model update command, triggering the local system's action executor to complete the modification of invoice status or adjustment of credit parameters.

4. The invoice risk control management method based on multi-source data linkage of claim 3, wherein, The secure sharing risk gradient information includes: A temporary session key is generated, the risk gradient information is symmetrically encrypted, and then transmitted to the participants through a secure channel. After the participants decrypt the gradient information using their private keys, they perform model inference in the sandbox environment and output disposal suggestions. Record decision logs of handling recommendations and generate audit trail records for subsequent tracing.

5. The invoice risk control management method based on multi-source data linkage of claim 1, wherein, The process of transforming review feedback into reinforcement learning rewards through knowledge self-updating technology, and fine-tuning the risk model online and hot-updating the rule base includes: The results of manual review are converted into reward signals, with positive rewards corresponding to correct risk identification and negative rewards corresponding to false alarms or omissions. The policy gradient algorithm of reinforcement learning is used to adjust the parameters of the risk model to maximize the long-term reward accumulation. The confidence threshold and entity parsing rules of the rule base are dynamically updated through an online learning framework and deployed to the production environment in real time.

6. The invoice risk control management method based on multi-source data linkage of claim 5, wherein, The process of converting review feedback into reinforcement learning rewards includes: Define a reward function, where risk identification accuracy, response time, and resource consumption are multi-dimensional reward indicators; Using a convolutional neural network to perform sentiment analysis on the review feedback text and extract implicit reward weights; By employing reward pruning techniques, the sparse reward problem can be balanced, ensuring the model's convergence stability.

7. A multi-source data linkage-based invoice risk control management system, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the invoice risk control management method based on multi-source data linkage as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • All-electricity invoice management monitoring platform

    CN118071432A

  • Electronic bill data management system and method based on artificial intelligence

    CN119205383A