Invoice risk control management method and system based on multi-source data linkage

By integrating multi-source data and graph neural network analysis, combined with a federal collaborative processing mechanism, the problems of data isolation and low risk identification efficiency in traditional invoice management have been solved. This has enabled real-time identification and coordinated processing of false invoices and duplicate financing, ensuring information security and data privacy.

CN121120289AActive Publication Date: 2025-12-12GUANGZHOU LESHUI INFORMATION TECH CO LTD

Patent Information

Application Number
CN202511649154.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-12
Publication Date
2025-12-12
Estimated Expiration
2045-11-12

AI Technical Summary

Technical Problem

Traditional invoice management relies on a single data source, making it difficult to identify risks such as fraudulent invoices and duplicate financing. Data processing efficiency is low, and information is isolated between departments, making it impossible to share information in a timely manner and coordinate risk management.

Method used

The system collects real-time data from tax bureau ledgers, bank statements, logistics GPS signals, contract blockchain records, and ERP image data through a multi-source data fusion gateway. It generates an encrypted unified data view, performs minute-level analysis and multi-scale anomaly analysis using graph neural networks and heterogeneous time-series Transformer models, and securely shares risk information and triggers coordinated response actions using a federated collaborative handling mechanism.

Benefits of technology

It enables real-time monitoring of the entire invoice business process, accurately identifies fraudulent invoice issuance groups and risks of repeated financing, and takes timely measures to reduce tax losses and financial risks, while ensuring information security and data privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121120289A_ABST
    Figure CN121120289A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of invoice risk control management, and discloses an invoice risk control management method and system based on multi-source data linkage. According to the method, multi-source data of the whole invoice business process is collected in real time through a multi-source data fusion gateway. And performing streaming ETL processing and encryption integration on the multi-source data to generate an encrypted unified data view. And based on the view, constructing an association hypergraph by utilizing a dynamic graph association technology, analyzing an entity relationship by virtue of a graph neural network, and identifying false invoice making gang and repeated financing risks. And calculating a dynamic risk score by adopting a risk quantification and traceability technology, and outputting a risk level, a responsibility node identifier and a capital flow thermodynamic diagram. When the risk score exceeds a threshold value, risk gradient information is safely shared among the group headquarters, the molecular companies and the tax-bank institutions, and linkage processing actions such as invoice freezing, red character notification or credit adjustment are triggered. According to a manual reexamination result, the risk model is finely adjusted through a knowledge self-updating technology, and the rule base is hot-updated.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of invoice risk control management, in particular to an invoice risk control management method and system based on multi-source data linkage. BACKGROUND

[0002] As a key financial and tax document in economic activities, the invoice plays a vital role. From the perspective of internal financial management, it is the legal document for financial revenue and expenditure. Each income and expenditure is recorded based on the invoice, so that the financial status of the enterprise can be clearly presented. When purchasing raw materials, the invoice issued by the supplier records the variety, quantity, price and other information of the purchase, and the enterprise financial personnel bases on it to process the account and accurately calculate the cost. At the same time, the invoice is also the basis for accounting, providing indispensable data support for preparing financial statements, helping the management to understand the business results and make scientific decisions.

[0003] In the field of tax collection and management, the invoice is also an important basis for tax authorities to control the tax source and collect tax. Through the monitoring of the issuance and circulation of invoices, the tax authorities ensure that enterprises pay taxes according to law and prevent tax loss. For value-added tax general taxpayers, the input tax and output tax marked on the value-added tax invoice are directly related to the amount of value-added tax that the enterprise should pay, and the tax authorities use this as a clue for tax inspection to ensure tax fairness. The invoice is also an important means for the state to supervise economic activities, maintain economic order and protect national property safety, and plays a key role in promoting the healthy and stable development of the market economy.

[0004] Traditional invoice management often relies on a single data source, which greatly limits the effective identification of invoice risks. For example, only according to the tax bureau's bottom account, although the basic issuance information of the invoice can be obtained, it is difficult to detect the false opening behavior hidden by the complex fund flow and contract relationship of the enterprise. The enterprise may sign false contracts with multiple related parties, and the funds circulate among these related parties, creating false transaction illusions, while issuing corresponding invoices. Due to the lack of comprehensive analysis of multi-source data such as bank flow and logistics GPS signals, it is difficult to find these abnormalities from the tax bureau's bottom account data, leading to long-term concealment of false invoice behavior, causing serious harm to national tax revenue and market economic order. Similarly, when identifying the risk of repeated financing, relying on a single data source cannot fully understand the asset and liability situation and the direction of the enterprise's fund flow, making it difficult to accurately determine whether the enterprise has the problem of using the same asset to repeat financing.

[0005] The traditional invoice management mode has obvious deficiencies in data processing efficiency and cannot realize real-time monitoring of the whole process of invoice business. In the data acquisition link, it may rely on manual input or periodic batch import, which not only consumes time and effort, but also is prone to data errors. After the invoice is issued, it may take several days or even longer to enter the relevant data into the tax system or the enterprise internal financial system. In the data processing stage, the traditional processing technology and process are difficult to cope with the rapid analysis demand of massive invoice data, leading to the risk being detected only after the risk occurs. When the enterprise commits the act of fictitious invoice, due to the lag of data processing, the tax authorities may discover the problem when they conduct routine inspection several months later, at which time the enterprise may have caused a large amount of tax loss, and the relevant person in charge may have transferred assets, increasing the difficulty of subsequent pursuit and punishment.

[0006] The information of each department and institution is isolated in the invoice management, and cannot form a combined force for disposing risks. The enterprise, tax bureau and bank, etc. are each responsible for their own affairs, and cannot share risk information in time. When the enterprise issues an invoice, the tax authorities cannot obtain the bank fund flow information and logistics of the enterprise in real time, and it is difficult to judge the authenticity and reasonableness of the invoice issuance. When the enterprise is suspected of fictitious invoice, after the tax authorities discover the problem, due to the lack of effective information sharing mechanism between the bank, it is difficult to obtain the fund flow data of the enterprise, and it is difficult to track and freeze the fund, affecting the efficiency of case handling. At the same time, there is a problem of poor information communication between different departments in the enterprise. The financial department only focuses on the account processing of the invoice, while the business department is responsible for the authenticity of the business behind the invoice, but there is a lack of effective cooperation and information sharing between the two, so that the invoice risk is difficult to be discovered and solved in time in the enterprise. SUMMARY

[0007] The purpose of the present application is to provide an invoice risk control management method and system based on multi-source data linkage to solve the problems raised in the above background art.

[0008] To achieve the above-mentioned purpose, the present application provides an invoice risk control management method based on multi-source data linkage, which comprises: Real-time multi-source data of the whole process of invoice business is collected through a multi-source data fusion gateway, including tax bureau bottom account, bank flow, logistics GPS signal, contract blockchain storage, business judicial record and ERP image data, and the multi-source data is processed by stream ETL and encrypted and integrated to generate an encrypted unified data view; Based on the encrypted unified data view, a dynamic graph correlation technology is used to construct the correlation hypergraph between the four tuples of invoice, enterprise, personnel and asset, and the entity relationship is analyzed by minute level through graph neural network to identify fictitious invoice gang and repeated financing risk; Adopting risk quantification and traceability technology, using heterogeneous time series Transformer model to perform multi-scale anomaly analysis on the edge weight of the associated hypergraph, calculate dynamic risk score, and output risk level, responsibility node identification and fund flow heat map; When the dynamic risk score exceeds the predefined risk threshold, through the federal collaborative disposal mechanism, based on differential privacy and zero-knowledge proof, the risk gradient information is safely shared between the group headquarters, the subsidiary company and the tax and financial institutions, triggering the linkage disposal actions of invoice freezing, red notice or credit adjustment; According to the artificial review results of the linkage disposal actions, the review feedback is converted into reinforcement learning rewards through knowledge self-updating technology, the risk model is fine-tuned online, and the rule library is hot-updated.

[0009] Preferably, the real-time collection of multi-source data of invoice business full process through the multi-source data fusion gateway includes: Configure multi-source data access rules, specify real-time push interface of tax bureau bottom account data, timing pull period of bank flow data, streaming reception frequency of logistics GPS signal, smart contract event listening of contract blockchain storage, batch update trigger condition of industrial and commercial judicial records, and asynchronous upload channel of ERP image data; Perform format standardization, field mapping and data verification on the multi-source data, eliminate semantic conflicts between heterogeneous data sources; Perform homomorphic encryption on the standardized multi-source data through encryption algorithm, generate the encrypted unified data view, and store it in distributed database.

[0010] Preferably, the construction of the associated hypergraph between the invoice, enterprise, personnel and asset four-tuples includes: Extract invoice code, enterprise unified social credit code, personnel ID number and asset serial number as entity nodes from the encrypted unified data view, and extract invoice issuing time, transaction amount, logistics track and contract terms as edge attributes; Use the neighborhood aggregation function of graph neural network to perform embedding representation learning on the entity nodes, and calculate the similarity matrix between nodes; Based on the similarity matrix, dynamically update the hypergraph structure, add virtual edges to capture potential associations across entity types, and identify subgraph patterns of risk gangs through community detection algorithm.

[0011] Preferably, the minute-level analysis of entity relationship through graph neural network includes: Schedule the graph neural network model to load pre-trained weights, perform multiple rounds of message passing iterations on the associated hypergraph, and update node embeddings; Use attention mechanism to calculate the importance score of edge weight, and filter high weight edges as key paths of risk association; Real-time clustering of entities on the critical path, generating a topological fingerprint of the risk group, and matching it with the historical risk pattern library.

[0012] Preferably, the multi-scale anomaly analysis of the edge weight of the association hypergraph comprises: Convert the edge weight data into a time series sequence and input it into the encoder layer of the heterogeneous time series Transformer model to capture long-term dependencies; Calculate the anomaly features of different time scales through the multi-head self-attention mechanism, including short-term fluctuations, medium-term trends, and long-term cycles; Fuse multi-scale anomaly features, use a fully connected layer to output a dynamic risk score, and generate a responsible node identifier through gradient backpropagation.

[0013] Preferably, the triggering of joint handling actions through the federal collaborative handling mechanism comprises: Under the federal learning framework, initialize the local model of each participant and upload the encrypted risk gradient information to the coordination server; Add noise to the gradient information based on differential privacy to meet the privacy budget constraint and verify the authenticity of the gradient data through zero-knowledge proof; The coordination server aggregates the gradients and issues global model update instructions to trigger the handling action executor of the local system to complete invoice status modification or credit parameter adjustment.

[0014] Preferably, the secure sharing of risk gradient information comprises: Generate a temporary session key, symmetrically encrypt the risk gradient information, and transmit it to the participants through a secure channel; The participants use the private key to decrypt the gradient information and execute model inference in a sandbox environment to output handling suggestions; Record the decision log of the handling suggestions and generate audit tracking records for subsequent tracing.

[0015] Preferably, the conversion of review feedback into reinforcement learning rewards through knowledge self-updating technology, online fine-tuning of risk models, and hot updating of rule libraries comprises: Convert the artificial review results into reward signals, where positive rewards correspond to correct risk identification and negative rewards correspond to false positives or false negatives; Use the policy gradient algorithm of reinforcement learning to adjust the parameters of the risk model to maximize the long-term reward accumulation value; Dynamically update the confidence threshold and entity parsing rules of the rule library through the online learning framework and deploy them to the production environment in real time.

[0016] Preferably, the conversion of review feedback into reinforcement learning rewards comprises: Define the reward function, where risk identification accuracy, treatment response time and resource consumption are multi-dimensional reward indicators; Use a convolutional neural network to perform sentiment analysis on the review feedback text and extract implicit reward weights. Balance the sparse reward problem through reward clipping technology to ensure the stability of model convergence.

[0017] Preferably, the present application also includes an invoice risk control management system based on multi-source data linkage, comprising a memory, a processor and a computer program stored in the memory and running on the processor, wherein the processor, when executing the computer program, realizes the steps of the above-mentioned invoice risk control management method based on multi-source data linkage.

[0018] Compared with the prior art, the present application has the following advantages: The present application realizes real-time collection of multi-source data of invoice business full process through a multi-source data fusion gateway, covering tax bureau bottom account, bank flow, logistics GPS signal, contract blockchain storage, business judicial record and ERP image data, etc. These data sources are extensive and reflect relevant information of invoice business from different angles. The tax bureau bottom account records the basic information of invoice issuing, the bank flow reflects the flow of funds, the logistics GPS signal shows the transportation track of goods, the contract blockchain storage ensures the authenticity and non-tamperability of the contract, the business judicial record reflects the legality and credit status of the enterprise, and the ERP image data contains relevant image data of internal business processing. Through stream ETL processing and encryption integration of these multi-source data, an encrypted unified data view is generated. This process organically combines the originally scattered and isolated data together, provides a comprehensive and accurate data basis for subsequent risk analysis, and enables relevant personnel to grasp the overall situation of invoice business, understand each link and related factors of the business, and thus more effectively identify potential risk points.

[0019] Based on the unified data view based on encryption, the association hypergraph between the four-tuple of invoices, enterprises, personnel, and assets is constructed using dynamic graph association technology. This hypergraph can clearly show the complex relationships between entities, closely linking invoices with the enterprises that issued them, the personnel involved, and related assets. By using graph neural networks to analyze entity relationships at a minute level, potential information in these relationships can be deeply mined. When analyzing the relationship between invoices and enterprises, it can be determined whether there are abnormal invoice issuance and receipt patterns between enterprises; when studying the association between personnel and invoices, it can be identified whether personnel frequently participate in abnormal invoice business. This precise analysis can accurately identify fake invoice gangs and repeated financing risks. For fake invoice gangs, by analyzing their invoice circulation, financial transactions, and personnel associations, etc., the gang members and their modus operandi can be quickly locked in, and timely measures can be taken to combat them, avoiding loss of state tax revenue. For repeated financing risks, by correlating enterprise assets and financing data, it can be accurately determined whether the enterprise has engaged in repeated financing using the same assets, providing accurate risk warnings for financial institutions, reducing credit risks for financial institutions, and ensuring the stability of the financial market.

[0020] Using risk quantification and traceability technology, a heterogeneous time series Transformer model is used to perform multi-scale anomaly analysis on the edge weights of the association hypergraph. This process can deeply analyze risks from multiple dimensions, taking into account the changes in risks at different time scales and the strength and trends of relationships between entities. Through this analysis, a dynamic risk score is calculated, which can intuitively reflect the severity of the risk. At the same time, the risk level, responsibility node identification, and fund flow heat map are output. The division of risk levels allows managers to quickly understand the degree of risk, so that appropriate measures can be taken according to the severity of the risk. The responsibility node identification clearly identifies the source of the risk and the relevant responsible parties, allowing managers to quickly locate the problem and investigate and handle the responsible parties. The fund flow heat map visually displays the flow path and key flow areas of funds, helping managers clearly understand the destination of funds and track the flow of funds in the entire business process, thereby better grasping the propagation path and impact range of risks, and providing strong support for developing effective risk response strategies.

[0021] When the dynamic risk score exceeds the predefined risk threshold, the risk gradient information is securely shared among the group headquarters, subsidiary companies and tax agencies based on differential privacy and zero-knowledge proof through a federal collaborative disposal mechanism. The differential privacy technology can protect sensitive information while sharing information, ensuring data security; the zero-knowledge proof enables the receiving party to verify the authenticity and validity of the information without obtaining the specific data content. This secure information sharing mechanism enables relevant parties to learn about the risk situation in a timely manner, breaking down information barriers. On this basis, the linkage disposal actions of invoice freezing, red notice or credit adjustment are triggered. When a risk is found in an invoice, the invoice is frozen in time to prevent its further circulation and use, avoiding the expansion of the risk; a red notice is issued to correct the wrong or problematic invoice, ensuring the accuracy of the financial data; for enterprises involved in financing risks, the credit limit is adjusted to reduce the risk exposure of financial institutions. These linkage disposal actions can quickly and effectively respond to risks, reduce the scope of influence and losses of risks on enterprises and financial institutions, and protect the legitimate rights and interests of all parties. BRIEF DESCRIPTION OF DRAWINGS

[0022] Figure 1 A working principle diagram of the invoice risk control management method based on multi-source data linkage described in the present application; Figure 2 A flowchart of multi-source data acquisition and encryption integration; Figure 3 A flowchart of invoice, enterprise, personnel and asset four-tuple association hypergraph construction; Figure 4 A time series edge weight multi-scale anomaly detection and risk assessment diagram; Figure 5 A federal collaborative disposal mechanism performance monitoring diagram. DETAILED DESCRIPTION

[0023] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0024] Please refer to Figure 1The application provides a kind of invoice risk control management method and system based on multi-source data linkage, the method includes: through the multi-source data fusion gateway, the multi-source data of invoice business whole process is collected, multi-source data covers tax bureau bottom account, bank flow, logistics GPS signal, contract block chain storage, business judicial record and ERP image data.Multi-source data fusion gateway carries out stream ETL processing and encryption integration to multi-source data, stream ETL processing includes data analysis, format conversion, field mapping and validity check, encryption integration uses homomorphic encryption algorithm to encrypt the data processed, generates encrypted uniform data view and persists in distributed database.Based on the uniform data view of encryption, the correlation hypergraph between invoice, enterprise, personnel and asset four tuples is constructed using dynamic graph correlation technology.Graph neural network model carries out minute-level analysis to correlation hypergraph, learns entity relationship through message passing and node embedding, identifies virtual invoice gang and repeated financing risk mode.Adopt risk quantization and traceability technology, the edge weight data of correlation hypergraph is carried out multi-scale anomaly analysis using heterogeneous time sequence Transformer model, calculates dynamic risk score, outputs risk level, responsibility node identification and fund flow direction heat map.When dynamic risk score exceeds predefined risk threshold, federal collaborative disposal mechanism starts, based on differential privacy and zero-knowledge proof technology, risk gradient information is safely shared between group headquarters, subsidiary company, tax and bank institutions.Safety sharing triggers invoice freezing, red letter notification or credit adjustment linkage disposal action.Linkage disposal action produces artificial review result, knowledge self-updating technology converts review feedback into reinforcement learning reward signal, adjusts risk model parameters online through policy gradient algorithm, and hot updates confidence threshold and analysis rules in rule base.

[0025] Example 1: see Figure 2The configuration module of the multi-source data fusion gateway pre-defines multi-source data access rules, the access rules are stored and managed in the form of XML configuration files, the real-time push interface configuration of the tax bureau's basic account data includes URL endpoints, authentication tokens, data format protocols, and timeout retry mechanisms, the timing pull cycle setting of bank transaction data depends on the batch processing window of the bank business system, usually set to full pull at low business load period in the early morning every day, and incremental pull at specific hours after the transaction peak, the streaming reception frequency threshold of logistics GPS signals is dynamically adjusted according to the moving speed of logistics vehicles and business intensity, high frequency reporting is used in high-speed trunk transportation stage, and lower frequency is used in urban distribution stage to balance data accuracy and network load, the smart contract event listening configuration of the contract blockchain storage is configured with event signature and contract address filter, the listening node captures the on-chain events related to invoice business immediately after the block is generated in the blockchain network, the batch update trigger condition of the industrial and commercial judicial records is associated with the update announcement of the external government data publishing platform, the platform triggers batch data download and comparison process after publishing the update notification, and the asynchronous upload channel of ERP image data opens an independent bandwidth resource pool, the channel is configured with a breakpoint resume mechanism and a file integrity check hash function to prevent interruption or damage of large-capacity image files during transmission. The data processing core of the multi-source data fusion gateway performs streaming ETL operations, the data parsing engine performs preliminary unpacking on the accessed raw byte stream, identifies the encoding format and compression algorithm of the data packet, and applies a pattern mapping table in the format standardization stage to convert heterogeneous data into an internal standard data model, converts date and time information into ISO8601 standard format, converts currency amounts into RMB units and retains two decimal places, converts geographic location coordinates into WGS84 coordinate system, and performs field mapping according to a central metadata warehouse, the metadata warehouse maintains a field mapping dictionary for each data source, which clearly records the correspondence between source system field names and standard model field names, multi-layer rule checking is implemented in the data verification link, integrity checking verifies whether the mandatory fields exist, logical consistency checking verifies business rules, such as invoice amount must match contract amount, bank transaction direction must match transaction type, business rule compliance judgment refers to the latest tax regulations and financial systems, invalid data is marked and transferred to the dead letter queue, the dead letter queue is equipped with monitoring and alarm functions, and notifies the data administrator for manual intervention and repair.

[0026] In the encryption integration stage, the processed clean data is protected by an encryption algorithm with homomorphic properties. The Paillier encryption scheme is selected for the encryption of numerical data. The Paillier encryption scheme allows addition operations in the ciphertext state, meeting the demand for numerical aggregation such as transaction amount in subsequent risk analysis. Text and identification data are symmetrically encrypted using the AES algorithm. The encryption process is performed in memory to avoid information leakage risks caused by the storage of plaintext data. The generated encrypted unified data view is a logical data layer. The view provides a unified SQL-like query interface for the upper-layer application. The query interface supports equality queries and range queries in the ciphertext state. The distributed database uses a columnar storage structure to store the encrypted unified data view. Columnar storage improves the performance of analytical queries. The data partitioning strategy performs horizontal sharding according to time ranges and business entities. The multi-replica mechanism ensures data consistency through the Raft consensus protocol. Each data replica is stored in a different availability zone. The data access interface implements role-based permission control. The access permissions are fine-grained to the row level and column level. Audit logs record all query operations on the encrypted unified data view. The operation of the multi-source data fusion gateway relies on a highly available microservice architecture. The gateway core components are deployed on a container orchestration platform. The container orchestration platform implements service auto-scaling and failover. Each data processing module runs as an independent microservice. Microservices communicate with each other through a lightweight RPC framework. Service mesh technology provides load balancing, service discovery, and circuit breaking mechanisms. A monitoring system collects performance indicators of the multi-source data fusion gateway, including data throughput, processing delay, and error rate. A log aggregation center collects the running logs of all microservices for troubleshooting and performance optimization. Configuration information is stored in a distributed configuration center. Configuration changes can be pushed to each microservice instance in real time without restarting the service. Security modules are integrated into the traffic entrance of the multi-source data fusion gateway. The security modules are responsible for network firewalls, DDoS attack protection, and API access throttling.

[0027] The data quality governance process runs through the entire life cycle of the multi-source data fusion gateway. The data bloodline tracking tool records the complete transformation path from the data source to the encrypted unified data view. The data quality rule engine performs periodic data quality assessment, and the evaluation dimensions include accuracy, timeliness, and completeness. The data quality board visualizes the compliance of various quality indicators. The data quality exception triggers an alarm to notify the data governance team. The data governance team analyzes the root cause and optimizes the data access rules or stream ETL processing logic. The data standard management committee is responsible for maintaining and updating the standard data model and field mapping relationship in the central metadata warehouse to adapt to business changes and the needs of new data sources. The integration of the multi-source data fusion gateway and the downstream risk analysis system is achieved through an event-driven architecture. Once a batch of new multi-source data is processed by stream ETL and integrated and encrypted, and successfully persisted to the encrypted unified data view, the multi-source data fusion gateway publishes a data ready event to a highly reliable message queue. The event message contains the data batch identifier, data time range, and data source system identifier. The downstream risk analysis system subscribes to this message topic and triggers the subsequent dynamic graph correlation construction process after receiving the event. This loosely coupled integration method ensures the asynchrony and scalability of the data processing flow and avoids the bottleneck caused by strong dependence between systems.

[0028] Example 2: see Figure 3The process of extracting entity identifiers from the encrypted unified data view is executed by a dedicated entity parsing engine. This engine connects to a distributed database and uses pre-compiled query statements to batch extract invoice codes, enterprise unified social credit codes, personnel ID numbers, and asset serial numbers. These identifiers, after decryption and formatting, serve as the basic entity nodes of the associative hypergraph. Edge attribute extraction is a parallel data processing flow. Invoice issuance timestamps are extracted from tax bureau ledger data and converted to Unix timestamp format. Transaction amounts are obtained after cross-validation of bank statements and invoice amount fields. Logistics trajectory coordinate sequences are reconstructed from GPS signal streams to reconstruct the complete transportation path. Contract term hash values ​​are directly read from blockchain storage to ensure immutability. Entity nodes and edge attributes are assembled in an in-memory graph computation engine. The initial associative hypergraph is represented as an attribute graph containing multiple node and edge types. The graph structure is maintained in memory using adjacency lists and multi-dimensional indexes to enable fast traversal and querying. Learning node embedding representations using the neighborhood aggregation function of a graph neural network is a computationally intensive process. The graph attention network mechanism is instantiated as a differentiable message-passing function. For each entity node in the graph, the graph attention network calculates its attention coefficients with all its first-order neighbors. These attention coefficients are calculated by a small feedforward neural network that takes the feature vectors of the source and target nodes as input. In the neighborhood aggregation stage, each node weights and sums the feature vectors of its neighbors according to the attention coefficients. The aggregated neighbor features are then concatenated with or averaged with the node's own features, and a new node embedding vector is generated through a nonlinear transformation layer. The node embedding vector is a low-dimensional, dense floating-point vector. The geometric distance in the vector space reflects the semantic similarity between nodes. The similarity matrix is ​​constructed by calculating the cosine similarity of all nodes with respect to their embedding vectors. The cosine similarity value ranges from -1 to 1; the closer the value is to 1, the more similar the node features are. Dynamically updating the hypergraph structure based on the similarity matrix is ​​an iterative optimization step. The system sets a dynamically adjusted similarity threshold. For node pairs belonging to different entity types and whose similarity exceeds the threshold, a virtual edge is automatically added to the hypergraph. The virtual edge represents a potential, indirect association relationship. For example, if an individual acts as the legal representative of multiple newly registered companies within a short period, virtual edges will be established between these companies. The community detection algorithm uses the Louvain method on the updated hypergraph. The Louvain method is a hierarchical clustering algorithm based on modularity optimization. The algorithm continuously moves nodes to neighboring communities to maximize the overall modularity, thereby identifying node clusters with tight internal connections and sparse external connections. These clusters correspond to potential fraudulent invoicing groups or units with repeated financing risks.

[0029] Minute-level parsing of entity relationships using graph neural networks requires efficient model scheduling and inference pipelines. The scheduler monitors the frequency of changes in the association hypergraph. When new data injection causes the graph structure to update to a certain scale, the scheduler triggers the graph neural network model to load pre-trained weights. These pre-trained weights are stored in a model repository, which supports version management and fast loading. The graph neural network model architecture includes multiple graph convolutional layers and attention layers. After loading, the model resides in GPU memory to accelerate inference. Multiple rounds of message passing iterations unfold on the complete association hypergraph. In each iteration, each node receives messages from its neighboring nodes. The message content is the current embedding representation of the neighboring node. Nodes use a learnable aggregation function to integrate all received messages and update their own state. Typically, after 2 to 3 iterations, the node embeddings tend to stabilize. The importance score calculation using the attention mechanism is specifically manifested in calculating the attention weight attached to each edge in the associated hypergraph. This attention weight is calculated by an independent attention network, which takes the embedding vectors of the two connected nodes and the attribute features of the edge as input, and outputs a scalar score. Edges with high scores are considered critical paths; for example, the path traversed by an invoice frequently flowing between multiple companies will receive a high attention score. The selected high-weight edges and their associated nodes are extracted to form a subgraph, which is considered the critical path of risk associations. The critical path reveals the core links of abnormal patterns such as fund circulation, centralized invoice issuance, and goods idleness. Real-time clustering of entities on the critical path uses the density-based DBSCAN algorithm. The DBSCAN algorithm can discover clusters of arbitrary shapes and effectively identify noise points. The clustering results divide the entities on the critical path into several groups. The topological fingerprint of the risk group is generated based on the graph structure features of each group. The topological fingerprint features include the average clustering coefficient, average path length, and node degree distribution entropy of the group, which are encoded into a fixed-length feature vector. The real-time matching engine calculates the similarity between the topological fingerprint feature vector and the feature vector in the historical risk pattern library. The historical risk pattern library stores the topological fingerprints of previously confirmed fraudulent invoicing gangs and tax evasion gangs. The matching process uses an approximate nearest neighbor search algorithm to balance accuracy and speed. The successfully matched groups are marked as high-risk gangs and an early warning event is generated.

[0030] The entire graph computation process relies on a distributed graph computation framework. This framework partitions the associative hypergraph and distributes the partitions across multiple computing nodes. Each node is responsible for storing and computing a subgraph, and message passing between nodes is conducted via a high-speed network. A monitoring system tracks the inference latency, memory usage, and convergence of the community detection algorithm in real time. The resource manager dynamically allocates computing resources based on load, such as increasing the parallelism of graph computation tasks during peak business periods. Every graph structure update and risk identification result is recorded in an audit log. The audit log includes the graph version identifier, model inference parameters, a list of identified risk groups, and their topological fingerprints, used to meet compliance requirements and for subsequent model performance evaluation and optimization. Persistent storage of the associative hypergraph uses a snapshot plus incremental log approach. Full graph data snapshots are periodically saved to the distributed file system, and incremental changes during this period are recorded in log form.

[0031] Example 3: The heterogeneous temporal Transformer model performs multi-scale anomaly analysis on the edge weight data of the associated hypergraph. The edge weight data comes from the connection weights between entities in the associated hypergraph. These weights change continuously over time, forming a data stream with temporal characteristics. The edge weight data is converted into a time series with equal intervals. The conversion process involves dividing the time window. The window size is dynamically adjusted according to the business scenario. The short-term window is set to several hours to capture immediate fluctuations, the medium-term window is set to several days to several weeks to observe trends, and the long-term window is set to several months to identify periodic patterns. Each data point in the time series contains the edge weight value, timestamp, and associated entity identifier. Missing values ​​or outliers are filled and smoothed using linear interpolation or moving average methods to ensure the continuity and integrity of the sequence. The encoder layer of the heterogeneous temporal Transformer model receives a preprocessed temporal sequence as input. The encoder consists of multiple identical layers stacked together. Each layer contains a multi-head self-attention mechanism and a feedforward neural network. The self-attention mechanism calculates the correlation strength between each position in the sequence and all other positions, capturing long-term dependencies, such as the transaction patterns of an invoice with multiple companies over a year. Positional encoding information is added to the input sequence to inject temporal order information. Sine and cosine functions are used to generate positional codes. The encoder output is a high-dimensional feature representation of each time point, which contains deep semantic information about the changes in edge weights.

[0032] Anomaly features at different time scales are calculated using a multi-head self-attention mechanism. This mechanism projects the input features into multiple subspaces, each focusing on a different feature dimension. Short-term anomalies are extracted using a small attention window covering the most recent time steps, focusing on transient changes or spikes in edge weights. Mid-term anomalies are extracted using a moving average attention weight, smoothing short-term noise and revealing trends over several weeks. Long-term anomalies utilize a global attention mechanism to analyze the overall shape and periodicity of the sequence. The feature vectors output by each attention head are concatenated and linearly transformed to form a multi-scale feature tensor, the dimension of which is related to the number of time steps and the feature dimension. The multi-scale anomaly features are fused using a weighted summation method, with the following formula: ; in: Represents a dynamic risk score. Representing the Fusion weights for each time scale Representing the Abnormal feature values ​​at each time scale, Value , , These correspond to short-term, medium-term, and long-term scales, respectively.

[0033] The gradient backpropagation process is used not only for parameter optimization during model training but also for generating responsibility node identifiers during inference. It calculates the gradient of the dynamic risk score with respect to the input edge weights. The gradient value reflects the contribution of each edge weight to the risk score. The entity nodes corresponding to edge weights with larger gradient magnitudes are marked as responsibility nodes. The responsibility node identifiers indicate the key links in the risk transmission path. The gradient calculation is based on the chain rule, propagating back from the output layer to the input layer. The gradient values ​​of intermediate layers are obtained through automatic differentiation. The responsibility node identifiers are output in list form, containing node identifiers and corresponding gradient contribution scores. The fund flow heatmap is rendered based on the edge weight gradient values. The color depth in the heatmap is proportional to the gradient magnitude, intuitively displaying areas of concentrated risk. The entire analysis process is deployed on a distributed computing framework that supports model parallelism and data parallelism. It automatically scales computing resources when processing massive amounts of edge-weighted data. Model inference results are written to a risk database in real time. Database index optimization supports fast query and aggregation operations. The monitoring module tracks model performance metrics, including inference latency, accuracy, and resource utilization. When metrics are abnormal, alarms are triggered to notify the operations team. The system regularly backs up model parameters and configuration information to ensure rapid reconstruction of the analysis environment after fault recovery. The training of the heterogeneous time-series Transformer model uses historical edge-weighted data as samples. Sample labels are based on confirmed risk events. The training objective is to minimize the cross-entropy loss between predicted scores and true labels. The optimizer employs... The algorithm adaptively adjusts the learning rate, employs early stopping during training to prevent overfitting, and determines the final model version based on performance on the validation set. Model updates are synchronized with business data update frequencies. New models are rolled over to the production environment after validation in the testing environment. A version control system manages model iteration history and supports rapid rollback to a stable version. The results of multi-scale anomaly analysis are integrated with downstream risk management systems. When dynamic risk scores exceed thresholds, early warning processes are automatically triggered. Responsibility node identification and heatmap visualization assist human decision-making. Analysis logs record the complete data processing path, including input data hashes, model version, output scores, and timestamps, meeting audit and compliance requirements. Log analysis tools regularly generate analysis reports summarizing changes in risk patterns and trends in model performance.

[0034] See Figure 4 The chart illustrates the changes in connection weights between entities over a complete time period, and the dynamic risk score calculated based on these changes. The blue curve represents the time-series changes in edge weight values, reflecting the dynamic fluctuations in connection strength between entities. Three colored areas identify different types of anomaly patterns: the red area shows short-term anomaly patterns, characterized by instantaneous abrupt changes and spikes in edge weight values, which typically correspond to immediate risk events in the business; the yellow area shows medium-term trend anomalies, characterized by persistent changes in edge weight values ​​over a longer period, reflecting abnormal deviations in business trends; and the green area presents long-term periodic anomalies, showing the disruption of the periodic change pattern of edge weight values ​​on a longer time scale. The red dashed line represents the dynamic risk score curve, calculated by fusing anomaly features from different time scales using a multi-head self-attention mechanism. The risk score directly reflects the severity of business risks at the current point in time, providing a quantitative basis for risk warning and response decisions. The chart clearly demonstrates the correlation between edge weight changes and the risk score; when edge weights experience abnormal fluctuations, the risk score rises accordingly, validating the effectiveness of multi-scale anomaly analysis.

[0035] Example 4: The federated collaborative handling mechanism achieves secure risk information sharing and coordinated control in a distributed environment. Under the federated learning framework, the mechanism initializes the local risk identification models of each participant. Participants include the group headquarters risk control system, subsidiary business systems, tax bureau systems, and the bank's core system. Each participant deploys an identical model skeleton locally, containing feature extraction and classification layers, but with initially randomly generated weights. Each participant uses its own encrypted data to calculate risk gradient information, which is the partial derivative of the model parameters with respect to the loss function, reflecting the model's optimization direction under the current data. The process of adding noise to the gradient information based on differential privacy follows strict privacy budget constraints. Noise addition uses a Gaussian mechanism, and the standard deviation of the noise is related to the gradient sensitivity and privacy budget parameters. The privacy budget is controlled by the epsilon-differential privacy parameter, with the epsilon value set at a low level to limit the risk of information leakage from individual data points. Noise addition is completed before gradient aggregation. The gradient vector generated locally by each participant is added to the Gaussian noise vector, preserving the statistical properties of the perturbed gradient information while protecting individual privacy. The zero-knowledge proof protocol verifies the authenticity and compliance of gradient data. Participants need to prove to the coordination server that the gradient information they submit is calculated from compliant local data without tampering with the original data or the calculation process. The proof process uses non-interactive zero-knowledge proof. Participants generate a proof string about the correctness of the gradient calculation. The coordination server verifies the validity of the string without knowing the specific local data. The verified encrypted gradient information is uploaded to the federated learning coordination server through a TLS-based secure channel.

[0036] After aggregating gradients, the coordination server issues a global model update command. Gradient aggregation uses the FedAvg algorithm to perform a weighted average of encrypted gradients from multiple participants. The weights are allocated according to the amount of data from each participant; participants with larger data volumes contribute more to the global model update. The aggregated global gradients are used to update a central global model. The coordination server encodes the global model update command into a message with a specific format, containing the updated parameter tensor, version number, and timestamp. The update command is broadcast to all participants via a message queue. The global model update command triggers the local system's action executor, which is a predefined rule engine. The rule engine parses the received update command and executes specific risk handling actions based on the command code and local business context. For example, modifying invoice status calls the invoice management system's voiding interface, adjusting credit parameters connects to the credit system's limit management module, and red-letter notifications are sent to the tax bureau system via a direct tax connection. The secure sharing of risk gradient information involves cryptographic operations and secure communication. The generation of temporary session keys employs the Elliptic Curve Diffie-Hellman key exchange protocol. Each participant and the coordinating server generates a pair of temporary public and private keys. By exchanging public key materials, they independently calculate an identical shared secret as the session key, which is valid only within a single communication session. The risk gradient information is symmetrically encrypted using the AES-256 algorithm, with GCM mode selected to provide both confidentiality and integrity protection. The encrypted ciphertext, along with the authentication tag, is transmitted to the target participant through a secure channel. Each participant uses its private key to decrypt the session key, and then uses the session key to decrypt the gradient information. This decryption operation is performed within a hardware security module to protect the key materials. The decrypted gradient information is then used for model inference in a memory-isolated sandbox environment. This sandbox environment restricts network access and file write permissions to prevent the leakage of sensitive gradient data. The model inference outputs disposal suggestions, which are returned in structured JSON format. The decision log for the proposed actions is recorded in detail. The decision log includes timestamps, participant identifiers, action types, and risk scoring criteria. Audit trail records are generated based on the decision logs and are used to meet regulatory requirements and support post-event traceability analysis.

[0037] The operation of the federal collaborative processing mechanism relies on a series of configuration parameters that determine the mechanism's privacy protection strength, communication efficiency, and model performance. See Table 1 for a list of key configuration parameters and their typical values.

[0038]

[0039] Table 1: Key Configuration Parameters of the Federal Coordination Mechanism The coordination server's architecture supports horizontal scaling and adopts a microservice architecture. Core components include a gradient aggregator, a model updater, a communication coordinator, and a security manager. The gradient aggregator receives and aggregates encrypted gradients; the model updater maintains the global model version and calculates parameter updates; the communication coordinator manages participant registration, heartbeat detection, and message routing; and the security manager implements cryptographic operations and access control logic. Communication between participants and the coordination server is asynchronous. Participants upload gradients immediately after local training without waiting for other nodes. The coordination server triggers aggregation operations after collecting a sufficient number of gradients. Asynchronous communication improves system throughput and avoids the overhead of synchronous waiting. Fault tolerance mechanisms handle node failures and network partitioning. The coordination server monitors the online status of participants; gradients from failed nodes are excluded from aggregation. Model update commands support retransmission mechanisms to ensure eventual consistency, and checkpoint mechanisms save intermediate states for easy fault recovery. The performance monitoring of the federated collaborative handling mechanism is achieved through a distributed tracing system. This system injects a tracing identifier into each processing step, recording end-to-end latency for gradient calculation, encryption, transmission, aggregation, and model updates. Monitoring metrics include privacy budget consumption, model convergence curves, and communication overhead statistics. These metrics are visualized on a monitoring dashboard, and anomaly detection rules trigger alarms to notify system administrators. The mechanism's security is formally verified; the strength of differential privacy protection is guaranteed by mathematical proof; the security of the zero-knowledge proof protocol is based on the assumption of computational complexity; and the security of the communication channel relies on the TLS protocol standard. Regular security audits ensure consistency between the code and design specifications. Integration of the federated collaborative handling mechanism with existing risk control systems is achieved through an adapter layer. This layer converts data formats and API interfaces, ensuring seamless integration of the new mechanism into the enterprise's existing IT architecture and gradually replacing the traditional centralized risk processing model.

[0040] See Figure 5 This section showcases the performance of risk information sharing and coordinated control. The blue curve represents the trend of model recognition accuracy with training epochs, reflecting the gradual improvement of global model performance during federated learning. The convergence curve of model accuracy demonstrates the effect of model optimization through gradient aggregation while protecting data privacy. The red dashed line represents the consumption of the privacy budget, reflecting the strength of the differential privacy protection mechanism in protecting individual privacy information. As the training epochs increase, the privacy budget is gradually consumed, ensuring that the risk of information leakage of individual data points is effectively controlled. The green dotted line shows the changes in communication resource overhead, indicating the data transmission cost between participants and the coordination server during federated learning. Fluctuations in communication overhead reflect changes in the system's resource requirements at different training stages. The background bar chart shows the number of active participants in each training epoch, reflecting the collaborative work of participating nodes in the distributed system.

[0041] Example 5: Manual review results are obtained by risk control experts who conduct secondary confirmation of early warning cases triggered by the federal collaborative handling mechanism. Experts review the related hypergraphs, risk scores, fund flow heatmaps, and original transaction voucher images provided by the system through the review workbench. Correctly identified risk cases, such as a genuine fraudulent invoicing group, are accurately marked by the system, and the expert clicks the confirmation button to generate a positive reward signal. False alarm cases, such as a normal trading company being misjudged as risky due to its special business model, generate a negative reward signal by clicking the false alarm button. Missed cases are discovered by post-audit and manually entered into the system, generating a stronger negative reward signal. The defined reward function includes multi-dimensional reward indicators. Risk identification accuracy is the core indicator, calculated by comparing system early warnings and expert confirmation results. Handling response time measures the time interval from risk identification to the completion of the handling action. Resource consumption considers the computing and storage resources used by the system during data analysis and model inference. The multi-dimensional reward indicators are combined into a comprehensive reward value through a weighted summation method, with weight coefficients set by domain experts according to business priorities.

[0042] Convolutional neural networks (CNNs) are used to perform sentiment analysis on the review feedback text to extract implicit reward weights. Risk control experts often leave text annotations in the system when making review decisions. These annotations include the reasons for the decision and subjective evaluations, such as "Although the company has concentrated transactions, it has a real logistics background" or "The obvious traces of fund repatriation support the high-risk judgment." The CNN model takes the word vector sequence of these annotation texts as input. The word vectors are obtained through a pre-trained language model. The convolutional kernel performs a sliding convolution operation on the word vector sequence to extract local semantic features. Max pooling layers select the most significant features, and fully connected layers map the features to sentiment polarity scores. The sentiment polarity score acts as an implicit reward weight multiplier, adjusting the base reward value. Positive comments such as "accurate judgment" generate a multiplier greater than 1 to amplify the reward, while negative comments such as "insufficient evidence" generate a multiplier less than 1 to decay the reward. To balance the sparse reward problem, a reward pruning technique is used. This technique sets an upper and lower reward threshold; rewards exceeding the upper threshold are pruned to the upper limit, preventing excessive impact of extreme cases on the model. Reward pruning also ensures stable policy gradient update steps, preventing drastic fluctuations during training. A reinforcement learning-based policy gradient algorithm is used to adjust the parameters of the risk model. This algorithm directly optimizes the policy function of the risk model, which takes the system state as input. State features include topological indicators of the hypergraph, statistical characteristics of edge weight time series, and entity historical behavior profiles. The policy function outputs the probability distribution of different risk judgment actions for the target entity given the state. The action space includes marking as high-risk, marking as low-risk, and suggesting manual review. The optimization objective is to maximize the long-term cumulative reward. The long-term reward considers the sum of reward discounts over multiple future time steps, with a discount factor set to less than 1 to reduce the present value contribution of future rewards. The policy gradient algorithm calculates the gradient of the objective function with respect to the model parameters and updates the parameters along the gradient direction to improve the expected reward. Gradient estimation uses the Monte Carlo method, calculating the average gradient based on a complete set of decision trajectories.

[0043] The rule base is dynamically updated through an online learning framework, which supports hot updates of model parameters and business rules without interrupting risk control services. Updates include adjusting the confidence threshold for risk assessment. This threshold, the minimum score boundary for risk determination, is dynamically adjusted based on the model's accuracy and recall on recent data. The threshold is raised to reduce false positives when accuracy declines and lowered to capture more risks when recall is insufficient. The logical conditions of entity parsing rules are optimized based on changes in entity relationship patterns. For example, a new holding relationship pattern is added to identify related companies under the same ultimate controlling shareholder. Risk pattern feature templates are expanded based on newly emerging risk methods, such as adding a feature vector to identify fraudulent invoicing through cross-border e-commerce platforms. Updated rules and model parameters are deployed to the production environment in real-time after integrity verification. Integrity verification includes syntax checking, logical conflict detection, and performance benchmarking. A blue-green deployment strategy is used, maintaining two identical production environments: one running the old version and the other deploying the new version. Seamless upgrades are achieved through traffic switching, and a version rollback mechanism automatically switches back to the old version when performance metrics are abnormal.

[0044] A specific example illustrates the operation of the knowledge self-updating technology. The system generates a high-risk warning for a company called "ABC Trading Co., Ltd.", with a dynamic risk score of 0.92, triggering an invoice freezing action. Risk control expert Zhang Gong retrieves comprehensive data on ABC Trading Co., Ltd. on the review workbench. The hypergraph shows that the company has frequent financial transactions with three newly registered companies, and bank statements show a clear rapid inflow and outflow characteristic. However, the contract blockchain evidence is complete, and logistics GPS signals show actual goods movement. Based on experience, Zhang Gong judges that this is a typical financing trade based on real transactions rather than fictitious invoicing. Therefore, he clicks the "false alarm" button in the system and enters the text in the comment box: "Abnormal cash flow but real goods flow, complete contract, recommended for key monitoring but not frozen for now." The system records the review result and generates a negative reward signal. The convolutional neural network performs sentiment analysis on the comment text "Abnormal cash flow but real goods flow, complete contract, recommended for key monitoring but not frozen for now." The analysis results identify positive words such as "real" and "complete" and cautious words such as "abnormal" and "monitoring," ultimately outputting a neutral to negative sentiment polarity strength score of 0.8. The reward function integrates accuracy, response time, and sentiment weights to calculate a negative overall reward. The policy gradient algorithm uses this reward to update the parameters of the risk model, adjusting the weighting of patterns like "abnormal cash flow but genuine goods flow," reducing the contribution of risk scoring solely based on cash flow characteristics. Simultaneously, based on this feedback, the rule base update module adds a new rule to the entity parsing rules: for enterprises with genuine logistics and complete contracts but abnormal cash flow, the risk level is adjusted to "monitoring" instead of "high risk," and the confidence threshold is adjusted from 0.9 to 0.95. This update will take effect in the next model release cycle, and the system will handle similar cases more accurately in the future.

[0045] The infrastructure for knowledge self-updating technology includes a versioned feedback knowledge base that stores all historical review records, corresponding reward signals, model parameter versions, and rule snapshots. Data traceability supports analysis of the root causes of model performance changes. An A / B testing framework conducts small-scale experiments before rule updates, redirecting some traffic to the new rule version and comparing business metrics between the old and new versions. Experimental data assists in deciding the timing of large-scale deployment. A monitoring system tracks key indicators of knowledge self-updating technology, including model stability, rule effectiveness rate, and feedback loop latency. Anomaly monitoring ensures the self-learning process remains under control. Knowledge self-updating technology enables the risk control system to continuously evolve from real-world experience, constantly absorbing expert experience to optimize judgment logic and adapt to increasingly complex financial violation methods.

[0046] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0047] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. An invoice risk control management method based on multi-source data linkage, characterized in that, The method includes: The system collects multi-source data from the entire invoice business process in real time through a multi-source data fusion gateway, including tax bureau ledgers, bank statements, logistics GPS signals, contract blockchain evidence, industrial and commercial judicial records, and ERP image data. The system then performs streaming ETL processing and encryption integration on the multi-source data to generate an encrypted unified data view. Based on the encrypted unified data view, a hypergraph of relationships between invoices, enterprises, personnel, and assets is constructed using dynamic graph association technology. Entity relationships are then analyzed in minutes using graph neural networks to identify fraudulent invoice groups and risks of repeated financing. Using risk quantification and tracing techniques, a heterogeneous time-series Transformer model is used to perform multi-scale anomaly analysis on the edge weights of the associated hypergraph, calculate dynamic risk scores, and output risk levels, responsibility node identifiers, and heatmaps of fund flows. When the dynamic risk score exceeds the predefined risk threshold, the risk gradient information is securely shared among the group headquarters, subsidiaries, and tax and banking institutions through a federal collaborative handling mechanism based on differential privacy and zero-knowledge proof, triggering coordinated handling actions such as invoice freezing, red-letter notification, or credit adjustment. Based on the results of manual review of the coordinated response actions, the review feedback is transformed into reinforcement learning rewards through knowledge self-updating technology, and the risk model is fine-tuned online and the rule base is updated hot.

2. The invoice risk control management method based on multi-source data linkage as described in claim 1, characterized in that, The real-time collection of multi-source data throughout the entire invoice business process via the multi-source data fusion gateway includes: Configure multi-source data access rules, specifying the real-time push interface for tax bureau ledger data, the timed retrieval cycle for bank transaction data, the streaming reception frequency for logistics GPS signals, the smart contract event monitoring for contract blockchain notarization, the batch update trigger conditions for industrial and commercial judicial records, and the asynchronous upload channel for ERP image data. Perform format standardization, field mapping, and data validation on the multi-source data to eliminate semantic conflicts between heterogeneous data sources; Homomorphic encryption is performed on the standardized multi-source data using an encryption algorithm to generate the encrypted unified data view, which is then stored in a distributed database.

3. The invoice risk control management method based on multi-source data linkage as described in claim 1, characterized in that, The construction of the hypergraph connecting the four tuples of invoices, enterprises, personnel, and assets includes: Extract invoice codes, enterprise unified social credit codes, personnel ID numbers, and asset serial numbers from the encrypted unified data view as entity nodes, and extract invoice issuance time, transaction amount, logistics trajectory, and contract terms as edge attributes; The embedding representation of entity nodes is learned using the neighborhood aggregation function of graph neural network, and the similarity matrix between nodes is calculated. The hypergraph structure is dynamically updated based on the similarity matrix, virtual edges are added to capture potential associations across entity types, and subgraph patterns of risky groups are identified through community detection algorithms.

4. The invoice risk control management method based on multi-source data linkage as described in claim 3, characterized in that, The minute-level parsing of entity relationships using graph neural networks includes: The scheduling graph neural network model loads pre-trained weights, performs multiple rounds of message passing iterations on the associated hypergraph, and updates node embeddings. An attention mechanism is used to calculate the importance score of edge weights, and high-weight edges are selected as key paths for risk association. Entities on the critical path are clustered in real time to generate topological fingerprints of risky groups and matched with a historical risk pattern library.

5. The invoice risk control management method based on multi-source data linkage as described in claim 1, characterized in that, The multi-scale anomaly analysis of the edge weights of the associated hypergraph includes: The edge weight data is converted into a time series and input into the encoder layer of the heterogeneous temporal Transformer model to capture long-term dependencies. Anomalies at different time scales, including short-term fluctuations, medium-term trends, and long-term cycles, are calculated using a multi-head self-attention mechanism. By fusing multi-scale anomaly features, a dynamic risk score is output using a fully connected layer, and a responsibility node identifier is generated through gradient backpropagation.

6. The invoice risk control management method based on multi-source data linkage as described in claim 1, characterized in that, The triggering of coordinated response actions through the federal collaborative response mechanism includes: Under the federated learning framework, each participant's local model is initialized, and the encrypted risk gradient information is uploaded to the coordination server. Noise is added to gradient information based on differential privacy to meet privacy budget constraints, and the authenticity of gradient data is verified by zero-knowledge proof. After the coordinating server aggregates gradients, it issues a global model update command, triggering the local system's action executor to complete the modification of invoice status or adjustment of credit parameters.

7. The invoice risk control management method based on multi-source data linkage as described in claim 6, characterized in that, The secure sharing risk gradient information includes: A temporary session key is generated, the risk gradient information is symmetrically encrypted, and then transmitted to the participants through a secure channel. After the participants decrypt the gradient information using their private keys, they perform model inference in the sandbox environment and output disposal suggestions. Record decision logs of handling recommendations and generate audit trail records for subsequent tracing.

8. The invoice risk control management method based on multi-source data linkage as described in claim 1, characterized in that, The process of transforming review feedback into reinforcement learning rewards through knowledge self-updating technology, and fine-tuning the risk model online and hot-updating the rule base includes: The results of manual review are converted into reward signals, with positive rewards corresponding to correct risk identification and negative rewards corresponding to false alarms or omissions. The policy gradient algorithm of reinforcement learning is used to adjust the parameters of the risk model to maximize the long-term reward accumulation. The confidence threshold and entity parsing rules of the rule base are dynamically updated through an online learning framework and deployed to the production environment in real time.

9. The invoice risk control management method based on multi-source data linkage as described in claim 8, characterized in that, The process of converting review feedback into reinforcement learning rewards includes: Define a reward function, where risk identification accuracy, response time, and resource consumption are multi-dimensional reward indicators; Using a convolutional neural network to perform sentiment analysis on the review feedback text and extract implicit reward weights; By employing reward pruning techniques, the sparse reward problem can be balanced, ensuring the model's convergence stability.

10. An invoice risk control management system based on multi-source data linkage, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the invoice risk control management method based on multi-source data linkage as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Block chain-based electronic invoice full life cycle management method and management server

    CN114693223A

  • All-electricity invoice management monitoring platform

    CN118071432A

  • Electronic bill data management system and method based on artificial intelligence

    CN119205383A

  • Electronic report bill generation method and device and storage medium

    CN120876124A

  • A system for automatically gathering and managing electric tax bill

    KR1020100132625A

Cited By

  • Financial business risk control management system based on big data model

    CN121724736A

  • Financial network isolation control method and system based on service awareness

    CN121967089A