Transaction risk detection method and device, electronic equipment and storage medium
By constructing a heterogeneous information network and a fusion model, the problems of missed detection of collusion and abnormal fund loops in existing technologies have been solved, achieving more efficient identification of financial transaction risks and anti-fraud prevention.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA CONSTRUCTION BANK
- Filing Date
- 2025-12-08
- Publication Date
- 2026-04-14
AI Technical Summary
Existing methods for monitoring serial number transactions rely on single-point statistical feature analysis, which leads to missed detections of collusive behavior and an inability to capture structural anomalies such as closed-loop fund flows, thus failing to comprehensively cover abnormal risks in financial transactions.
A heterogeneous information network is constructed. By extracting graph structure features and combining unsupervised anomaly detection and graph structure reconstruction models, a comprehensive risk index is generated, which fully covers single-point anomalies and collusive association anomalies.
It improves the comprehensiveness and accuracy of identifying abnormal risks in financial transactions, reduces the probability of missed detection of collusive fraud, accurately captures structural anomalies, and strengthens the ability to prevent and control financial fraud.
Smart Images

Figure CN121860740A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing technology, and in particular to a method and apparatus for detecting transaction risks, an electronic device, and a storage medium. Background Technology
[0002] The serial number, as the unique identifier of physical banknotes in bank cash transactions, is a core element of the financial anti-fraud system. In related technologies, a risk identification system based on statistical features has been constructed through the collaborative operation of transaction flow analysis and unsupervised learning. Specifically, this technical system covers the entire process from data extraction and feature engineering to model inference, including key steps such as transaction amount statistics, time series analysis, and isolated forest detection. Existing methods mostly employ... The batch processing mode identifies isolated outliers that deviate from the norm by calculating characteristics such as the transaction frequency and amount distribution of individual employee accounts.
[0003] Existing methods for monitoring serial number transactions directly employ single-point statistical feature analysis, which may lead to missed detection of collusive behavior or failure to capture structural anomalies such as closed-loop fund flows. Summary of the Invention
[0004] This disclosure provides a method, apparatus, electronic device, and storage medium for detecting transaction risks.
[0005] According to a first aspect of this disclosure, a method for detecting transaction risk is provided, comprising: Based on financial transaction flow data containing transaction subjects, transaction objects and related identification information, a heterogeneous information network is constructed to represent the transaction relationships between multiple types of entities; The graph structure features of the target entity nodes in the heterogeneous information network are extracted, and the graph structure features are used to quantify the structural role and behavior pattern of the target entity nodes in the network. The graph structure features are input into a fusion model for anomaly risk determination. The fusion model combines the outputs of at least an unsupervised anomaly detection model based on node features and an anomaly detection model based on graph structure reconstruction to generate a comprehensive risk index for the target entity node.
[0006] Optionally, the construction of the heterogeneous information network for representing transaction relationships among multiple types of entities includes: The transaction subject, transaction object, and associated identification information are respectively mapped to different types of entity nodes; Based on the relationships recorded in the financial transaction flow data, relationship edges with type and direction are constructed between the corresponding entity nodes.
[0007] Optionally, the entity node includes at least an account node and a physical circulation identifier node; the relationship edge includes at least a transfer edge representing the direction of fund flow and an ownership edge representing the ownership status of the physical circulation identifier.
[0008] Optionally, extracting graph structure features includes extracting the community structure features and path features of the target entity nodes; wherein, The community structure features include the proportion of transaction connections between the target entity node and other entity nodes within the community. The path characteristics include the number of closed-loop funding paths that pass through the target entity node and meet preset constraints.
[0009] Optionally, the fusion model combines the first anomaly score of the node feature-based unsupervised anomaly detection model with the second anomaly score of the graph structure reconstruction-based anomaly detection model using a weighted summation method to obtain the comprehensive risk index.
[0010] Optionally, the method further includes: The financial transaction flow data is divided according to a preset time window, and the heterogeneous information network and its graph structure features for the corresponding time window are updated to perform near real-time risk detection.
[0011] According to a second aspect of this disclosure, a device for detecting transaction risk is provided, comprising: The building unit is also used to construct a heterogeneous information network that represents the transaction relationships between multiple entities based on financial transaction flow data containing transaction subjects, transaction objects and related identification information; The extraction unit is also used to extract the graph structure features of the target entity nodes in the heterogeneous information network, and the graph structure features are used to quantify the structural role and behavior pattern of the target entity nodes in the network. The generation unit is also used to input the graph structure features into the fusion model for anomaly risk determination. The fusion model combines at least the outputs of the unsupervised anomaly detection model based on node features and the anomaly detection model based on graph structure reconstruction to generate a comprehensive risk index for the target entity node.
[0012] Optionally, the building unit is further configured to: The transaction subject, transaction object, and associated identification information are respectively mapped to different types of entity nodes; Based on the relationships recorded in the financial transaction flow data, relationship edges with type and direction are constructed between the corresponding entity nodes.
[0013] Optionally, the entity node includes at least an account node and a physical circulation identifier node; the relationship edge includes at least a transfer edge representing the direction of fund flow and an ownership edge representing the ownership status of the physical circulation identifier.
[0014] Optionally, the extraction unit is further configured to extract the community structure features and path features of the target entity node; wherein, The community structure features include the proportion of transaction connections between the target entity node and other entity nodes within the community. The path characteristics include the number of closed-loop funding paths that pass through the target entity node and meet preset constraints.
[0015] Optionally, the generation unit is further configured to: The comprehensive risk index is obtained by fusing the first anomaly score of the node feature-based unsupervised anomaly detection model with the second anomaly score of the graph structure reconstruction-based anomaly detection model using a weighted summation method.
[0016] Optional, also includes: The updating unit is also used to divide the financial transaction flow data according to a preset time window, and update the heterogeneous information network and its graph structure features in the corresponding time window to perform near real-time risk detection.
[0017] According to a third aspect of this disclosure, an electronic device is provided, comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method described in the first aspect above.
[0018] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are configured to cause the computer to perform the method described in the first aspect above.
[0019] According to a fifth aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method described in the first aspect above.
[0020] The transaction risk detection method, apparatus, electronic device, and storage medium disclosed herein, through this application, can construct a heterogeneous information network representing the transaction relationships of multiple entities based on the transaction subjects, objects, and related identification information in financial transaction flow data (rather than analyzing individual nodes in isolation). By extracting graph structure features, it quantifies the structural roles and behavioral patterns of entities in the network, accurately captures structural features such as fund loops, and then uses a model that integrates unsupervised anomaly detection and graph structure reconstruction detection for comprehensive risk assessment. This comprehensively covers single-point anomalies and collusive association anomalies, making up for the limitations of single-point statistical feature analysis. Therefore, it can solve the technical problems of existing serial number transaction monitoring methods that use single-point statistical feature analysis, leading to missed detection of collusive behavior and inability to capture structural anomalies such as fund loops. It achieves the technical effects of improving the comprehensiveness and accuracy of financial transaction anomaly risk identification, reducing the probability of missed detection of collusive fraud, accurately capturing structural anomalies, strengthening financial anti-fraud prevention and control capabilities, and improving the serial number transaction monitoring system.
[0021] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description
[0022] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein: Figure 1 A schematic flowchart illustrating a method for detecting transaction risk provided in an embodiment of this disclosure; Figure 2 A schematic diagram of the structure of a transaction risk detection device provided in an embodiment of this disclosure; Figure 3 A schematic diagram of the structure of a transaction risk detection device provided in an embodiment of this disclosure; Figure 4 A schematic block diagram of an example electronic device provided for embodiments of this disclosure. Detailed Implementation
[0023] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0024] The following description, with reference to the accompanying drawings, outlines a method, apparatus, electronic device, and storage medium for detecting transaction risks according to embodiments of this disclosure.
[0025] Figure 1 This is a schematic flowchart illustrating a method for detecting transaction risks provided in an embodiment of this disclosure.
[0026] like Figure 1 As shown, the method includes the following steps: Step 101: Based on financial transaction flow data containing transaction subjects, transaction objects and related identification information, construct a heterogeneous information network to represent the transaction relationships between multiple types of entities; The transaction subject typically refers to the entity that initiates or executes the transaction, such as a bank employee account. The transaction object typically refers to the other party in the transaction, such as a customer account. The associated identification information is used to uniquely identify and associate a specific transaction object; in this method, it specifically refers to the serial number. Constructing a heterogeneous information network is fundamental to shifting from single-transaction analysis to complex relationship analysis. Its core lies in transforming discrete transaction data into a structured graph model. This heterogeneous information network consists of a set of nodes and a set of edges. The node set contains various types of entity nodes; for example, the transaction subject is mapped to employee account nodes, and the transaction object is mapped to customer account nodes. Simultaneously, the associated identification information, i.e., each unique serial number, is mapped to a serial number node. Furthermore, other entity types, such as physical or logical branches, can be introduced as branch nodes depending on the actual risk control scenario.
[0027] The edge set is used to define and connect various relationships between these different types of nodes. The main edge type includes flow edges, which connect two account nodes to represent fund transfer relationships. Their direction represents the flow of funds, and transaction-related attribute information can be attached to the edges. In this way, each financial transaction is deconstructed and mapped into nodes and edges in a heterogeneous network, thus forming a complex relationship network that can fully depict the fund flow path, serial number ownership, and transaction operation behavior. Furthermore, to capture the temporal evolution characteristics of transaction relationships, a dynamic dimension is introduced into the construction process. That is, the transaction data is sliced according to a preset time window, and a heterogeneous information network snapshot is constructed for each time window, thus forming a network graph sequence arranged in chronological order, thereby dynamically representing the changes in transaction relationships between multiple entities over time.
[0028] Step 102: Extract the graph structure features of the target entity nodes in the heterogeneous information network. The graph structure features are used to quantify the structural role and behavior pattern of the target entity nodes in the network. The target entity node specifically refers to a particular type of node that requires risk assessment; in this method, it is the employee account node. Extracting graph structure features is a crucial step in transforming the complex topological relationships inherent in heterogeneous information networks into computable and comparable numerical vectors. These features, from a network science perspective, deeply characterize the target node's position and connection patterns within the global network, thereby revealing hidden behavioral patterns that cannot be reflected by simple transaction statistics. The extracted graph structure features mainly include multiple dimensions such as centrality, community, and path characteristics.
[0029] Centrality features are used to measure a node's importance or influence in a network. For example, degree centrality reflects the number of other nodes directly connected to the target entity node, characterizing its transaction activity or breadth of direct connections. Betweenness centrality calculates the frequency with which the node appears on the shortest path between any two other nodes in the network, measuring its control as a bridge or hub for fund transfers. Community features, based on network community partitioning algorithms such as the Louvain algorithm, divide the entire network into several tightly connected communities and calculate the cross-community transaction ratio of the target entity node, i.e., the proportion of edges pointing to different communities in the connections sent by the node. This feature effectively quantifies the node's tendency to connect to different groups; an abnormally high ratio may indicate abnormal behavior such as fund transfers between multiple groups. Path features focus on path patterns formed by the flow of association identifier information. For example, it calculates the number of closed transaction loops that pass through the target entity node and whose length does not exceed a preset threshold. Closed loops may indicate an abnormal pattern of funds circulating among a limited number of nodes. At the same time, it calculates the average shortest path length from the node to all other nodes in its community to measure its information transmission efficiency or tightness in the local network. By systematically extracting the aforementioned multi-dimensional graph structure features, comprehensive and in-depth input can be provided for subsequent anomaly detection models, thereby enabling accurate quantification and anomaly identification of target entity node behavior patterns.
[0030] Step 103: Input the graph structure features into the fusion model for anomaly risk determination. The fusion model combines the outputs of at least the unsupervised anomaly detection model based on node features and the anomaly detection model based on graph structure reconstruction to generate a comprehensive risk index for the target entity node.
[0031] For unsupervised anomaly detection models based on node features, the Isolation Forest model can be used. This model takes the graph structure feature vectors of all target entity nodes as input, constructs isolation trees, and calculates the path length required to isolate each node to evaluate its anomaly level. It outputs an Isolation Forest anomaly score that reflects the deviation of node behavior from the overall pattern. For anomaly detection models based on graph structure reconstruction, the Graph Autoencoder model can be used. This model includes an encoder and a decoder. The encoder typically uses a graph convolutional network to learn the topology and node features of the heterogeneous information network to obtain a low-dimensional vector embedding representation of each target entity node. The decoder then attempts to reconstruct the connection relationships between nodes, i.e., the adjacency matrix, based on this embedding representation. By calculating the reconstruction error of each node, the anomaly of its connection pattern is evaluated, and the Graph Autoencoder anomaly score is output.
[0032] The fusion model integrates the abnormal scores output by the two models through a specific score fusion mechanism. For example, a weighted average strategy is used to assign preset weights to the two scores and perform a linear combination to obtain a comprehensive risk index. This comprehensive risk index combines the ability to detect global outliers based on node feature distribution with the ability to identify abnormal patterns based on network local structure reconstruction. It can more comprehensively and robustly quantify the abnormal risk level of target entity nodes, effectively reducing false positives or false negatives that may be caused by a single model. Finally, by setting a risk threshold or ranking, high-risk target entity nodes can be screened based on this comprehensive risk index to complete the determination of abnormal transaction behavior.
[0033] Optionally, the construction of the heterogeneous information network for representing transaction relationships among multiple types of entities includes: The transaction subject, transaction object, and associated identification information are respectively mapped to different types of entity nodes; Based on the relationships recorded in the financial transaction flow data, relationship edges with type and direction are constructed between the corresponding entity nodes.
[0034] Mapping the transaction subject, transaction object, and associated identification information to different types of entity nodes means that, based on the specific information in the financial transaction flow data, the parties involved in the transaction and the identifiers used for associated transaction items are abstracted into nodes with clearly defined types in a heterogeneous information network.
[0035] Specifically, the transaction entity is typically mapped to employee account nodes, representing the bank's internal personnel account that executes or initiates the transaction; the transaction object is typically mapped to customer account nodes, representing the external customer account with which the transaction takes place; the associated identification information is specifically the serial number, with each unique serial number mapped to an independent serial number node. Furthermore, depending on actual business and risk control needs, other entity types such as physical or logical branches may be included as branch nodes. Based on the relationships recorded in the financial transaction flow data, constructing directional and typed relationship edges between corresponding entity nodes refers to creating directed edges connecting these entity nodes in the network based on the transaction behavior, attribution relationships, and operational facts recorded in the flow data. Each edge type corresponds to a specific relationship semantic. The main relationship edge types include flow edges, attribution edges, and operational edges.
[0036] Flow edges connect two account nodes, representing the flow of funds from one to the other. These edges can carry attributes such as the set of serial numbers involved in the flow, the total amount, and the transaction timestamp. Attribution edges connect serial number nodes to account nodes, indicating that at a specific point in time or within a transaction context, the serial number belongs to that account. Operation edges connect employee account nodes to a specific flow edge, indicating that the employee executed the fund transfer transaction. Through this process of node mapping and relational edge construction, the original linear transaction log data is systematically transformed into a structured, heterogeneous information network containing multiple types of entities and their rich semantic relationships, laying the structural foundation for subsequent graph feature extraction and deep analysis.
[0037] Optionally, the entity node includes at least an account node and a physical circulation identifier node; the relationship edge includes at least a transfer edge representing the direction of fund flow and an ownership edge representing the ownership status of the physical circulation identifier.
[0038] The entity nodes include at least account nodes and physical circulation identification nodes, meaning that in the constructed heterogeneous information network, the core set of node types consists of nodes representing financial accounts and nodes representing unique identifiers of physical circulation items.
[0039] Account nodes specifically map to the transaction subjects and objects in financial transaction flow data, such as bank employee account nodes and external customer account nodes, representing the endpoints of fund storage and transfer. In this specific implementation, the physical circulation identifier node corresponds to the serial number node; each node uniquely represents a physical banknote with a specific serial number, serving as a key abstraction for tracking the physical circulation of currency. The relationship edges, including at least flow edges representing the direction of fund flow and ownership edges representing the ownership status of physical circulation identifiers, refer to the two most basic and core edge types defined and constructed in the network. Flow edges directly connect two account nodes, clearly indicating the direction of funds from the transferring account node to the transferring account node. This edge carries attributes such as transaction amount and time, intuitively depicting the movement path of funds in the network.
[0040] The ownership edge connects the physical circulation identifier node (i.e., the serial number node) to the account node. Its direction indicates which account the physical item represented by the physical circulation identifier belongs to at a given moment or transaction context, thus establishing a dynamic ownership relationship between physical currency and its financial account carrier. By including at least these two node and edge types, the constructed heterogeneous information network can form the basic framework for describing the flow of funds and changes in the ownership of physical currency, capturing the core elements of a transaction: whose money, to whom, and which specific banknote is involved. This provides a structured data foundation for expressing transaction relationships and the circulation of goods for subsequent analysis.
[0041] Optionally, extracting graph structure features includes extracting the community structure features and path features of the target entity nodes; wherein, The community structure features include the proportion of transaction connections between the target entity node and other entity nodes within the community. The path characteristics include the number of closed-loop funding paths that pass through the target entity node and meet preset constraints.
[0042] Extracting the community structure characteristics and path characteristics of the target entity node refers to calculating two key indicators from the topology of the heterogeneous information network that can quantify the node's community connection pattern and specific fund flow path. Specifically, the transaction connection ratio between the target entity node and other entity nodes within other communities, as part of the community structure characteristics, is called the cross-community transaction ratio. Its calculation first requires dividing the entire heterogeneous information network into communities. A community discovery algorithm divides the nodes in the network into several internally tightly connected community substructures. Then, for the target entity node, all its transaction connection edges (flow edges) are counted, and the proportion of edges connecting to different communities is calculated. The level of this characteristic value directly reflects the degree to which the target node plays a role as a fund bridge or intermediary between multiple communities. An abnormally high proportion often suggests that it may be engaged in abnormal cross-group fund integration or transfer behavior.
[0043] The number of closed-loop funding paths that pass through the target entity node and meet preset constraints in the path features specifically refers to the number of closed-loop funding flows in the heterogeneous information network that have the target entity node as a necessary node and whose path length does not exceed a preset number of steps, such as four steps. This number is counted by traversing and analyzing finite-length transaction paths with the target node as the key link. A closed-loop funding path represents a cyclical process in which funds start from a certain account, go through several transfers, and finally return to the original starting account. The preset constraints are mainly used to limit the search depth and time range of the closed-loop path to ensure computational efficiency and business relevance. This feature can effectively reveal potential abnormal patterns of fund circulation among finite nodes, such as potentially indicating fraudulent transactions, inflated transaction volume, or idle fund transfers. By comprehensively calculating these two types of features, the structural role and behavioral patterns of the target entity node in the transaction network can be deeply characterized from the two dimensions of the community traversal of node connections and the cyclical nature of participation, providing key and direct quantitative evidence for subsequent abnormal risk assessment.
[0044] Optionally, the fusion model combines the first anomaly score of the node feature-based unsupervised anomaly detection model with the second anomaly score of the graph structure reconstruction-based anomaly detection model using a weighted summation method to obtain the comprehensive risk index.
[0045] The fusion model combines the first anomaly score of the node-feature-based unsupervised anomaly detection model with the second anomaly score of the graph structure reconstruction-based anomaly detection model using a weighted summation method to obtain the comprehensive risk index. The weighted summation is a linear fusion strategy, the core of which involves assigning appropriate weight coefficients to the first and second anomaly scores respectively, and then summing the weighted scores. The weight allocation can be preset to fixed values based on prior knowledge, such as 0.5% each, to achieve a balanced consideration, or it can be dynamically adjusted based on the model's performance on historical validation datasets to optimize the fusion effect. Before actual calculation, the first and second anomaly scores usually need to be standardized to eliminate differences in the units and distribution ranges of the output scores from different models, ensuring the fairness and effectiveness of the weighted summation.
[0046] The weighted summation process integrates two anomaly detection perspectives based on different principles. The first anomaly score primarily reflects the outlier degree of the target entity node within its individual feature space, while the second anomaly score focuses on measuring the reconfigurability anomaly of the node in the local network connectivity structure. Through linear fusion, the final comprehensive risk index simultaneously encompasses information from both node behavioral anomalies and network structural anomalies, thereby improving the overall robustness and accuracy of risk assessment. This comprehensive risk index, as a single, comparable quantitative value, facilitates subsequent screening and ranking of high-risk nodes based on thresholds.
[0047] Optionally, the method further includes: The financial transaction flow data is divided according to a preset time window, and the heterogeneous information network and its graph structure features for the corresponding time window are updated to perform near real-time risk detection.
[0048] The financial transaction flow data is divided according to a preset time window, and the heterogeneous information network and its graph structure characteristics for the corresponding time window are updated to perform near real-time risk detection. The preset time window is a time interval length pre-set based on the timeliness requirements of actual risk monitoring and the system's processing capacity, such as one hour or one day. Data division refers to using this time window as a benchmark to cut the continuously generated financial transaction flow data into continuous data slices according to transaction timestamps. Each data slice contains all transaction records that occurred within the window period. Updating the heterogeneous information network for the corresponding time window refers to dynamically maintaining the network structure based on the latest acquired data slices for the current time window. Updates can be incremental, that is, adding newly appearing nodes and relationship edges in the current window period to the network snapshot of the previous time window, while correcting changed relationship states or archiving expired relationships according to business rules, thereby generating a snapshot of the heterogeneous information network for the current time window reflecting the latest transaction relationship state.
[0049] The subsequent updating of graph structure features refers to recalculating the graph structure feature values of target entity nodes in the current network topology based on the newly generated network snapshot, in order to capture the latest changes in their behavioral patterns. Through this periodic, time-window-synchronized data partitioning, network update, and feature update process, the entire detection method can continuously operate in sync with the generation rhythm of transaction data. This transforms the basis of risk analysis from static historical snapshots to the flowing, latest state, thereby significantly shortening the risk detection cycle from the traditional T+1 batch model to achieve near real-time hourly or even minute-level risk scanning and early warning, significantly improving the timeliness of risk discovery and enabling timely intervention in potential illegal transactions. This dynamic processing mechanism also allows the model to adapt to the temporal evolution of transaction entity behavioral patterns, enhancing the risk detection system's adaptability to changing environments, while balancing computational load and real-time requirements through periodic windowing processing.
[0050] This invention relates to a method for detecting abnormal serial number transactions based on dynamic heterogeneous graphs. It aims to address the problems of existing technologies, such as a single perspective failing to identify collusion, neglecting relationship information, high false negative rates, and slow response lacking real-time performance. The method efficiently, accurately, and in real-time identifies coordinated violations by bank employees through serial number transactions. The core idea of this method is to construct a dynamically evolving heterogeneous information network from discrete serial number transaction data. The heterogeneous graph allows for various types of nodes and edges, thus comprehensively representing the entities and relationships in the transaction. The dynamic aspect adds a time dimension, using time windows to form a graph sequence to capture network changes over time. First, cash transaction data is obtained from bank data sources, including transaction time, operator ID, local account, counterparty account, transaction amount, and associated serial number set. After data cleaning and standardization, a dynamic heterogeneous information network is constructed. This network includes a set of nodes and a set of edges. The node set includes serial number nodes, employee account nodes, customer account nodes, and branch nodes. The edge set includes flow edges, attribution edges, and operation edges. Flow edges connect two account nodes to represent the flow of funds and include the total amount of the serial number set and the transaction timestamp. Attribution edges connect serial number nodes and account nodes to represent the attribution relationship of serial numbers. Operation edges connect employee account nodes and flow edges to represent transaction operations. The dynamism is reflected by constructing a graph sequence according to time windows, such as daily.
[0051] Next, graph structure features are calculated for employee account nodes in the network. These features quantify node behavior patterns from a global network perspective, including centrality features such as degree centrality and betweenness centrality, community features such as cross-community transaction ratio, and path features such as the number of closed loops and average path length. By calculating these features, abnormal nodes such as bridge accounts and fund transfer stations can be identified. Then, machine learning models are used for anomaly detection. A model fusion strategy is used, combining unsupervised learning models and graph neural network models. The unsupervised learning model outputs anomaly scores based on graph structure feature vectors, while the graph neural network model calculates reconstruction errors as anomaly scores by learning node embeddings and reconstructing the network. Finally, the two anomaly scores are weighted and averaged to obtain a comprehensive anomaly risk score. Finally, a threshold is set to filter out the employee accounts with the highest risk scores to generate the detection results. This method visualizes hidden transaction relationships by constructing a heterogeneous information network, accurately locates abnormal nodes by calculating graph structure features, and improves detection accuracy and efficiency by fusing machine learning models. It achieves near real-time risk scanning, reduces false positive and false negative rates, reduces labor costs, and enhances collusion detection and networked risk visualization capabilities, bringing significant improvements to the bank's risk control system.
[0052] Corresponding to the aforementioned method for detecting transaction risks, this invention also proposes a device for detecting transaction risks. Since the device embodiments of this invention correspond to the method embodiments described above, details not disclosed in the device embodiments can be referred to in the method embodiments, and will not be repeated here.
[0053] Figure 2 This is a schematic diagram of the structure of a transaction risk detection device provided in an embodiment of the present disclosure, as shown below. Figure 2 As shown, it includes: Construction unit 21 is also used to construct a heterogeneous information network to represent the transaction relationship between multiple entities based on financial transaction flow data containing transaction subjects, transaction objects and related identification information; Extraction unit 22 is also used to extract graph structure features of target entity nodes in the heterogeneous information network, wherein the graph structure features are used to quantify the structural role and behavior pattern of the target entity nodes in the network; The generation unit 23 is also used to input the graph structure features into the fusion model for anomaly risk determination. The fusion model combines at least the outputs of the unsupervised anomaly detection model based on node features and the anomaly detection model based on graph structure reconstruction to generate a comprehensive risk index for the target entity node.
[0054] Furthermore, in one possible implementation of this disclosure embodiment, the construction unit 21 is further configured to: The transaction subject, transaction object, and associated identification information are respectively mapped to different types of entity nodes; Based on the relationships recorded in the financial transaction flow data, relationship edges with type and direction are constructed between the corresponding entity nodes.
[0055] Furthermore, in one possible implementation of this disclosure, the entity node includes at least an account node and a physical circulation identifier node; the relation edge includes at least a flow edge representing the direction of fund flow and an ownership edge representing the ownership status of the physical circulation identifier.
[0056] Furthermore, in one possible implementation of this embodiment, the extraction unit 22 is further configured to extract the community structure features and path features of the target entity node; wherein, The community structure features include the proportion of transaction connections between the target entity node and other entity nodes within the community. The path characteristics include the number of closed-loop funding paths that pass through the target entity node and meet preset constraints.
[0057] Furthermore, in one possible implementation of this disclosure embodiment, the generation unit 23 is further configured to: The comprehensive risk index is obtained by fusing the first anomaly score of the node feature-based unsupervised anomaly detection model with the second anomaly score of the graph structure reconstruction-based anomaly detection model using a weighted summation method.
[0058] Furthermore, in one possible implementation of the embodiments of this disclosure, such as Figure 3 The diagram also includes: The updating unit 24 is also used to divide the financial transaction flow data according to a preset time window, and update the heterogeneous information network and its graph structure features in the corresponding time window to perform near real-time risk detection.
[0059] It should be noted that the foregoing explanation of the method embodiments also applies to the apparatus of the embodiments of this disclosure, and the principle is the same. Therefore, the embodiments of this disclosure are not limited thereto.
[0060] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0061] Figure 4 A schematic block diagram of an example electronic device 400 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0062] like Figure 4 As shown, device 400 includes a computing unit 401, which can perform various appropriate actions and processes based on a computer program stored in ROM (Read-Only Memory) 402 or a computer program loaded from storage unit 408 into RAM (Random Access Memory) 403. RAM 403 may also store various programs and data required for the operation of device 400. The computing unit 401, ROM 402, and RAM 403 are interconnected via bus 404. I / O (Input / Output) interface 405 is also connected to bus 404.
[0063] Multiple components in device 400 are connected to I / O interface 405, including: input unit 406, such as keyboard, mouse, etc.; output unit 407, such as various types of monitors, speakers, etc.; storage unit 408, such as disk, optical disk, etc.; and communication unit 409, such as network card, modem, wireless transceiver, etc. Communication unit 409 allows device 400 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0064] The computing unit 401 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, CPUs (Central Processing Units), GPUs (Graphics Processing Units), various special-purpose AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, DSPs (Digital Signal Processors), and any suitable processor, controller, microcontroller, etc. The computing unit 401 performs the various methods and processes described above, such as methods for detecting transaction risks. For example, in some embodiments, the method for detecting transaction risks may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 408. In some embodiments, part or all of the computer program may be loaded and / or installed on device 400 via ROM 402 and / or communication unit 409. When the computer program is loaded into RAM 403 and executed by the computing unit 401, one or more steps of the methods described above may be performed. Alternatively, in other embodiments, the computing unit 401 may be configured to perform the aforementioned method for detecting transaction risks by any other suitable means (e.g., by means of firmware).
[0065] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application-Specific Standard Products), SOCs (System-on-Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0066] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0067] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, RAM, ROM, EPROM (Electrically Programmable Read-Only Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0068] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0069] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include LANs (Local Area Networks), WANs (Wide Area Networks), the Internet, and blockchain networks.
[0070] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service system that addresses the shortcomings of traditional physical hosts and VPS (Virtual Private Server) services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers incorporating blockchain technology.
[0071] It's important to note that artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.
[0072] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0073] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for detecting transaction risk, characterized in that, include: Based on financial transaction flow data containing transaction subjects, transaction objects and related identification information, a heterogeneous information network is constructed to represent the transaction relationships between multiple types of entities; The graph structure features of the target entity nodes in the heterogeneous information network are extracted, and the graph structure features are used to quantify the structural role and behavior pattern of the target entity nodes in the network. The graph structure features are input into a fusion model for anomaly risk determination. The fusion model combines the outputs of at least an unsupervised anomaly detection model based on node features and an anomaly detection model based on graph structure reconstruction to generate a comprehensive risk index for the target entity node.
2. The method according to claim 1, characterized in that, The construction of the heterogeneous information network for representing transaction relationships among multiple types of entities includes: The transaction subject, transaction object, and associated identification information are respectively mapped to different types of entity nodes; Based on the relationships recorded in the financial transaction flow data, relationship edges with type and direction are constructed between the corresponding entity nodes.
3. The method according to claim 2, characterized in that, The entity node includes at least an account node and a physical circulation identifier node; the relation edge includes at least a transfer edge representing the direction of fund flow and an ownership edge representing the ownership status of the physical circulation identifier.
4. The method according to claim 1, characterized in that, Extracting graph structure features includes extracting the community structure features and path features of the target entity nodes; wherein, The community structure features include the proportion of transaction connections between the target entity node and other entity nodes within the community. The path characteristics include the number of closed-loop funding paths that pass through the target entity node and meet preset constraints.
5. The method according to claim 1, characterized in that, The fusion model combines the first anomaly score of the node feature-based unsupervised anomaly detection model with the second anomaly score of the graph structure reconstruction-based anomaly detection model using a weighted summation method to obtain the comprehensive risk index.
6. The method according to any one of claims 1-5, characterized in that, The method further includes: The financial transaction flow data is divided according to a preset time window, and the heterogeneous information network and its graph structure features for the corresponding time window are updated to perform near real-time risk detection.
7. A device for detecting transaction risk, characterized in that, include: The building unit is also used to construct a heterogeneous information network that represents the transaction relationships between multiple entities based on financial transaction flow data containing transaction subjects, transaction objects and related identification information; The extraction unit is also used to extract the graph structure features of the target entity nodes in the heterogeneous information network, and the graph structure features are used to quantify the structural role and behavior pattern of the target entity nodes in the network. The generation unit is also used to input the graph structure features into the fusion model for anomaly risk determination. The fusion model combines at least the outputs of the unsupervised anomaly detection model based on node features and the anomaly detection model based on graph structure reconstruction to generate a comprehensive risk index for the target entity node.
8. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-6.
9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-6.
10. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method according to any one of claims 1-6.