Data matching method and system in big data environment
By establishing a cross-platform transaction behavior database and building a global transaction chart, combining dynamic rule databases and real-time risk thresholds, identifying and warning cross-platform and cross-account dynamic fraud behaviors, cross-platform fraud problems that are difficult to identify and warning in the existing technology are solved, and efficient fraud detection and risk management are achieved.
Patent Information
- Application Number
- CN202510176001.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-23
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing technology is difficult to effectively identify and warn of dynamic fraud across platforms and cross-accounts, and lacks chain risk analysis capabilities, which leads to fraud being detected only after it spreads, increasing economic losses and risk control pressure.
By collecting transaction data from all users’ platforms, pre-processing and trading feature extraction, establishing a cross-platform trading behavior database, eliminating data silos, and identifying potential risk accounts through fraud detection models, building a global transaction chart, generating risk scores and fraud chains, setting a dynamic rule base and real-time risk thresholds, and triggering risk alerts and blocking mechanisms.
It realizes the capture of complex network relationships and hidden patterns, and identifies cross-platform potential risk accounts and fraud chains, improving the accuracy and comprehensiveness of fraud detection, ensuring that even fraudulent behaviors that are disguised are difficult to escape detection.
Smart Images

Figure CN120030363A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of data matching, and in particular to a data matching method and system in a big data environment. Background Art
[0002] The field of data matching technology is a key component of the big data processing and analysis system, and is widely used in many industries such as financial risk control, e-commerce, social networks, healthcare, and supply chain management. Its core goal is to discover the correlation and risk characteristics hidden in complex data relationships through unified identification, integration, and analysis of multi-source heterogeneous data, and provide a scientific basis for business decision-making. Data matching technology involves multiple technical links such as data preprocessing, feature extraction, model construction, and real-time computing, and plays a vital role in dealing with massive, complex, and dynamic data.
[0003] In existing technologies, data between multiple platforms are usually isolated and lack a unified integration and analysis mechanism, which makes it difficult to effectively discover cross-platform relationships. This data island problem greatly limits the comprehensiveness of risk identification. At the same time, many methods only use a single account as an analysis perspective and are unable to identify complex relationships between accounts and potential fraud networks. In particular, traditional methods are unable to cope with dynamic fraud behaviors across accounts and platforms.
[0004] In addition, most existing risk control systems rely on fixed static rules. This single-dimensional analysis method lacks flexibility and adaptability, making it difficult to cope with the continuous evolution of new fraud behaviors, limiting the system's dynamic identification capabilities. Insufficient early warning capabilities are also a major problem. Traditional methods can often only detect abnormal behaviors that have already occurred, but it is difficult to capture potential risk accounts in a timely manner, resulting in fraudulent behaviors being detected only after they spread, increasing economic losses and risk control pressure. More importantly, existing technologies lack chain-based risk analysis capabilities, and are unable to expand from a single high-risk account to the identification of the entire fraud chain, making it difficult to achieve risk control from local to global.
[0005] For carefully disguised fraudulent activities, traditional systems are unable to effectively identify hidden abnormal activities due to their single feature extraction dimension, making it easier for such fraudulent activities to evade detection, further exacerbating the difficulty of risk management. Summary of the invention
[0006] The purpose of the present invention is to provide a data matching method and system in a big data environment, aiming to solve the technical problems existing in the prior art identified in the background technology.
[0007] The present invention is implemented in this way: a data matching method in a big data environment, the method comprising:
[0008] Collect transaction data of all users on all platforms, pre-process the acquired data, extract transaction features from the processed data, establish a cross-platform transaction behavior database, and store the extracted transaction feature data;
[0009] Integrate the extracted transaction feature data through the cross-platform transaction behavior database to eliminate data silos, and perform cross-platform data matching based on account identifiers and transaction features to establish primary account associations, and build fraud detection models to identify cross-platform potential risk accounts;
[0010] With accounts as nodes and transaction records as edges, a global transaction graph is constructed, and the features of nodes, edges and their connections are extracted. The fraud detection model identifies the existing fraud patterns, generates potential risk scores for risky accounts and potential fraud chains, and identifies hidden abnormal activities.
[0011] Set a potential risk threshold, extract transaction information and account features with potential risk scores higher than the potential risk threshold based on the identification results of fraud patterns, and convert them into a dynamic rule base;
[0012] Input real-time transaction data into the dynamic rule base for rapid matching, identify high-risk transactions and generate real-time risk scores, and set real-time risk thresholds. For transactions with real-time risk scores higher than the real-time risk thresholds, trigger risk alerts and real-time blocking mechanisms, and record transaction behaviors;
[0013] Utilizing risky transaction records in the real-time matching process, new fraud behavior patterns are fed back into the global transaction graph and dynamic rule base to update the global transaction graph and dynamic rule base.
[0014] As a further solution of the present invention, the transaction data of all platforms of the user are collected, and the acquired data is preprocessed, and transaction features are extracted from the processed data, and a cross-platform transaction behavior database is established to store the extracted transaction feature data, which specifically includes:
[0015] Obtain transaction data from all trading platforms, covering the user's full life cycle transaction behavior, and identify and distinguish data types from different sources, including: basic account information, transaction records, device information, and behavior logs;
[0016] Map fields that describe the same attributes in different platforms to unified field names, standardize numerical features, and generate time series features based on transaction timestamps;
[0017] Extract transaction behavior features and conduct cross-platform data integration, integrating transaction records from different platforms into a unified cross-platform transaction behavior database.
[0018] As a further solution of the present invention, the establishment of a fraud detection model to identify cross-platform potential risk accounts specifically includes:
[0019] Build fraud detection models based on historical data and flagged fraudulent behavior;
[0020] The fraud detection model is used to analyze the transaction behavior characteristics stored in the cross-platform transaction behavior database, identify and mark risk accounts with abnormal transactions based on transaction amount and frequency, device switching and geographical span.
[0021] As a further solution of the present invention, the global transaction graph is constructed, and the features of nodes, edges and their connection relationships are extracted, the existing fraud patterns are identified through the fraud detection model, the potential risk scores of risk accounts and potential fraud chains are generated, and hidden abnormal activities are identified, which specifically include:
[0022] Extract data from the cross-platform transaction behavior database, use accounts as nodes, each node is accompanied by basic account information, and transaction records are used as edges. The direction of the edge indicates the direction of capital flow, and each edge is accompanied by relevant transaction features, so as to construct a global transaction graph;
[0023] Calculate the static and dynamic characteristics of each node, analyze the ratio of edges to nodes in the graph, calculate the concentration of transactions, and identify sub-communities and high-frequency cyclic transaction areas in the transaction network;
[0024] Analyze nodes and edges to identify abnormal behaviors, and generate potential risk scores for each node through fraud detection models based on node characteristics and associated paths;
[0025] Construct a fraud chain based on the associated paths in the global transaction graph, set a potential risk threshold, define nodes with potential risk scores higher than the potential risk threshold as high-risk nodes, and read chains with multiple high-risk nodes from the fraud chain;
[0026] Based on the chain of high-risk nodes and potential risk scores, accounts are grouped to obtain high-risk account groups, and the correlation between high-risk account groups and fraud chains is analyzed to mark hidden abnormal activities.
[0027] As a further solution of the present invention, the identification result based on the fraud pattern extracts the transaction information and account features with potential risk scores higher than the potential risk threshold and converts them into a dynamic rule base, specifically including:
[0028] For account transaction features with potential risk scores higher than the threshold, including transaction records, extract time features, amount features, geographic features, and behavior features, and extract abnormal transaction behavior features for high-risk accounts;
[0029] Based on the abnormal transaction behavior characteristics and account transaction characteristics, dynamic rule templates are generated, including transaction amount rules, high-frequency transaction rules within a time period, geographic location jump rules and device switching rules.
[0030] As a further solution of the present invention, the real-time transaction data is input into a dynamic rule base for rapid matching, high-risk transactions are identified and real-time risk scores are generated, and a real-time risk threshold is set. For transactions with real-time risk scores higher than the real-time risk threshold, risk alarms and real-time blocking mechanisms are triggered, and transaction behaviors are recorded;
[0031] Access transaction data from various platforms and match the accessed real-time transaction data with the rules in the dynamic rule base one by one;
[0032] Generate real-time risk scores based on the number of rules matched to the transaction, rule weights, and the degree of deviation of transaction characteristics :
[0033]
[0034] in, represents the risk score of the current transaction when it meets rule i, W i is the weight of rule i, and n represents the number of rules in the dynamic rule base matched by the current transaction;
[0035] According to business needs and historical data analysis, set a warning real-time risk threshold and a blocking real-time risk threshold respectively. The warning real-time risk threshold is smaller than the blocking real-time risk threshold. Transactions with real-time risk scores exceeding the warning real-time risk threshold are defined as high-risk transactions.
[0036] For transactions that are above the warning real-time risk threshold but below the blocking real-time risk threshold, a system alarm is triggered and the transaction is marked as;
[0037] For transactions that are higher than the real-time risk threshold, the transactions will be blocked and the related accounts will be frozen.
[0038] Another object of the present invention is to provide a data matching system in a big data environment, the system comprising:
[0039] The transaction feature extraction module is used to collect the transaction data of all platforms of users, pre-process the acquired data, extract transaction features from the processed data, establish a cross-platform transaction behavior database, and store the extracted transaction feature data;
[0040] The cross-platform data integration module is used to integrate the extracted transaction feature data through the cross-platform transaction behavior database to eliminate data silos, and to match cross-platform data based on account identifiers and transaction features to establish primary account associations, and to establish fraud detection models to identify cross-platform potential risk accounts;
[0041] The global transaction graph construction module is used to construct a global transaction graph with accounts as nodes and transaction records as edges, and extract the features of nodes, edges and their connection relationships. It uses the fraud detection model to identify existing fraud patterns, generate potential risk scores for risky accounts and potential fraud chains, and identify hidden abnormal activities.
[0042] A dynamic rule base generation module is used to set a potential risk threshold, extract transaction information and account features with potential risk scores higher than the potential risk threshold based on the fraud pattern recognition results, and convert them into a dynamic rule base;
[0043] The real-time transaction matching module is used to input the real-time transaction data into the dynamic rule base for rapid matching, identify high-risk transactions and generate real-time risk scores, and set real-time risk thresholds. For transactions with real-time risk scores higher than the real-time risk thresholds, risk alarms and real-time blocking mechanisms are triggered, and transaction behaviors are recorded;
[0044] The fraud pattern feedback module is used to utilize the risk transaction records in the real-time matching process to feed back new fraud behavior patterns to the global transaction graph and dynamic rule base, and update the global transaction graph and dynamic rule base.
[0045] The beneficial effects of the present invention are:
[0046] This method can capture complex network relationships and hidden patterns that cannot be discovered based on table data alone, and is particularly effective when dealing with cross-platform, multi-account related fund flows. By combining static and dynamic features, this method can not only identify current abnormal behaviors, but also discover accounts with potential risks, which is of great significance for early warning of fraud.
[0047] In addition, the construction of the fraud chain provides the system with a higher level of analytical capabilities, so that the risk of a single account can be associated with the entire chain, achieving risk management from the individual to the whole. This is particularly important for the identification of complex and diverse fraudulent behaviors, and can significantly improve the accuracy and comprehensiveness of fraud detection. At the same time, based on the correlation analysis between high-risk account groups and chains, the system can mark hidden abnormal activities, ensuring that even well-disguised fraudulent behaviors are difficult to escape detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 A flowchart of a data matching method in a big data environment provided by an embodiment of the present invention;
[0049] Figure 2 A flowchart for collecting transaction data of all users on all platforms provided by an embodiment of the present invention;
[0050] Figure 3 A flowchart of establishing a fraud detection model and identifying cross-platform potential risk accounts provided by an embodiment of the present invention;
[0051] Figure 4 A flowchart for generating potential risk scores for risky accounts and potential fraud chains and identifying hidden abnormal activities provided by an embodiment of the present invention;
[0052] Figure 5 A flowchart for extracting transaction information and account features with potential risk scores higher than a potential risk threshold and converting them into a dynamic rule base provided by an embodiment of the present invention;
[0053] Figure 6 A flow chart of triggering a risk alert and real-time blocking mechanism for a transaction whose real-time risk score is higher than a real-time risk threshold provided by an embodiment of the present invention;
[0054] Figure 7 A structural block diagram of a data matching system in a big data environment provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0055] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0056] It is understood that the terms "first", "second", etc. used in this application may be used herein to describe various elements, but unless otherwise specified, these elements are not limited by these terms. These terms are only used to distinguish a first element from another element. For example, without departing from the scope of this application, a first xx script may be referred to as a second xx script, and similarly, a second xx script may be referred to as a first xx script.
[0057] Figure 1 A flowchart of a data matching method in a big data environment provided by an embodiment of the present invention is shown in FIG. Figure 1 As shown, the method includes:
[0058] S100, collecting transaction data of all platforms of the user, and preprocessing the acquired data, extracting transaction features from the processed data, and establishing a cross-platform transaction behavior database to store the extracted transaction feature data;
[0059] This step will obtain user transaction data on all trading platforms, which is the basis of this step. The coverage must include the user's full life cycle transaction behavior, which means not only obtaining real-time transaction data, but also historical transaction records, from account creation to the latest current transaction behavior. This comprehensive data collection method can ensure the integrity and accuracy of user behavior. In addition, data sources include but are not limited to basic account information (such as account name, account ID, binding information, etc.), transaction records (such as transaction amount, transaction time, counterparty, etc.), device information (such as device model, device IP address, operating system, etc.) and behavior logs (such as user login times, click behavior, operating habits, etc.).
[0060] In data preprocessing, data from different trading platforms need to be standardized. Since the data structures and field names of different platforms may differ, field mapping must be used to align fields that express the same attributes into a unified field name.
[0061] For example, one platform may name the user's transaction amount field "amount", while another platform may name it "transaction_value", which needs to be unified through field mapping rules.
[0062] For numerical features, standardization is also required, such as normalizing the transaction amount to eliminate differences in transaction size between different platforms or users. In addition, generating time series features based on transaction timestamps is a key step, which allows us to better characterize the time patterns of users' transaction behaviors and discover users' transaction habits, preferences, and abnormal behaviors in the time dimension.
[0063] Through feature extraction, we conduct in-depth analysis of transaction behaviors. Transaction behavior feature extraction includes the extraction of multi-dimensional features such as user transaction frequency, amount distribution, transaction object diversity, transaction period characteristics, etc. On this basis, we integrate transaction records from different platforms into a unified cross-platform transaction behavior database.
[0064] like Figure 2 As shown, the transaction data of all platforms of the user are collected, and the acquired data is preprocessed, and transaction features are extracted from the processed data, and a cross-platform transaction behavior database is established to store the extracted transaction feature data, which specifically includes:
[0065] S110, obtaining transaction data of all trading platforms, covering the user's full life cycle transaction behavior, and identifying and distinguishing data types from different sources, including: basic account information, transaction records, device information and behavior logs;
[0066] S120, mapping fields that express the same attributes in different platforms to a unified field name, and standardizing the numerical features, while generating time series features according to the transaction timestamp;
[0067] S130, extracting transaction behavior features and performing cross-platform data integration, integrating transaction records from different platforms into a unified cross-platform transaction behavior database.
[0068] S200, integrating the extracted transaction feature data through the cross-platform transaction behavior database to eliminate data silos, and matching cross-platform data based on account identifiers and transaction features to establish primary account associations, and establish a fraud detection model to identify cross-platform potential risk accounts;
[0069] This step will build an efficient and accurate fraud detection model based on historical data and combined with marked fraudulent behaviors. In the process of model building, machine learning and big data analysis technologies are used to extract characteristic patterns of fraudulent behaviors through in-depth research on known fraudulent behaviors, including abnormal fluctuations in transaction amounts, irregularities in transaction frequency, suddenness of device switching, and characteristics of large transaction geographic locations. These characteristic patterns provide clear directions and high-quality samples for model training.
[0070] After the fraud detection model was built, it was used to conduct a comprehensive analysis of the data in the cross-platform transaction behavior database. The focus of the analysis was to discover possible abnormal patterns in transaction behavior and identify potential risky accounts.
[0071] For example, analysis based on transaction amount and frequency can help identify accounts with a sudden surge in transaction amount and abnormally high transaction frequency, as these may be associated with fraudulent activities such as money laundering and fraud; analysis based on device switching can identify situations where users frequently change devices in an abnormally short period of time, which may be a signal that hackers are using multiple devices to take over accounts; analysis based on geographic span can detect possible account theft or false transactions by identifying users appearing in multiple geographical locations far apart within unreasonable time periods.
[0072] These analytical dimensions are closely integrated with the actual characteristics of fraudulent behavior and can effectively improve the accuracy of risk identification.
[0073] Through the above multi-dimensional analysis, the system can mark high-risk accounts and establish primary account association relationships. This cross-platform account matching can not only detect abnormal behavior within a single platform, but also associate seemingly unrelated accounts on multiple platforms and identify potential fraud chains hidden between multiple platforms. In this way, the system realizes a panoramic identification of risky accounts, providing an important foundation for the subsequent construction of a dynamic rule base and real-time transaction monitoring.
[0074] like Figure 3 As shown, the establishment of a fraud detection model to identify potential risk accounts across platforms specifically includes:
[0075] S210, building a fraud detection model based on historical data and marked fraudulent behaviors;
[0076] S220, analyzing the transaction behavior characteristics stored in the cross-platform transaction behavior database through a fraud detection model, identifying risk accounts with abnormal transactions based on transaction amount and frequency, device switching and geographical span, and marking them.
[0077] S300, with accounts as nodes and transaction records as edges, constructs a global transaction graph and extracts the features of nodes, edges and their connections. It uses the fraud detection model to identify existing fraud patterns, generate potential risk scores for risky accounts and potential fraud chains, and identify hidden abnormal activities.
[0078] This step extracts data from the cross-platform transaction behavior database, with accounts as nodes in the graph. Each node carries basic account information, such as account type, historical transaction statistics, device information, etc. Transaction records are represented as edges in the graph, with the direction of the edge pointing to the inflow or outflow of funds. Each edge carries transaction-related features, such as transaction amount, transaction time, transaction category, etc. This way of constructing a graph model can clearly show the relationship between fund flows between accounts, capturing not only a single transaction behavior, but also the structure of the overall transaction network.
[0079] After the global transaction graph is constructed, the system further calculates the static and dynamic features of the nodes and edges.
[0080] For example, static features include the account's total transaction amount, total number of transactions, income-to-expenditure ratio, number of connections with other accounts, etc., while dynamic features involve the time dimension, such as the transaction frequency in the recent period, fluctuations in transaction amounts in a specific time period, etc.
[0081] In addition, by analyzing the ratio of edges to nodes in the graph, the concentration of transactions can be calculated, revealing whether certain accounts have excessive inflows or outflows of funds. At the same time, network partition analysis is performed on the global transaction graph to identify subcommunities and high-frequency cyclic transaction areas in the transaction network. The identification of subcommunities helps reveal the aggregation or isolation between transaction accounts, while high-frequency cyclic transaction areas may point to abnormal behaviors such as money laundering or false transactions.
[0082] After completing the network structure analysis, the system further identifies abnormal behaviors based on the characteristics of nodes and edges, combined with the associated path information. Through the fraud detection model, a potential risk score is generated for each node.
[0083] The calculation basis of the potential risk score includes not only the characteristic value of a single node, but also takes into account the characteristics of the associated accounts on the path where the node is located. For example, the risk score of a node may be increased by the number of high-risk nodes it is directly connected to. This method can effectively identify potential risks that cannot be explained by isolated behavior. In addition, based on the associated paths in the global transaction graph, the system can build a fraud chain. In the process of identifying fraud chains, a potential risk threshold is set, and nodes with potential risk scores above the threshold are defined as high-risk nodes. If there are multiple high-risk nodes in the chain and the interaction behavior of these nodes conforms to known fraud patterns, the chain is marked as a high-risk fraud chain.
[0084] The system groups accounts based on high-risk nodes and their associated chains to form high-risk account groups. The system then analyzes the correlation between the high-risk account group and the fraud chain to mark hidden abnormal activities. For example, if an account is not directly involved in high-frequency trading, but it has frequent transactions with multiple high-risk nodes, then the account may also be an important link hidden in the chain. In this way, the system can dig out deep-level abnormal behaviors and risky accounts from seemingly ordinary transaction networks, forming a precise positioning of fraudulent behavior.
[0085] This method can capture complex network relationships and hidden patterns that cannot be discovered based on table data alone, and is particularly effective when dealing with cross-platform, multi-account related fund flows. By combining static and dynamic features, this method can not only identify current abnormal behaviors, but also discover accounts with potential risks, which is of great significance for early warning of fraud. In addition, the construction of the fraud chain provides the system with a higher level of analytical capabilities, so that the risk of a single account can be associated with the entire chain, realizing risk management from the individual to the whole. This is particularly important for the identification of complex and diverse fraudulent behaviors, and can significantly improve the accuracy and comprehensiveness of fraud detection. At the same time, based on the correlation analysis between high-risk account groups and chains, the system can mark hidden abnormal activities to ensure that even well-disguised fraudulent behaviors are difficult to escape detection.
[0086] like Figure 4 As shown, the global transaction graph is constructed, and the features of nodes, edges and their connection relationships are extracted. The fraud detection model is used to identify the existing fraud patterns, generate the potential risk scores of risk accounts and potential fraud chains, and identify hidden abnormal activities, including:
[0087] S310, extracting data from a cross-platform transaction behavior database, taking accounts as nodes, each node with basic account information, taking transaction records as edges, the direction of the edge indicating the direction of fund flow, and each edge with relevant transaction features, thereby constructing a global transaction graph;
[0088] S320, calculating the static characteristics and dynamic characteristics of each node, analyzing the ratio of edges to nodes in the graph, calculating the concentration of transactions, and identifying sub-communities and high-frequency cycle transaction areas in the transaction network;
[0089] S330, analyzing nodes and edges, identifying abnormal behaviors, and generating a potential risk score for each node through a fraud detection model based on node features and associated paths;
[0090] S340, constructing a fraud chain based on the associated paths in the global transaction graph, setting a potential risk threshold, defining nodes with potential risk scores higher than the potential risk threshold as high-risk nodes, and reading a chain with multiple high-risk nodes from the fraud chain;
[0091] S350, based on the chain of high-risk nodes and potential risk scores, groups accounts to obtain high-risk account groups, analyzes the correlation between high-risk account groups and fraud chains, and marks hidden abnormal activities.
[0092] S400, setting a potential risk threshold, extracting transaction information and account features with potential risk scores higher than the potential risk threshold based on the identification results of the fraud pattern, and converting them into a dynamic rule base;
[0093] This step will screen out accounts with potential risk scores higher than the set threshold and their related transaction behaviors as the focus of analysis. These accounts usually show abnormal transaction behavior patterns, such as sharp fluctuations in transaction amounts, sudden increases in transaction frequency, large jumps in transaction geographic locations, and frequent device switching. In order to fully characterize the characteristics of these high-risk accounts, the system will extract features from their transaction records from multiple dimensions. The extraction of time features includes analyzing the transaction time distribution of accounts, such as whether there are abnormal nighttime transaction peaks or multiple transactions in a very short period of time; the amount feature focuses on the distribution of transaction amounts, whether there are abnormal large transactions or behaviors that are seriously inconsistent with the historical transaction amounts of the account; the geographic feature analyzes the geographical distribution of transactions to determine whether there are transactions in multiple distant locations in a short period of time; the behavioral feature focuses on the operating behavior of the account, such as whether the device is frequently switched, or whether the login and logout are repeated. The extraction of these features can help the system characterize the behavior patterns of high-risk accounts in a more refined way.
[0094] After obtaining the detailed characteristics of high-risk accounts, the system further classifies and analyzes the abnormal characteristics in their transaction behaviors, and uses these abnormal characteristics to generate dynamic rule templates. The dynamic rule template is the core of the rule base, which mainly includes: transaction amount rules, such as identifying transactions where the amount of a single transaction exceeds a certain multiple of the account's historical average; high-frequency transaction rules within a time period, such as the occurrence of multiple consecutive transactions in a very short period of time; geographic location jump rules, such as transaction activities across multiple geographic locations within an unreasonably expected time range; device switching rules, such as the frequent change of devices or IP addresses by user accounts in a short period of time. The design of the dynamic rule template fully considers the diversity and complexity of fraud patterns to ensure that common fraud behaviors and potential new fraud patterns can be covered.
[0095] In addition, the generation process of dynamic rule templates is not static. When generating rule templates, the system will also combine historical data and the latest fraud behavior patterns to ensure the flexibility and effectiveness of the dynamic rule base. This dynamic update mechanism ensures that the rule base can continuously adapt to new risk scenarios and always maintain sensitivity to high-risk transactions.
[0096] like Figure 5 As shown, the identification result based on the fraud pattern extracts the transaction information and account features with potential risk scores higher than the potential risk threshold and converts them into a dynamic rule base, which specifically includes:
[0097] S410, extracting transaction features of accounts with potential risk scores higher than a threshold, including: transaction records, time features, amount features, geographic features, and behavior features, and extracting abnormal transaction behavior features for high-risk accounts;
[0098] S420, based on the abnormal transaction behavior characteristics and account transaction characteristics, a dynamic rule template is generated, including a transaction amount rule, a high-frequency transaction rule within a time period, a geographic location jump rule, and a device switching rule.
[0099] S500, inputs real-time transaction data into the dynamic rule base for rapid matching, identifies high-risk transactions and generates real-time risk scores, and sets real-time risk thresholds. For transactions with real-time risk scores higher than the real-time risk thresholds, triggers risk alerts and real-time blocking mechanisms, and records transaction behaviors;
[0100] This step can connect real-time transaction data from various platforms to the system. In this process, the system must ensure efficient data transmission and compatibility. Since real-time transaction data may come from multiple different platforms and in different formats, it is necessary to receive and parse the data through a unified interface protocol. After completing data preprocessing, each transaction data will be entered into the dynamic rule base for matching operations. The dynamic rule base was previously generated by the previous step, including rule definitions for multi-dimensional features such as transaction amount, time, geographic location, and device switching. The matching process is the process of comparing each transaction data with each rule in the rule base. Each rule not only defines the specific characteristics of a potential risk behavior, but also comes with weights and constraints to characterize its contribution to the risk score. For example, a rule may define the behavioral characteristics of "cross-border transactions in a unified account within a short period of time" and assign a higher weight to indicate its high sensitivity to fraud.
[0101] After the transaction data has completed the rule matching, the system will generate a real-time risk score based on the number of rules matched to the transaction, the rule weights, and the degree to which the transaction characteristics deviate from the rules. The calculation of the real-time risk score is a comprehensive process that dynamically combines the results of multiple matching rules. For transactions that meet multiple high-weight rules, their risk scores will be significantly improved; while if only a small number of low-weight rules are met, the risk score may remain at a lower level. At the same time, the calculation of the degree of deviation also further adjusts the risk score. For example, the more significant the deviation of the transaction behavior from the rule characteristics, the higher the score may be. The final real-time risk score is a quantitative indicator that can intuitively reflect the potential risk level of the transaction.
[0102] Next, the system will set two risk thresholds based on business needs and historical data analysis: a warning real-time risk threshold and a blocking real-time risk threshold. The warning threshold is lower than the blocking threshold to achieve risk classification management. For transactions with real-time risk scores higher than the warning threshold but lower than the blocking threshold, the system will define them as high-risk transactions and trigger an alarm mechanism. This alarm mechanism usually informs the risk control team in various forms (such as notification messages, email reminders, etc.), and marks the transaction as a "transaction to be observed" and records the detailed characteristics of the transaction for subsequent analysis and processing. This processing method reserves intervention time for the risk control team to avoid excessive intervention problems caused by false alarms.
[0103] For transactions whose real-time risk scores exceed the blocking threshold, the system will directly trigger the blocking mechanism. This mechanism will not only immediately terminate the relevant transactions, but also freeze the accounts associated with the transaction to prevent further financial losses or the spread of fraud. The implementation of the blocking mechanism requires a high degree of real-time and accuracy, so the system will give priority to high-performance computing resources during this process to ensure immediate response to high-risk transactions. At the same time, all blocked transactions and accounts will be recorded and further sent to the subsequent risk analysis link.
[0104] like Figure 6 As shown, the real-time transaction data is input into the dynamic rule base for rapid matching, high-risk transactions are identified and real-time risk scores are generated, and real-time risk thresholds are set. For transactions with real-time risk scores higher than the real-time risk thresholds, risk alarms and real-time blocking mechanisms are triggered, and transaction behaviors are recorded;
[0105] S510, accessing transaction data from each platform, and matching the accessed real-time transaction data with the rules in the dynamic rule base one by one;
[0106] S520: Generate a real-time risk score based on the number of rules matched by the transaction, the rule weights, and the degree of deviation of transaction characteristics :
[0107]
[0108] in, represents the risk score of the current transaction when it meets rule i, W i is the weight of rule i, and n represents the number of rules in the dynamic rule base matched by the current transaction;
[0109] S530, according to business needs and historical data analysis, set a warning real-time risk threshold and a blocking real-time risk threshold respectively, and the warning real-time risk threshold is less than the blocking real-time risk threshold, and the transaction corresponding to the real-time risk score exceeding the warning real-time risk threshold is defined as a high-risk transaction;
[0110] S540, for transactions that are higher than the warning real-time risk threshold but have not reached the blocking real-time risk threshold, trigger a system alarm and mark the transaction as;
[0111] For transactions that are higher than the real-time risk threshold, the transactions will be blocked and the related accounts will be frozen.
[0112] S600, utilizes the risk transaction records in the real-time matching process to feed back new fraud behavior patterns into the global transaction graph and dynamic rule base, and updates the global transaction graph and dynamic rule base.
[0113] Figure 7A structural block diagram of a data matching system in a big data environment provided by an embodiment of the present invention, such as Figure 7 As shown, the system comprises:
[0114] The transaction feature extraction module 100 is used to collect the transaction data of all platforms of the user, pre-process the acquired data, extract transaction features from the processed data, establish a cross-platform transaction behavior database, and store the extracted transaction feature data;
[0115] The cross-platform data integration module 200 is used to integrate the extracted transaction feature data through the cross-platform transaction behavior database to eliminate data silos, and to perform cross-platform data matching based on account identifiers and transaction features to establish primary account associations, and to establish a fraud detection model to identify cross-platform potential risk accounts;
[0116] A global transaction graph construction module 300 is used to construct a global transaction graph with accounts as nodes and transaction records as edges, and to extract features of nodes, edges and their connection relationships, identify existing fraud patterns through a fraud detection model, generate potential risk scores for risky accounts and potential fraud chains, and identify hidden abnormal activities;
[0117] A dynamic rule base generation module 400 is used to set a potential risk threshold, extract transaction information and account features with a potential risk score higher than the potential risk threshold based on the identification result of the fraud pattern, and convert them into a dynamic rule base;
[0118] The real-time transaction matching module 500 is used to input the real-time transaction data into the dynamic rule base for rapid matching, identify high-risk transactions and generate real-time risk scores, and set real-time risk thresholds. For transactions with real-time risk scores higher than the real-time risk thresholds, risk alarms and real-time blocking mechanisms are triggered, and transaction behaviors are recorded;
[0119] The fraud pattern feedback module 600 is used to utilize the risk transaction records in the real-time matching process to feed back new fraud behavior patterns to the global transaction graph and the dynamic rule base, and update the global transaction graph and the dynamic rule base.
[0120] It should be understood that, although each step in the flow chart of each embodiment of the present invention is shown in sequence according to the indication of the arrow, these steps are not necessarily performed in sequence according to the order indicated by the arrow. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be performed in other orders. Moreover, at least a portion of the steps in each embodiment may include a plurality of sub-steps or a plurality of stages, and these sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these sub-steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of other steps or sub-steps or stages of other steps.
[0121] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0122] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0123] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.
[0124] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A data matching method in a big data environment, characterized in that: The method comprises: Collect transaction data of all users on all platforms, pre-process the acquired data, extract transaction features from the processed data, establish a cross-platform transaction behavior database, and store the extracted transaction feature data; Integrate the extracted transaction feature data through the cross-platform transaction behavior database to eliminate data silos, and perform cross-platform data matching based on account identifiers and transaction features to establish primary account associations, and build fraud detection models to identify cross-platform potential risk accounts; With accounts as nodes and transaction records as edges, a global transaction graph is constructed, and the features of nodes, edges and their connections are extracted. The fraud detection model identifies the existing fraud patterns, generates potential risk scores for risky accounts and potential fraud chains, and identifies hidden abnormal activities. Set a potential risk threshold, extract transaction information and account features with potential risk scores higher than the potential risk threshold based on the identification results of fraud patterns, and convert them into a dynamic rule base; Input real-time transaction data into the dynamic rule base for rapid matching, identify high-risk transactions and generate real-time risk scores, and set real-time risk thresholds. For transactions with real-time risk scores higher than the real-time risk thresholds, trigger risk alerts and real-time blocking mechanisms, and record transaction behaviors; Utilizing risky transaction records in the real-time matching process, new fraud behavior patterns are fed back into the global transaction graph and dynamic rule base to update the global transaction graph and dynamic rule base.
2. The method according to claim 1, characterized in that The collecting of transaction data of all platforms of users, preprocessing of the acquired data, extraction of transaction features from the processed data, establishment of a cross-platform transaction behavior database, and storage of the extracted transaction feature data specifically include: Obtain transaction data from all trading platforms, covering the user's full life cycle transaction behavior, and identify and distinguish data types from different sources, including: basic account information, transaction records, device information, and behavior logs; Map fields that describe the same attributes in different platforms to unified field names, standardize numerical features, and generate time series features based on transaction timestamps; Extract transaction behavior features and conduct cross-platform data integration, integrating transaction records from different platforms into a unified cross-platform transaction behavior database.
3. The method according to claim 2, characterized in that The establishment of a fraud detection model to identify potential risk accounts across platforms specifically includes: Build fraud detection models based on historical data and flagged fraudulent behavior; The fraud detection model is used to analyze the transaction behavior characteristics stored in the cross-platform transaction behavior database, identify and mark risk accounts with abnormal transactions based on transaction amount and frequency, device switching and geographical span.
4. The method according to claim 3, characterized in that The global transaction graph is constructed, and the features of nodes, edges and their connection relationships are extracted. The fraud detection model is used to identify the existing fraud patterns, generate the potential risk scores of risk accounts and potential fraud chains, and identify hidden abnormal activities, including: Extract data from the cross-platform transaction behavior database, use accounts as nodes, each node is accompanied by basic account information, and transaction records are used as edges. The direction of the edge indicates the direction of capital flow, and each edge is accompanied by relevant transaction features, so as to construct a global transaction graph; Calculate the static and dynamic characteristics of each node, analyze the ratio of edges to nodes in the graph, calculate the concentration of transactions, and identify sub-communities and high-frequency cyclic transaction areas in the transaction network; Analyze nodes and edges to identify abnormal behaviors, and generate potential risk scores for each node through fraud detection models based on node characteristics and associated paths; Construct a fraud chain based on the associated paths in the global transaction graph, set a potential risk threshold, define nodes with potential risk scores higher than the potential risk threshold as high-risk nodes, and read chains with multiple high-risk nodes from the fraud chain; Based on the chain of high-risk nodes and potential risk scores, accounts are grouped to obtain high-risk account groups, and the correlation between high-risk account groups and fraud chains is analyzed to mark hidden abnormal activities.
5. The method according to claim 4, characterized in that The identification result based on the fraud pattern extracts the transaction information and account features with potential risk scores higher than the potential risk threshold and converts them into a dynamic rule base, specifically including: For account transaction features with potential risk scores higher than the threshold, including transaction records, extract time features, amount features, geographic features, and behavior features, and extract abnormal transaction behavior features for high-risk accounts; Based on the abnormal transaction behavior characteristics and account transaction characteristics, dynamic rule templates are generated, including transaction amount rules, high-frequency transaction rules within a time period, geographic location jump rules and device switching rules.
6. The method according to claim 5, characterized in that The real-time transaction data is input into the dynamic rule base for rapid matching, high-risk transactions are identified and real-time risk scores are generated, and real-time risk thresholds are set. For transactions with real-time risk scores higher than the real-time risk thresholds, risk alarms and real-time blocking mechanisms are triggered, and transaction behaviors are recorded; Access transaction data from various platforms and match the accessed real-time transaction data with the rules in the dynamic rule base one by one; Generate real-time risk scores based on the number of rules matched to the transaction, rule weights, and the degree of deviation of transaction characteristics : in, represents the risk score of the current transaction when it meets rule i, W i is the weight of rule i, and n represents the number of rules in the dynamic rule base matched by the current transaction; According to business needs and historical data analysis, set a warning real-time risk threshold and a blocking real-time risk threshold respectively. The warning real-time risk threshold is smaller than the blocking real-time risk threshold. Transactions with real-time risk scores exceeding the warning real-time risk threshold are defined as high-risk transactions. For transactions that are above the warning real-time risk threshold but below the blocking real-time risk threshold, a system alarm is triggered and the transaction is marked as; For transactions that are higher than the real-time risk threshold, the transactions will be blocked and the related accounts will be frozen.
7. The data matching system in the big data environment is characterized by: The system comprises: The transaction feature extraction module is used to collect the transaction data of all platforms of users, pre-process the acquired data, extract transaction features from the processed data, establish a cross-platform transaction behavior database, and store the extracted transaction feature data; The cross-platform data integration module is used to integrate the extracted transaction feature data through the cross-platform transaction behavior database to eliminate data silos, and to match cross-platform data based on account identifiers and transaction features to establish primary account associations, and to establish fraud detection models to identify cross-platform potential risk accounts; The global transaction graph construction module is used to construct a global transaction graph with accounts as nodes and transaction records as edges, and extract the features of nodes, edges and their connection relationships. It uses the fraud detection model to identify existing fraud patterns, generate potential risk scores for risky accounts and potential fraud chains, and identify hidden abnormal activities. A dynamic rule base generation module is used to set a potential risk threshold, extract transaction information and account features with potential risk scores higher than the potential risk threshold based on the fraud pattern recognition results, and convert them into a dynamic rule base; The real-time transaction matching module is used to input the real-time transaction data into the dynamic rule base for rapid matching, identify high-risk transactions and generate real-time risk scores, and set real-time risk thresholds. For transactions with real-time risk scores higher than the real-time risk thresholds, risk alarms and real-time blocking mechanisms are triggered, and transaction behaviors are recorded; The fraud pattern feedback module is used to utilize the risk transaction records in the real-time matching process to feed back new fraud behavior patterns to the global transaction graph and dynamic rule base, and update the global transaction graph and dynamic rule base.
Citation Information
Patent Citations
Intelligent centralized monitor method and system for bank personal business fraudulent conducts
CN103714479A
Real-time monitoring, management and control method and system for reverse division
CN118657600A
Financial anti-fraud business database construction method and system
CN118916348A
Cited By
Enterprise bogus transaction detection method and system based on quantum LSTM
CN120410703A
Intelligent management system for acquiring payment data of international card
CN120746565A
Financial business risk control management system based on big data model
CN121724736A