A specific fund flow tracking method and system based on artificial intelligence
By applying artificial intelligence technology in capital flow tracking, using the isolated forest model and automatic encoder model combined with the graph database method, the problem of insufficient real-time and accuracy of capital flow in the existing technology is solved, and more efficient capital risk control and abnormal detection is achieved.
Patent Information
- Application Number
- CN202510272469.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-03-10
AI Technical Summary
When facing large-scale and cross-platform capital flow, existing capital flow tracking methods have problems such as insufficient real-time, limited accuracy and difficulty in adapting to complex scenarios. In addition, a single anomaly detection algorithm has high false alarm rates and underreport.
Using an artificial intelligence-based method, the capital transaction data is trained through the isolated forest model and the automatic encoder model, and a fund flow network diagram is built in combination with the graph database, and the capital flow path is identified and abnormal detection is performed to achieve a balance between real-time and analysis depth.
It improves the real-time and analysis depth of capital risk control, reduces the false alarm rate and missed response of abnormal detection, and enhances the ability to identify complex capital flow patterns.
Smart Images

Figure CN119782951B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and system for tracking specific fund flows based on artificial intelligence, and belongs to the technical field of fund flow tracking. Background Art
[0002] The current mainstream methods for tracking fund flows mainly rely on rule-based analysis and manual auditing. These methods are effective in dealing with simple single abnormal transactions, but when faced with large-scale, cross-platform fund flows, they have problems such as insufficient real-time performance, limited accuracy, and difficulty adapting to complex scenarios. At the same time, the diversity of data sources and the multi-dimensional characteristics of fund flow patterns increase the difficulty of analysis. Traditional technologies are unable to identify highly hidden abnormal behaviors and dynamically track fund flows, and cannot meet the needs of an increasingly complex financial environment.
[0003] In recent years, with the rapid development of artificial intelligence (AI) technology, the application of unsupervised learning algorithms in anomaly detection and behavior pattern analysis has become increasingly mature. However, the application of existing artificial intelligence technology in financial funds tracking platforms still has many shortcomings. First, there is a lack of unified standards in the data preprocessing stage, which makes it difficult to ensure high-quality data input, affecting the accuracy of subsequent analysis. Secondly, when faced with complex and diverse abnormal behaviors, a single anomaly detection algorithm has a high false alarm rate and missed reports. At the same time, in the process of tracking the flow of funds, how to strike a balance between real-time performance and analysis depth is still a difficult point in technical implementation. Summary of the invention
[0004] In order to overcome the above problems, the present disclosure provides a method and system for tracking specific fund flows based on artificial intelligence.
[0005] In a first aspect, the present disclosure provides a method for tracking specific fund flows based on artificial intelligence, comprising the following steps:
[0006] Collecting fund transaction data from multiple platforms, cleaning and preprocessing the fund transaction data, and storing the preprocessed fund transaction data in a structured database;
[0007] Using the fund transaction data, a transaction record sample is created and marked as a normal transaction sample or an abnormal transaction sample. Using the transaction record sample, an isolation forest model and an autoencoder model are trained to obtain a trained isolation forest model and autoencoder model. The isolation forest model and the autoencoder model are combined with a preset rule base to obtain an abnormal fund transaction data detection model. The fund transaction data to be detected is detected by using the abnormal fund transaction data detection model, and an early warning is issued for the abnormal fund transaction data.
[0008] Obtain the risk accounts involved in the abnormal fund transaction data, obtain the accounts and fund transaction data associated with the risk accounts, use the abnormal fund transaction data and risk accounts as the data, build a fund flow network diagram through the graph database, abnormal fund transaction data and risk accounts, and obtain the fund flow path through the fund flow network diagram.
[0009] Furthermore, the transaction record sample includes statistical characteristics, behavioral characteristics and correlation characteristics. The statistical characteristics include transaction amount and transaction frequency. The behavioral characteristics include transaction time, business type and transaction location. The correlation characteristics include the correlation relationship between the two transaction parties and the degree of correlation between the two transaction parties and the historical abnormal transaction subjects.
[0010] Furthermore, when the isolation forest model and the autoencoder model are trained respectively by the transaction record samples, cross-validation is used to evaluate the model performance and to tune the hyperparameters.
[0011] Furthermore, the isolation forest model is an improved isolation forest model, specifically:
[0012] Dynamically select the features of each node in the isolation forest model, specifically:
[0013] Calculate features F Feature score F score :
[0014]
[0015] in, α is a weight parameter used to balance the influence of information gain and variance. IG ( D , F ) is the information gain, V F Features F The variance of H ( D ) is the data set D The entropy of For the i subsets, N is the number of subsets;
[0016] The feature with the highest score m features as candidate features, randomly selected k candidate features as the features of the node, m>k ;
[0017] Randomly obtain the segmentation value of each feature of the node, and take the weighted sum of the segmentation values of each feature as the comprehensive segmentation value of the node. The weight of each feature segmentation value satisfies:
[0018]
[0019] in, F i For the i The feature score of each feature, β i For the i The weight of each feature;
[0020] Obtaining a comprehensive feature score of the node, wherein the comprehensive feature score is obtained by weighted summation of the features of the node, and the weight of each feature is the same as the weight of each feature segmentation value;
[0021] The node divides the left subtree and the right subtree according to the comprehensive feature score and the comprehensive segmentation value.
[0022] Furthermore, the path length of the isolation forest model h ( x ) is calculated as follows:
[0023]
[0024] in, K For sample x The length of the path, x i For sample x No. i The weight of the node, c ( n ) is the correction value, F j For sample x No. i On the node j The value of the feature, is the feature in all samples F j The average value of n is the size of the dataset, H ( n- 1) is a harmonic number.
[0025] Furthermore, an anomaly score is constructed based on the accuracy of the selected features in the path, specifically:
[0026] sample x Anomaly score s ( x , n )as follows:
[0027]
[0028] in, E ( h (x )) is the path processing function, which processes the path length as follows:
[0029] path T For sample x A path in isolated forest, get path T Medium Features F j The number of selections and weights of each selection are summed to obtain the feature F j The comprehensive weight of F Tj :
[0030]
[0031] in, F ji Features F j On the path T Middle i The weight when it is selected, u Features F j On the path T The number of times selected;
[0032] Structural features F j Relative to the sample x The comprehensive weight matrix E j :
[0033]
[0034] in, M For sample x The number of paths, F 1j Features F j In the sample x The comprehensive weight on the first path;
[0035] Constructing samples x The path length matrix h x :
[0036]
[0037] Calculate the comprehensive weight matrix E j and the path length matrix h x Correlation coefficient r j ;
[0038] Eliminate the correlation coefficient r j For features smaller than the preset value, the path length is calculated using the remaining features. h ’ ( x ):
[0039]
[0040] in, K ’ is the number of remaining features;
[0041] E ( h ( x ))= h ’ v ( x );
[0042] in, h ’ v ( x ) is a sample x All path lengths h ’ ( x )’s average value.
[0043] Furthermore, the rule base includes rules based on expert experience and rules based on statistical analysis.
[0044] Furthermore, a fund flow network diagram is constructed through the graph database, abnormal fund transaction data and risk accounts, and the fund flow path is obtained through the fund flow network diagram, including:
[0045] Construct a fund flow network diagram, define transaction accounts and transaction entities as node types, transaction account nodes include account ID, account type, account platform and transaction frequency, transaction entity nodes include merchant ID and wallet address, nodes are connected by directed edges, the direction of the edge is the direction of fund flow, and the attributes of the edge include transaction amount, transaction time, transaction channel and transaction location;
[0046] A breadth-first search algorithm and a bidirectional search algorithm are used to identify the capital flow path in the capital flow network diagram.
[0047] Furthermore, when obtaining the capital flow path through the capital flow network diagram, the A* algorithm is used to optimize the path;
[0048] For abnormal fund transaction data that may be money laundering, the Tarjan strongly connected component algorithm is used to identify loops in the fund flow network diagram to determine whether there is money laundering in the loop;
[0049] The Louvain algorithm and the label propagation algorithm are used to cluster the transaction accounts of the fund flow network diagram to identify illegal fund flow groups;
[0050] For abnormal fund transaction data involving blockchain, heuristic rules are used to cluster blockchain addresses, and multiple associated addresses are identified and processed as the same entity.
[0051] In a second aspect, the present disclosure further provides a specific fund flow tracking system based on artificial intelligence, including a data acquisition unit, a data storage unit, a data processing unit and a data output unit:
[0052] The data collection unit is used to collect fund transaction data of multiple platforms;
[0053] The data storage unit is used to store the data collected by the data collection unit and the data processed by the data processing unit;
[0054] The data processing unit processes the fund transaction data by using the specific fund flow tracking method based on artificial intelligence described in the first aspect;
[0055] The data output unit is used to output the fund transaction data or the data processed by the data processing unit.
[0056] The present disclosure has the following beneficial effects:
[0057] The present invention solves the problem of tracking specific fund flows in steps. First, the risk warning of abnormal fund transaction data is identified through the isolation forest model and the autoencoder model. Then, the fund flow is identified through the fund flow network diagram, ensuring the real-time and analysis depth of fund risk control.
[0058] The present invention improves the isolation forest model and introduces a dynamic feature selection mechanism. At the same time, by acting on multiple features together during a segmentation, compared with the random feature selection of the existing isolation forest model, the dynamic feature selection mechanism can avoid irrelevant features from reducing the accuracy of anomaly detection. When calculating the path length, a weighted path length is proposed, taking into account the different importance of nodes at different depths, making the path length more meaningful in the calculation of the anomaly score. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 A flow chart of the method disclosed in the present invention. DETAILED DESCRIPTION
[0060] In order to make the purpose, technical solution and advantages of the embodiments of the present disclosure clearer, the technical solution of the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings of the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the described embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.
[0061] Unless otherwise defined, the technical terms or scientific terms used in the present disclosure should be understood by people with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar words used in the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing in front of the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly. In order to keep the following description of the embodiments of the present disclosure clear and concise, the present disclosure omits the detailed description of some known functions and known components.
[0062] The present disclosure is described in detail below with reference to the accompanying drawings and specific embodiments.
[0063] First, refer to Figure 1 The present disclosure provides a method for tracking specific fund flows based on artificial intelligence, comprising the following steps:
[0064] Collecting fund transaction data from multiple platforms, cleaning and preprocessing the fund transaction data, and storing the preprocessed fund transaction data in a structured database;
[0065] Using the fund transaction data, a transaction record sample is created and marked as a normal transaction sample or an abnormal transaction sample. Using the transaction record sample, an isolation forest model and an autoencoder model are trained to obtain a trained isolation forest model and autoencoder model. The isolation forest model and the autoencoder model are combined with a preset rule base to obtain an abnormal fund transaction data detection model. The fund transaction data to be detected is detected by using the abnormal fund transaction data detection model, and an early warning is issued for the abnormal fund transaction data.
[0066] Obtain the risk accounts involved in the abnormal fund transaction data, obtain the accounts and fund transaction data associated with the risk accounts, use the abnormal fund transaction data and risk accounts as the data, build a fund flow network diagram through the graph database, abnormal fund transaction data and risk accounts, and obtain the fund flow path through the fund flow network diagram.
[0067] This embodiment handles the problem of tracking specific fund flows in steps. First, it uses the isolation forest model and the autoencoder model to identify risk warnings for abnormal fund transaction data, and then identifies fund flows through a fund flow network diagram, ensuring the real-time nature of fund risk control and the depth of analysis.
[0068] In one embodiment of the present disclosure, the transaction record sample includes statistical characteristics, behavioral characteristics and association characteristics. The statistical characteristics include transaction amount and transaction frequency. The behavioral characteristics include transaction time, business type and transaction location. The association characteristics include the association relationship between the two transaction parties and the degree of association between the two transaction parties and historical abnormal transaction entities.
[0069] In one embodiment of the present disclosure, when the isolation forest model and the autoencoder model are trained respectively by the transaction record samples, cross-validation is used to evaluate the model performance and to tune the hyperparameters.
[0070] In one embodiment of the present disclosure, the isolation forest model is an improved isolation forest model, specifically:
[0071] Dynamically select the features of each node in the isolation forest model, specifically:
[0072] Calculate features F Feature score F score :
[0073]
[0074] in, α is a weight parameter used to balance the influence of information gain and variance. IG ( D , F ) is the information gain, V F Features F The variance of H ( D ) is the data set D The entropy of For the i subsets, N is the number of subsets;
[0075] The feature with the highest score mfeatures as candidate features, randomly selected k candidate features as the features of the node, m>k ;
[0076] Randomly obtain the segmentation value of each feature of the node, and take the weighted sum of the segmentation values of each feature as the comprehensive segmentation value of the node. The weight of each feature segmentation value satisfies:
[0077]
[0078] in, F i For the i The feature score of each feature, β i For the i The weight of each feature;
[0079] Obtaining a comprehensive feature score of the node, wherein the comprehensive feature score is obtained by weighted summation of the features of the node, and the weight of each feature is the same as the weight of each feature segmentation value;
[0080] The node divides the left subtree and the right subtree according to the comprehensive feature score and the comprehensive segmentation value.
[0081] This embodiment introduces a dynamic feature selection mechanism into the isolation forest model. At the same time, by acting on multiple features together during one segmentation, compared with the random feature selection of the existing isolation forest model, the dynamic feature selection mechanism can avoid irrelevant features from reducing the accuracy of anomaly detection.
[0082] In one embodiment of the present disclosure, the path length of the isolation forest model is h ( x ) is calculated as follows:
[0083]
[0084] in, K For sample x The length of the path, x i For sample x No. i The weight of the node, c ( n ) is the correction value, F j For sample x No. i On the node j The value of the feature, is the feature in all samples F j The average value of n is the size of the dataset,H ( n- 1) is a harmonic number.
[0085] This embodiment proposes a weighted path length when calculating the path length, taking into account the different importance of nodes at different depths, so that the path length is more meaningful in the calculation of the anomaly score.
[0086] In one embodiment of the present disclosure, an anomaly score is constructed by the accuracy of the selected features in the path, specifically:
[0087] sample x Anomaly score s ( x , n )as follows:
[0088]
[0089] in, E ( h ( x )) is the path processing function, which processes the path length as follows:
[0090] path T For sample x A path in isolated forest, get path T Medium Features F j The number of selections and weights of each selection are summed up to get the feature F j The comprehensive weight of F Tj :
[0091]
[0092] in, F ji Features F j On the path T Middle i The weight when it is selected, u Features F j On the path T The number of times selected;
[0093] Structural features F j Relative to the sample x The comprehensive weight matrix E j :
[0094]
[0095] in, M For sample x The number of paths, F 1j Features F j In the sample x The comprehensive weight on the first path;
[0096] Constructing samples x The path length matrix h x :
[0097]
[0098] Calculate the comprehensive weight matrix E j and the path length matrix h x Correlation coefficient r j ;
[0099] Eliminate the correlation coefficient r j For features smaller than the preset value, the path length is calculated using the remaining features. h ’ ( x ):
[0100]
[0101] in, K ’ is the number of remaining features;
[0102] E ( h ( x ))= h ’ v ( x );
[0103] in, h ’ v ( x ) is a sample x All path lengths h ’ ( x )’s average value.
[0104] In one embodiment of the present disclosure, the rule base includes rules based on expert experience and rules based on statistical analysis.
[0105] This embodiment takes into account that some features have little influence on sample segmentation when calculating the anomaly score of the sample, and removes such features based on the relationship between the features and the path length, thereby improving the accuracy of the anomaly score.
[0106] In one embodiment of the present disclosure, a fund flow network diagram is constructed through a graph database, abnormal fund transaction data and risk accounts, and a fund flow path is obtained through the fund flow network diagram, including:
[0107] Construct a fund flow network diagram, define transaction accounts and transaction entities as node types, transaction account nodes include account ID, account type, account platform and transaction frequency, transaction entity nodes include merchant ID and wallet address, nodes are connected by directed edges, the direction of the edge is the direction of fund flow, and the attributes of the edge include transaction amount, transaction time, transaction channel and transaction location;
[0108] A breadth-first search algorithm and a bidirectional search algorithm are used to identify the capital flow path in the capital flow network diagram.
[0109] In one embodiment of the present disclosure, when obtaining the capital flow path through the capital flow network diagram, the A* algorithm is used to optimize the path; by setting the transaction amount and timestamp as the weight of the path, high-frequency and large-amount capital flow paths are preferentially identified. For example, when tracking cross-border capital transfers, the system will give priority to transaction paths with larger amounts, thereby ensuring the accuracy and importance of the capital flow path.
[0110] For abnormal fund transaction data that may be money laundering, the Tarjan strongly connected component algorithm is used to identify the loop in the fund flow network diagram to determine whether there is money laundering in the loop. The Tarjan algorithm can help quickly detect this loop and mark it as a suspicious transaction.
[0111] The Louvain algorithm and label propagation algorithm are used to cluster the transaction accounts of the fund flow network graph to identify illegal fund flow groups. In specific scenarios, the Louvain algorithm will find high-density communities in the graph by optimizing modularity, while the LPA algorithm identifies the relationship between nodes based on the label propagation mechanism. The system analyzes the transaction behavior of each account and uses these algorithms to identify potential illegal activities. For example, when multiple accounts trade frequently and there is a strong connection between these accounts, these accounts may be clustered into a community, thereby identifying a potential money laundering gang.
[0112] For abnormal fund transaction data involving blockchain, heuristic rules are used to cluster blockchain addresses, and multiple associated addresses are identified and processed as the same entity.
[0113] In a second aspect, the present disclosure further provides a specific fund flow tracking system based on artificial intelligence, including a data acquisition unit, a data storage unit, a data processing unit and a data output unit:
[0114] The data collection unit is used to collect fund transaction data of multiple platforms;
[0115] The data storage unit is used to store the data collected by the data collection unit and the data processed by the data processing unit;
[0116] The data processing unit processes the fund transaction data by using the specific fund flow tracking method based on artificial intelligence described in the first aspect;
[0117] The data output unit is used to output the fund transaction data or the data processed by the data processing unit.
[0118] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0119] The units involved in the embodiments described in the present disclosure may be implemented by software or hardware, wherein the name of a unit does not, in some cases, limit the unit itself.
[0120] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0121] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present disclosure (but not limited to) by each other to form a technical solution.
[0122] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0123] Although the subject matter has been described in language specific to structural features and / or methodological logical actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. On the contrary, the specific features and actions described above are merely example forms of implementing the claims.
[0124] There are a few points to note about this disclosure:
[0125] (1) The drawings of the embodiments of the present disclosure only relate to the structures related to the embodiments of the present disclosure. Other structures may refer to the general design.
[0126] (2) In the absence of conflict, the embodiments of the present disclosure and the features therein may be combined with each other to obtain new embodiments.
[0127] The above descriptions are merely embodiments of the present disclosure and are not intended to limit the patent scope of the present disclosure. Any equivalent structures made using the contents of the present disclosure and the drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present disclosure.
Claims
1. A specific fund flow tracking method based on artificial intelligence, characterized in that: The following steps are involved: Collecting fund transaction data from multiple platforms, cleaning and preprocessing the fund transaction data, and storing the preprocessed fund transaction data in a structured database; Using the fund transaction data, a transaction record sample is created and marked as a normal transaction sample or an abnormal transaction sample. Using the transaction record sample, an isolation forest model and an autoencoder model are trained to obtain a trained isolation forest model and autoencoder model. The isolation forest model and the autoencoder model are combined with a preset rule base to obtain an abnormal fund transaction data detection model. The fund transaction data to be detected is detected by using the abnormal fund transaction data detection model, and an early warning is issued for the abnormal fund transaction data. Obtain risk accounts involved in abnormal fund transaction data, obtain accounts and fund transaction data associated with risk accounts, use abnormal fund transaction data and risk accounts as data, construct a fund flow network graph through the graph database, abnormal fund transaction data and risk accounts, and obtain fund flow paths through the fund flow network graph; The isolation forest model is an improved isolation forest model, specifically: Dynamically select the features of each node in the isolation forest model, specifically: Calculate features F Feature score F score : in, α is a weight parameter used to balance the influence of information gain and variance. IG ( D , F ) is the information gain, V F Features F The variance of H ( D ) is the data set D The entropy of For the i subsets, N is the number of subsets; The feature with the highest score m features as candidate features, randomly selected k candidate features as the features of the node, m >k ; Randomly obtain the segmentation value of each feature of the node, and take the weighted sum of the segmentation values of each feature as the comprehensive segmentation value of the node. The weight of each feature segmentation value satisfies: in, F i For the i The feature score of each feature, β i For the i The weight of each feature; Obtaining a comprehensive feature score of the node, wherein the comprehensive feature score is obtained by weighted summation of the features of the node, and the weight of each feature is the same as the weight of each feature segmentation value; The node divides the left subtree and the right subtree according to the comprehensive feature score and the comprehensive segmentation value.
2. The specific fund flow tracking method based on artificial intelligence according to claim 1 is characterized in that: The transaction record sample includes statistical features, behavioral features and correlation features. The statistical features include transaction amount and transaction frequency. The behavioral features include transaction time, business type and transaction location. The correlation features include the correlation relationship between the two transaction parties and the degree of correlation between the two transaction parties and the historical abnormal transaction subjects.
3. The specific funds flow tracking method based on artificial intelligence according to claim 2 is characterized in that: When the isolation forest model and the autoencoder model are trained respectively by using the transaction record samples, cross-validation is used to evaluate the model performance and to tune the hyperparameters.
4. The specific fund flow tracking method based on artificial intelligence according to claim 3 is characterized in that: The path length of the isolation forest model h ( x ) is calculated as follows: in, K For sample x The length of the path, x i For sample x No. i The weight of the node, c ( n ) is the correction value, F j For sample x No. i On the node j The value of the feature, is the feature in all samples F j The average value of n is the size of the dataset, H ( n- 1) is a harmonic number.
5. The specific fund flow tracking method based on artificial intelligence according to claim 4 is characterized in that: The anomaly score is constructed by the accuracy of the features selected in the path, specifically: sample x Anomaly score s ( x , n )as follows: in, E ( h ( x )) is the path processing function, which processes the path length as follows: path T For sample x A path in isolated forest, get path T Medium Features F j The number of selections and weights of each selection are summed up to get the feature F j The comprehensive weight of F Tj : in, F ji Features F j On the path T Middle i The weight when it is selected, u Features F j On the path T The number of times selected; Structural features F j Relative to the sample x The comprehensive weight matrix E j : in, M For sample x The number of paths, F 1j Features F j In the sample x The comprehensive weight on the first path; Constructing samples x The path length matrix h x : Calculate the comprehensive weight matrix E j and the path length matrix h x Correlation coefficient r j ; Eliminate the correlation coefficient r j For features smaller than the preset value, the path length is calculated using the remaining features. h ’ ( x ): in, K ’ is the number of remaining features; E ( h ( x ))= h ’ v ( x ); in, h ’ v ( x ) is a sample x All path lengths h ’ ( x )’s average value.
6. The specific fund flow tracking method based on artificial intelligence according to claim 3 is characterized in that: The rule base includes rules based on expert experience and rules based on statistical analysis.
7. The specific fund flow tracking method based on artificial intelligence according to claim 3 is characterized in that: The fund flow network diagram is constructed through the graph database, abnormal fund transaction data and risk accounts, and the fund flow path is obtained through the fund flow network diagram, including: Construct a fund flow network diagram, define transaction accounts and transaction entities as node types, transaction account nodes include account ID, account type, account platform and transaction frequency, transaction entity nodes include merchant ID and wallet address, nodes are connected by directed edges, the direction of the edge is the direction of fund flow, and the attributes of the edge include transaction amount, transaction time, transaction channel and transaction location; A breadth-first search algorithm and a bidirectional search algorithm are used to identify the capital flow path in the capital flow network diagram.
8. The specific fund flow tracking method based on artificial intelligence according to claim 3 is characterized in that: When obtaining the capital flow path through the capital flow network diagram, the A* algorithm is used to optimize the path; For abnormal fund transaction data that may be money laundering, the Tarjan strongly connected component algorithm is used to identify loops in the fund flow network diagram to determine whether there is money laundering in the loop; The Louvain algorithm and the label propagation algorithm are used to cluster the transaction accounts of the fund flow network diagram to identify illegal fund flow groups; For abnormal fund transaction data involving blockchain, heuristic rules are used to cluster blockchain addresses, and multiple associated addresses are identified and processed as the same entity.
9. A specific fund flow tracking system based on artificial intelligence, characterized in that: It includes data acquisition unit, data storage unit, data processing unit and data output unit: The data collection unit is used to collect fund transaction data of multiple platforms; The data storage unit is used to store the data collected by the data collection unit and the data processed by the data processing unit; The data processing unit processes the fund transaction data by using the specific fund flow tracking method based on artificial intelligence according to any one of claims 1 to 8; The data output unit is used to output the fund transaction data or the data processed by the data processing unit.
Citation Information
Patent Citations
A method for detecting abnormal behavior of electricity consumption of consumers based on isolated forests
CN109308306A
Consumption credit fraud behavior detection method and system based on isolated forest
CN111833172A