Financial fraud detection method based on similar scene learning and pseudo label data generation
Through similar scenario learning and pseudo-label data generation methods, the detection accuracy and stability of the financial fraud detection model in various financial scenarios are improved, and the problems of low detection accuracy and poor generalization caused by insufficient data volume are solved, and efficient fraud detection is achieved.
Patent Information
- Application Number
- CN202510341619.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-18
AI Technical Summary
Due to insufficient data volume, the existing financial fraud detection methods have low detection accuracy and cannot be applied to multiple financial scenarios, which limits their reliability and generalization in different scenarios.
Using a method based on similar scenario learning and pseudo-label data generation, the GCN graph fraud detection model is used to learn prior knowledge in the source scenario, generate pseudo-label data, and train the model with step-by-step hybrid strategy to improve its fraud detection capabilities in the target scenario.
High-precision fraud detection in various financial scenarios is realized, the execution efficiency and stability of the detection is improved, and fraud behaviors in mastered and uncovered financial scenarios can be identified.
Smart Images

Figure CN120338813A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer big data technology, and specifically relates to a financial fraud detection method based on similar scenario learning and pseudo-label data generation.
Background Art
[0002] Fraud detection is one of the core technologies in the field of financial user behavior monitoring. However, due to problems such as personal privacy and conflicts of interest in financial scenarios, the high data sensitivity results in the inability to obtain a large amount of data. The existing fraud detection methods, due to the limitation of low data volume, directly lead to low detection accuracy and can only be applied to preset financial scenarios and cannot be applied to other financial scenarios. This limitation seriously increases the learning and design costs of the fraud detection model in practical applications, reduces its reliability and generalization ability in multiple scenarios, and thus limits the wide application of fraud detection.
[0003] GCN (Graph Convolutional Network) is a deep learning model for processing graph-structured data.
[0004] The Wasserstein distance, which is a way to measure the difference between two probability distributions and belongs to part of the optimal transport theory.
[0005] The present invention makes technical improvements to the financial fraud detection method for the technical problems of low detection accuracy and poor generalization caused by insufficient user category annotation in the field of financial fraud detection.
Summary of the Invention
[0006] The object of the present invention is to provide a financial fraud detection method that meets the high-precision fraud detection requirements of different financial scenarios, has high execution efficiency and good stability.
[0007] To achieve the above object, the technical solution adopted by the present invention is a financial fraud detection method based on similar scenario learning and pseudo-label data generation, including the following steps:
[0008] S1. First, use the graph fraud detection model based on GCN to learn the prior knowledge required for fraud detection in the source scenario, that is, the scenario similar to the target scenario, and evaluate the detection ability of the model, so that the graph fraud detection model has the general feature representation ability of fraudulent users;
[0009] S2. Then, improve the diversity of the prior knowledge of the source scenario by generating the connected subgraphs of the source scenario, and calculate the average feature representations of fraudulent and normal users in each source sub-scenario;
[0010] S3. Subsequently, generate pseudo-label data according to the Wasserstein distance in the target scenario and the source sub-scenarios;
[0011] S4. Next, further train the graph fraud detection model by using the generated pseudo-label data and fraud detection with a progressive mixing strategy, so that the graph fraud detection model has the ability to represent unique features of fraudulent users in the target scenario;
[0012] S5. Finally, deploy the trained graph fraud detection model to the actual application scenario to accurately detect fraudulent behaviors of users in the given user transaction network.
[0013] Preferably, for a financial fraud detection method based on similar scenario learning and pseudo-label data generation, the task definition of financial fraud detection is: given a user transaction behavior network where represents individual users, ε represents the transaction behaviors between users, represents the feature representation of the behaviors of all n individual users, and the goal is to detect whether there are abnormal transaction behaviors for each user; the financial fraud detection result is represented by where y i = 1 indicates the existence of abnormal behaviors, and y i = 0 indicates the non-existence of abnormal behaviors.
[0014] Preferably, for a financial fraud detection method based on similar scenario learning and pseudo-label data generation, step S1 specifically includes the following sub-steps:
[0015] S11. Use the data in the source scenario to initially train a graph fraud detection model of GCN, so that the graph fraud detection model of GCN has the ability to detect abnormal behaviors, and use cross-entropy loss as its loss function. The calculation formula is as follows:
[0016]
[0017] where y i is the model prediction result, is the actual situation;
[0018] S12. Evaluate the ability and reliability of the graph fraud detection model for the abnormal behavior detection task through the performance of the graph fraud detection model of GCN on the test data. The evaluation index is the proportion of the correctly detected abnormal TP and normal TN and the wrongly detected abnormal FP and normal FN in the source scenario by the graph fraud detection model in the total test samples, which are respectively represented as The specific calculation formula is as follows:
[0019]
[0020] where, represents the actual user detection result, Represents the actual detection result, Represents taking the inverse of the value in (·), n s Represents the number of user individuals in the source scenario test data.
[0021] Preferably, for the financial fraud detection method based on similar scenario learning and pseudo-label data generation, step S2 specifically includes the following sub-steps:
[0022] S21. To enhance the diversity of the learned feature representations, for the user transaction behavior network in the source scenario Generate several connected subgraphs The specific generation sub-steps of each connected subgraph,
[0023] First, start from a randomly selected user, query other users who have transaction records with this user, and add them to the user individual set of the subgraph
[0024] Then, repeat this process, gradually expand the subgraph until the number of users in the subgraph exceeds 60% of the total number of users in the source scenario. During this process, add the transaction behaviors between all these users to the edge set And add the behavior feature representations of these users to the feature set of the subgraph In particular, denote the set of normal user features as Denote the set of user features with fraud behaviors as In this way, generate multiple connected subgraphs, thereby increasing the diversity of feature representations and helping the graph fraud detection model better capture the complexity and heterogeneity of user behaviors;
[0025] S22. For each generated connected subgraph, calculate the means of the feature representations of users with abnormal behaviors and users without abnormal behaviors respectively, denoted as and And use these means as prior knowledge for the alignment of user behavior feature representations in the subsequent target scenario,
[0026]
[0027] Preferably, for the financial fraud detection method based on similar scenario learning and pseudo-label data generation, step S3 specifically includes the following sub-steps:
[0028] S31. Calculate the ratio of positive and negative pseudo-labels that should be generated in the target scenario through the performance evaluation index of the graph fraud detection model in the source scenario obtained in step S1. The calculation formula is as follows:
[0029]
[0030] Among them, is the fraud detection result of the model in the current round of training in the target scenario;
[0031] S32. To increase the diversity of pseudo-label samples and reduce the impact of the instability of fraud detection results on pseudo-label generation, a small-range random fluctuation is added to the calculated ratio. The specific calculation method is as follows:
[0032]
[0033] Among them, c is the small-range random fluctuation added, and the effective range is [0, 0.05];
[0034] S33. To ensure the quality of pseudo-labels, a confidence condition parameter α is introduced to calculate the proportion of positive and negative pseudo-labels generated in the total number of samples. The calculation formula is as follows:
[0035]
[0036] Among them, GCN(·) is the graph fraud detection model;
[0037] S34. Determine the number of positive and negative pseudo-labels to be generated respectively according to the ratio between positive and negative samples after adding fluctuations and their respective proportions in the total number of samples and The calculation formula is as follows:
[0038]
[0039] S35. Calculate the mean values of user feature representations in the target scenario and the source scenario and of the Wasserstein distance. The calculation formula is as follows:
[0040]
[0041] Among them, is x t and all possible combinations;
[0042] Sort the features of user individuals in the target scenario according to the calculated Wasserstein distance, and add positive pseudo-labels to the top features with the smallest distance to , and add negative pseudo-labels to the top features with the smallest distance to ; Considering the abnormal user representations unique to the target scenario that the graph fraud detection model fails to recognize, those related to and Add positive pseudo-labels to the top c user features with the smallest sum of distances.
[0043] Preferably, in the financial fraud detection method based on similar scenario learning and pseudo-label data generation, step S4 specifically includes the following sub-steps:
[0044] By linearly adding the abnormal and normal user individual features in the target scenario to those in the source scenario, gradually increasing the proportion of the target scenario features to further enhance the representation ability of the graph fraud detection model in the target scenario. The specific formula is as follows:
[0045]
[0046] Where, and are the user individual features with positive and negative pseudo-labels added in the target scenario; and are the mixing parameters for the e-th training round, and the calculation formula is as follows:
[0047]
[0048] Where, e is the current training round, E is the total number of training rounds, q P and q N are the weight factors for the user individual feature samples labeled with positive and negative pseudo-labels, and the calculation formula is as follows:
[0049]
[0050] Where, exp(x) represents e to the power of x.
[0051] The financial fraud detection method based on similar scenario learning and pseudo-label data generation of the present invention has the following beneficial effects: To solve the limitations of the prior art, especially for the detection requirements of fraudulent users in multiple similar scenarios, and to solve the problems of low detection accuracy and poor generalization caused by insufficient user category annotation in the financial field, the present invention proposes a brand-new financial fraud detection training framework based on similar scenario learning and pseudo-label data generation. By learning the existing fraud detection capabilities in similar financial scenarios and applying the prior knowledge therein to the target fraud detection scenario, high-quality pseudo-label data is generated through traditional empirical statistical models, thereby enhancing the fraud detection ability in the target scenario and increasing the coverage of detectable financial scenarios. Through the continuous transfer of prior knowledge between scenarios and the continuously generated pseudo-label data, the trained detection model can effectively handle the user fraud behavior detection tasks in multiple financial scenarios. The method of the present invention can learn or enhance the fraud detection ability in similar financial scenarios through financial scenarios with certain detection capabilities, and meet the high-precision fraud detection requirements in different financial scenarios. It can not only identify the mastered scenarios, but also effectively detect the originally uncovered financial scenarios, thereby effectively enhancing the intelligent detection and adaptation ability of the fraud behavior detection model in multiple financial scenarios. The method of the present invention provides intelligent support for tasks such as identifying or warning user fraud behaviors and tracking illegal financial criminals in the financial field, and effectively improves the execution efficiency and stability of fraud detection in the financial field.
Description of the Drawings
[0052] Figure 1 It is a flowchart of a financial fraud detection method framework based on similar scenario learning and pseudo-label data generation.
Specific Embodiments
[0053] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0054] Embodiment
[0055] This embodiment implements a financial fraud detection method based on similar scenario learning and pseudo-label data generation.
[0056] This embodiment proposes a financial fraud detection method based on similar scenario learning and pseudo-label data generation, aiming to solve the problems of low detection accuracy and poor generalization caused by insufficient user category annotation in the financial field. The method of this embodiment can learn or improve the fraud detection ability in similar financial scenarios through financial scenarios with certain detection capabilities, and meet the high-precision fraud detection requirements in different financial scenarios. The method of this embodiment provides intelligent support for tasks such as identifying or warning user fraud behaviors and tracking illegal financial criminals in the financial field, effectively improving the execution efficiency and stability of fraud detection in the financial field.
[0057] Figure 1 It is a framework flowchart of a financial fraud detection method based on similar scenario learning and pseudo-label data generation. As Figure 1 shown, the method of this embodiment: First, use the graph detection network based on GCN to learn the prior knowledge required for fraud detection in the target scenario and evaluate the detection ability of the model, so that it has the general feature representation ability of fraud users; then improve the diversity of the prior knowledge of the source scenario by generating the connected subgraph of the source scenario, and calculate the average feature representation of fraud and normal users in each source sub-scenario; subsequently, generate pseudo-label data according to the Wasserstein distance between the target scenario and the source sub-scenario; then, further train the fraud detection model through the generated pseudo-label data and using the step-by-step mixing strategy for fraud detection, so that it has the unique feature representation ability of fraud users in the target scenario; finally, deploy the trained fraud detection model to the actual application scenario to accurately detect the fraud behaviors of users in the given user transaction network. The method of this embodiment significantly improves the detection accuracy in the target scenario with low data volume, and can be extended to other new target financial scenarios, enhancing the accuracy and scenario generalization ability of fraud detection in multiple financial scenarios, and providing more reliable technical support for tasks such as identifying or warning user fraud behaviors and tracking illegal financial criminals in the financial field.
[0058] Specifically, a financial fraud detection method based on similar scenario learning and pseudo-label data generation in this embodiment is as follows:
[0059] 1. A financial fraud detection method based on similar scenario learning and pseudo-label data generation, which combines the prior knowledge of transfer learning and the data mining ability of deep learning models, aims to improve the detection accuracy and generalization of user fraud behavior in the financial field. Its overall process is as follows: First, considering the similarity of application scenarios in the financial field, a graph detection network based on GCN is used to preliminarily learn the general feature representation of user behavior in the source scenario (i.e., the scenario similar to the target scenario), and detect whether there is fraud behavior. Through the existing training data, the model's recognition ability for normal and abnormal user behaviors in the source scenario is evaluated, and then the sub-scenarios of the source scenario are generated and the mean value of their user characteristics is statistically calculated to provide prior knowledge for the target scenario. Subsequently, the model trained in the source scenario and the prior knowledge are applied to the target scenario (i.e., the actual scenario where fraud detection is required). According to the detection ability evaluated in the source scenario and the preliminary detection results in the target scenario, the ratio of normal and abnormal pseudo-labels generated in the target scenario is determined, and high-quality pseudo-labels are provided for the target scenario data through a pseudo-label generation method that combines this ratio constraint and confidence conditions. Finally, the model's learning ability for the unique characteristics of user behavior in the target scenario is enhanced through a step-by-step mixing strategy. This process helps the model better understand the user behavior pattern in the target scenario, thereby improving the accuracy of fraud detection. Through the learning of similar scenario knowledge, the transfer of prior knowledge, and the generation of pseudo-label data, this method significantly improves the understanding of user behavior patterns and the detection ability of abnormal behaviors in the financial field, and demonstrates high generalization and precision. This provides strong support for financial fraud early warning and illegal criminal activity tracking in practical applications.
[0060] The task definition of financial fraud detection is: Given a user transaction behavior network where represents individual users, ε represents the transaction behavior between users, represents the feature representation of the behavior of all n individual users, and the goal is to detect whether each user has abnormal transaction behavior. The detection result is represented by where y i = 1 indicates the existence of abnormal behavior, and y i = 0 indicates the non-existence of abnormal behavior.
[0061] 2. The method of this embodiment is based on a graph detection network of GCN. First, by learning the feature representation method in the source scenario, shared features that can be applied to the target scenario are obtained. Through the GCN network, the model can extract the feature information commonly existing in similar scenarios, thereby deeply learning the general representation of user behavior patterns and providing complex and diverse prior knowledge for the target scenario. Specifically, it includes the following steps:
[0062] Step A1: Initially train a GCN network using the data in the source scenario to endow it with the ability to detect abnormal behaviors. Use cross-entropy loss as its loss function, and the calculation formula is as follows:
[0063]
[0064] where, y i is the model prediction result, is the real situation.
[0065] Step A2: Evaluate the ability and reliability of this model for the abnormal behavior detection task based on its performance on the test data. The evaluation metrics are the proportions of correctly detecting abnormal (TP) and normal (TN) in the source scenario and misdetecting abnormal (FP) and normal (FN) in the total test samples by the detection network, which are respectively expressed as The specific calculation formulas are as follows:
[0066]
[0067] where, represents the real user detection result, represents the actual detection result, represents taking the inverse of the value in (·), and n s represents the number of user individuals in the source scenario test data. Through this metric, we can understand the detection accuracy, recall rate, and false detection rate of the initially trained detection model.
[0068] Step B: To enhance the diversity of the learned feature representations, for the user transaction behavior network in the source scenario generate several connected subgraphs The specific generation method for each connected subgraph is as follows: First, start from a randomly selected user, query other users who have transaction records with this user, and add them to the user individual set of the subgraph Then, repeat this process and gradually expand the subgraph until the number of users in the subgraph exceeds 60% of the total number of users in the source scenario. During this process, add all the transaction behaviors between these users to the edge set and add the behavior feature representations of these users to the feature set of the subgraph In particular, denote the set of normal user features as and denote the set of user features with fraud behaviors as In this way, we can generate multiple connected subgraphs, thereby increasing the diversity of feature representations and helping the model better capture the complexity and heterogeneity of user behaviors.
[0069] Step C: For each generated connected subgraph, calculate the mean of the feature representations of the users with and without abnormal behaviors, denoted as and and use these means as prior knowledge for aligning the user behavior feature representations in the subsequent target scenario. The specific calculation formula is as follows:
[0070]
[0071] 3. For the pseudo-label generation method combining the proportion constraint and the confidence condition described in this embodiment, calculate the proportion between the positive and negative pseudo-labels to be generated using the performance evaluation metrics of the above source scenario. Evaluate the performance of the model in the target scenario using the same metrics, and determine the number of pseudo-labels to be generated by introducing the confidence condition. Finally, generate the corresponding pseudo-labels based on the Wasserstein distance between the means and of the user feature representations in the target scenario and the source scenario. The specific steps are as follows:
[0072] Step D: For similar detection scenarios, there are a large number of common feature representations in the individual behaviors of users. After the model learns the prior knowledge in the source scenario, it will also have a certain fraud detection ability in the target scenario. Calculate the proportion of positive and negative pseudo-labels to be generated in the target scenario using the performance evaluation metrics of the model in the source scenario obtained in Step A. The calculation formula is as follows:
[0073]
[0074]
[0075] where is the fraud detection result of the model in the current round of training in the target scenario.
[0076] Step E1: To increase the diversity of the pseudo-label samples and reduce the impact of the instability of the fraud detection results on pseudo-label generation, add a small-range random fluctuation to the calculated proportion. Since the model initially does not learn the unique feature representations in the target scenario and cannot identify the specific abnormal features therein, this fluctuation will slightly increase the proportion of positive samples and decrease the proportion of negative samples. The specific calculation method is as follows:
[0077]
[0078] where c is the small-range random fluctuation added, and the effective range is [0, 0.05].
[0079] Step E2: To ensure the quality of the pseudo-labels, introduce the confidence condition parameter α to calculate the proportion of the generated positive and negative pseudo-labels in the total number of samples. The calculation formula is as follows:
[0080]
[0081] Among them, GCN(·) is a graph fraud detection model.
[0082] Step E3: According to the ratio between the positive and negative samples after adding fluctuations and their respective proportions in the total number of samples, the number of positive and negative pseudo-labels to be generated can be determined and The calculation formula is as follows:
[0083]
[0084] Step F: Calculate the means of the user feature representations in the target scenario and the source scenario and The Wasserstein distance, and the calculation formula is as follows:
[0085]
[0086] Among them, is all possible combinations of x t and .
[0087] Sort the features of user individuals in the target scenario according to the calculated Wasserstein distance, and add positive pseudo-labels to the top user features with the smallest distance to , and add negative pseudo-labels to the top user features with the smallest distance to . To consider the abnormal user representations unique to the target scenario that the model fails to identify, add positive pseudo-labels to the top c user features with the smallest sum of distances to and .
[0088] 4. The pseudo-label generation method described in this embodiment uses the step-by-step mixing strategy to further train the model to enhance the model's ability to represent the unique features of users in the target scenario. Finally, high-precision and high-reliability fraud detection of the model network in the target financial scenario is achieved. The specific steps are as follows:
[0089] Step G: By linearly adding the abnormal and normal user individual features in the target scenario and the source scenario, gradually increase the proportion of the target scenario features to further enhance the model's representation ability in the target scenario. The specific formula is as follows:
[0090]
[0091] Among them, and are the individual user characteristics with positive and negative pseudo-labels added in the target scenario. and are the mixing parameters for the e-th training round, and their calculation formulas are as follows:
[0092]
[0093] where e is the current training round, E is the total number of training rounds, q P and q N are the weight factors for the samples of individual user characteristics labeled with positive and negative pseudo-labels, and their calculation formulas are as follows:
[0094]
[0095] where exp(x) represents e to the power of x in natural logarithm.
[0096] Step H: Deploy the trained fraud detection network model in the target financial scenario by given the information or initial features of the users to be detected and their transaction behavior records. The model can quickly and accurately detect fraud in these users' past behaviors. The system has good scene generalization ability. For a newly added target scenario, the same training method can be used to improve the fraud detection ability of the model in the newly added scenario, thus effectively improving the execution efficiency and intelligent level of detecting fraud users in multiple financial scenarios.
[0097] The method of this embodiment proposes a comprehensive training framework that can accurately detect fraud users, overcoming the limitations of traditional methods in low data volume cases with low detection recall rate and inapplicability to multiple scenarios. Specifically, in this embodiment, the GCN graph neural network is used to initially learn the prior knowledge in the covered scenarios, and the average features of users with fraud behaviors are obtained by combining empirical statistical methods, enabling the detection model to learn general feature representations; then, confidence conditions are introduced to generate corresponding pseudo-labels in the target detection scenario to assist the model in learning the unique fraud user feature representations in the target scenario; finally, the learned feature representation methods are combined to improve the fraud detection ability of the model in the target scenario. By continuously introducing prior knowledge and generating pseudo-label data, this embodiment significantly improves the user fraud behavior detection ability of the fraud detection model in multiple financial scenarios, providing reliable technical support for tasks such as identifying or warning user fraud behaviors and tracking illegal financial criminals in the financial field, and has important research and application value.
[0098] Those of ordinary skill in the art will understand that all or part of the steps of implementing the above embodiments can be accomplished by hardware, or can be accomplished by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, where the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0099] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and supplements can be made, and these improvements and supplements should also be regarded as the protection scope of the present invention.
Claims
1. A financial fraud detection method based on similar scenario learning and pseudo-label data generation, characterized in that, It includes the following steps: S1. First, use the graph fraud detection model based on GCN to learn the prior knowledge required for fraud detection in the source scenario, i.e., the scenario similar to the target scenario, and evaluate the detection ability of the model, so that the graph fraud detection model has the general feature representation ability of fraudulent users; S2. Then, improve the diversity of the prior knowledge of the source scenario by generating the connected subgraphs of the source scenario, and calculate the average feature representations of fraudulent and normal users in each source sub-scenario; S3. Subsequently, generate pseudo-label data according to the Wasserstein distance in the target scenario and the source sub-scenarios; S4. Next, further train the graph fraud detection model by using the generated pseudo-label data and the step-by-step hybrid strategy fraud detection, so that the graph fraud detection model has the unique feature representation ability of fraudulent users in the target scenario; S5. Finally, deploy the trained graph fraud detection model to the actual application scenario to accurately detect the fraud behavior of users in the given user transaction network.
2. The financial fraud detection method based on similar scenario learning and pseudo-label data generation according to claim 1, wherein, The task of financial fraud detection is defined as: Given a user transaction behavior network where v represents individual users, and ε represents transaction behaviors between users, representing the feature representations of the behaviors of all n individual users, and the goal is to detect whether there are abnormal transaction behaviors for each user; The financial fraud detection results are represented by where y i = 1 indicates the existence of abnormal behavior, and y i = 0 indicates the non - existence of abnormal behavior.
3. The financial fraud detection method based on similar scenario learning and pseudo-label data generation according to claim 2, wherein Step S1 specifically includes the following sub-steps: S11. Use the data in the source scenario to initially train a graph fraud detection model of GCN, so that the graph fraud detection model of GCN has the ability to detect abnormal behaviors, and use the cross-entropy loss as its loss function. The calculation formula is as follows: Among them, y i is the model prediction result, and is the actual situation; S12. Evaluate the ability and reliability of the graph fraud detection model for the abnormal behavior detection task based on the performance of the graph fraud detection model of the GCN on the test data. The evaluation metrics are the proportions of the true positives (TP) of abnormal behaviors, true negatives (TN) of normal behaviors, false positives (FP) of abnormal behaviors, and false negatives (FN) of normal behaviors correctly or incorrectly detected by the graph fraud detection model in the source scenario in the total test samples, which are respectively expressed as The specific calculation formulas are as follows: Among them, represents the true user detection result, represents the actual detection result, represents taking the inverse of the value in (·), n s represents the number of user individuals in the source scenario test data.
4. A financial fraud detection method based on similar scenario learning and pseudo-label data generation according to claim 3, characterized in that, Step S2 specifically includes the following sub-steps: S21. To enhance the diversity of the learned feature representations, for the user transaction behavior network in the source scenario Generate several connected subgraphs The specific generation sub-steps of each connected subgraph First, start with a randomly selected user and query other users who have transaction records with this user. The user individual set Next, repeat this process to gradually expand the subgraph until the number of users in the subgraph exceeds 60% of the total number of users in the source scenario. During this process, add the transaction behaviors among all these users to the edge set. And add the representation of the behavioral characteristics of these users to the feature set of the subgraph. In particular, denote the set of normal user characteristics among them as Denote the set of user characteristics with fraudulent behaviors as In this way, generate multiple connected subgraphs to increase the diversity of feature representations and help the graph fraud detection model better capture the complexity and heterogeneity of user behaviors. S22. For each generated connected subgraph, calculate the mean values of the feature representations of users with abnormal behaviors and users without abnormal behaviors respectively, denoted as and and use these mean values as prior knowledge for aligning the feature representations of user behaviors in subsequent target scenarios.
5. A financial fraud detection method based on similar scenario learning and pseudo-label data generation according to claim 4, characterized in that Step S3 specifically includes the following sub-steps: S31. Calculate the ratio of positive and negative pseudo-labels that should be generated in the target scenario through the performance evaluation index of the graph fraud detection model in the source scenario obtained in step S1. The calculation formula is as follows: Among them, is the fraud detection result of the model in the current round of training in the target scenario; S32. To increase the diversity of pseudo-label samples and reduce the impact of the instability of fraud detection results on pseudo-label generation, add a small-range random fluctuation to the calculated ratio. The specific calculation method is as follows: where c is the added small-range random fluctuation, and the effective range is [0, 0.05]; S33. To ensure the quality of pseudo-labels, introduce the confidence condition parameter α to calculate the ratio of the generated positive and negative pseudo-labels to the total number of samples. The calculation formula is as follows: where GCN(·) is the graph fraud detection model; S34. Determine the number of positive and negative pseudo-labels to be generated respectively according to the ratio between positive and negative samples after adding fluctuations and their respective proportions in the total number of samples. and The calculation formula is as follows: S35. Calculate the mean values of the user feature representations of the target scenario and the source scenario and for the Wasserstein distance. The calculation formula is as follows: Among them, is x t and all possible combinations; Sort the features of user individuals in the target scenario according to the calculated Wasserstein distance, and add positive pseudo-labels to the top user features with the smallest distance to , and add negative pseudo-labels to the top user features with the smallest distance to ; Considering the abnormal user representations unique to the target scenario that the graph fraud detection model fails to recognize, add positive pseudo-labels to the top c user features with the smallest sum of distances to and .
6. The financial fraud detection method based on similar scenario learning and pseudo-label data generation according to claim 5, characterized in that, Step S4 specifically includes the following sub-steps: Gradually increase the proportion of the target scenario features by linearly adding the abnormal and normal user individual features in the target scenario and the source scenario to further improve the representation ability of the graph fraud detection model in the target scenario. The specific formula is as follows: Among them, and are the user individual characteristics with positive and negative pseudo-labels added in the target scenario; and are the mixing parameters for the e-th training round, and the calculation formula is as follows: where e is the current training round, E is the total number of training rounds, and q P and q N are the weight factors for positive and negative pseudo-labeled user individual feature samples, and the calculation formula is as follows: where exp(x) represents e to the power of x.