Transaction security authentication method based on multi-source information fusion

By obtaining and analyzing the evolution and growth of fraud samples, adjusting sample weights and conducting LightGBM model training, the problem of difficult fraudulent behavior in traditional models that traditional models can identify new and growing trends is solved, and the recognition accuracy and fraud prevention capabilities of the model are improved.

CN120218934AActive Publication Date: 2025-06-27INSPUR SMART SUPPLY CHAIN TECH (SHANDONG) CO LTD

Patent Information

Application Number
CN202510704699.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-06-27
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

The traditional LightGBM model fails to effectively identify new types of fraud and the growth trend of fraud, making it difficult to accurately identify changes and new types of fraud.

Method used

By collecting the feature vectors of the training fraud samples and normal transaction samples for the month, obtain the evolution and growth of each training fraud sample, adjust the sample weights, and train the LightGBM model based on these weights to improve the recognition accuracy of the model.

Benefits of technology

Improve the accuracy of the LightGBM model to identify new and growing fraud behaviors, and enhance the fraud prevention capabilities of transaction security certification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218934A_ABST
    Figure CN120218934A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a transaction security authentication method based on multi-source information fusion, and the method comprises the steps: collecting a current month training fraud sample and a current month normal transaction sample, obtaining the evolution degree of each training fraud sample in the current month according to the difference between the feature vectors of the training fraud samples in the current month and the feature vectors of the historical fraud samples in the historical time periods before the training fraud samples in the current month, and obtaining the growth degree of each training fraud sample in the current month; obtaining the sample weight of each training fraud sample in the current month according to the growth degree and the evolution degree, endowing the training fraud sample in the current month with the corresponding sample weight, inputting the training fraud sample in the current month and a normal transaction sample in the current month into a LightGBM model for training to obtain a trained LightGBM model, and identifying the fraud behavior of the latest transaction record according to the trained LightGBM model. According to the invention, the accuracy of fraudulent behavior identification is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing. More specifically, the present invention relates to a transaction security authentication method based on multi-source information fusion. Background Art

[0002] Transaction security authentication is the core mechanism to ensure the security of funds and information in the online and offline transaction processes. By means of identity verification, data encryption, risk monitoring, etc., it verifies the authenticity of the identities of both parties in the transaction and the compliance of transaction behaviors, and prevents malicious intrusion, data leakage and illegal transactions. Among them, identifying fraud behavior is a key link in transaction security authentication. Once it fails to be identified in time, it will lead to personal property losses, damage to the reputation of enterprises, and even trigger systemic financial risks; Multi-source information fusion plays the role of an "intelligent filter" in this process. By integrating multi-dimensional data such as user transaction data, geographical location information, and behavior patterns, a dynamic risk assessment model is constructed. By quickly cross-verifying, it accurately judges whether there is a fraud risk, significantly improving the accuracy and timeliness of risk identification and becoming a solid defense line against fraud behavior.

[0003] Since the LightGBM model identifies fraud behavior only relying on each training fraud sample, and the traditional LightGBM model assigns the same weight to each training fraud sample, but the manifestation forms of fraud behavior in the transaction security authentication scenario are not static, but are in continuous evolution. Therefore, the traditional LightGBM model does not take into account the emergence of new fraud behaviors and the growth trend of fraud behaviors, making it difficult for the model to accurately identify when facing different growth trend changes and new fraud behaviors. Summary of the Invention

[0004] In order to solve the problem that the traditional LightGBM model does not take into account the emergence of new fraud behaviors and the growth trend of fraud behaviors, making it difficult for the model to accurately identify when facing different growth trend changes and new fraud behaviors, the present invention proposes a transaction security authentication method based on multi-source information fusion, and the method includes the following steps: Collect the feature vectors of the training fraud samples and normal transaction samples of the current month; obtain the historical fraud samples of each historical time period before each training fraud sample of the current month; According to the differences between the feature vectors of each training fraud sample of the current month and the historical fraud samples of each historical time period before it, obtain the evolution degree of each training fraud sample of the current month, Obtain the neighboring fraud samples in each historical time period before each training fraud sample of the current month according to the Manhattan distance between each training fraud sample of the current month and each historical fraud sample in its previous historical time periods; obtain the growth rate of each training fraud sample of the current month according to the growth of the total number of neighboring fraud samples in all historical time periods before each training fraud sample of the current month; Obtain the sample weight of each training fraud sample of the current month , represents the sample weight of the i-th training fraud sample of the current month; represents the evolution degree of the i-th training fraud sample of the current month; represents the growth rate of the i-th training fraud sample of the current month; represents a preset coefficient; Train the LightGBM model based on the sample weights of each training fraud sample of the current month, the training fraud samples of the current month, and the normal transaction samples of the current month to obtain a trained LightGBM model; identify fraud behaviors in the latest transaction records according to the trained LightGBM model.

[0005] The innovation of the present invention lies in obtaining the evolution degree of each training fraud sample of the current month according to the difference in the feature vectors between each training fraud sample of the current month and the historical fraud samples in its previous historical time periods, which reflects whether each training fraud sample of the current month is a new type of fraud behavior. Then, obtain the neighboring fraud samples in each historical time period before each training fraud sample of the current month, and obtain the growth rate of each training fraud sample of the current month according to the growth of the total number of the neighboring fraud samples, which reflects the frequency of occurrence of each training fraud sample of the current month. Finally, obtain the sample weight of each training fraud sample of the current month according to the growth rate and the evolution degree, so that when training the LightGBM model according to each training fraud sample of the current month subsequently, it can focus on new and growing fraud behaviors, improve the accuracy of model training, and further improve the accuracy of identifying fraud behaviors in the latest transaction records using the trained LightGBM model.

[0006] Preferably, the collection of the feature vectors of the training fraud samples of the current month and the normal transaction samples of the current month includes: Obtain the feature vectors of all transaction records of the current month in the trading system, and the feature vectors of the transaction records include several feature components; Obtain the transaction labels of each transaction record of the current month confirmed by manual review; preset the sample quantity H, and select H transaction records with the transaction label of fraud from all transaction records of the current month as the training fraud samples of the current month, and select H transaction records with the transaction label of normal from all transaction records of the current month as the normal transaction samples of the current month.

[0007] Preferably, obtaining historical fraud samples for each historical time period before each training fraud sample of the current month includes: The preset historical time period unit is one month; according to the acquisition method of the training fraud samples of the current month, obtain the training fraud samples in each historical time period before the transaction time of the i-th training fraud sample of the current month, and record them as the historical fraud samples for each historical time period before the i-th training fraud sample of the current month.

[0008] Preferably, obtaining the evolution degree of each training fraud sample of the current month includes: ; In the formula, represents the evolution degree of the i-th training fraud sample of the current month; represents the trading time interval between the i-th training fraud sample of the current month and the last historical fraud sample in the m-th historical time period before it; and represents the value of the j-th feature component of the i-th fraudulent training sample of the current month and the value of the j-th feature component of the k-th historical fraud sample in the m-th historical time period before it; and represents the number of historical time periods before the i-th training fraud sample of the current month and the number of historical fraud samples in each historical time period; represents the variance of the j-th feature component of all training fraud samples of the current month; represents the number of feature components of the i-th training fraud sample of the current month; || is the absolute value symbol; exp() is the exponential function with the natural constant as the base.

[0009] The greater the evolution degree of the training fraud sample, the more likely it is that the training fraud sample belongs to a new type of fraud behavior.

[0010] Preferably, obtaining the neighboring fraud samples in each historical time period before each training fraud sample of the current month includes: Preset a distance threshold R, obtain the Manhattan distance between the feature vector of the i-th training fraud sample of the current month and the feature vectors of each historical fraud sample in each historical time period before the i-th training fraud sample of the current month, and record the historical fraud samples in each historical time period corresponding to when the Manhattan distance is less than or equal to R as the neighboring fraud samples in each historical time period before the i-th training fraud sample of the current month.

[0011] It is convenient to subsequently obtain the growth degree of each training fraud sample of the current month according to the growth of the number of neighboring fraud samples in all historical time periods.

[0012] Preferably, the obtaining of the growth degree of each training fraud sample in the current month includes:

[0013] In the formula, represents the growth degree of the i-th training fraud sample in the current month; represents the total number of neighboring fraud samples of the i-th fraudulent training sample in the current month; represents the total number of neighboring fraud samples in the neighboring historical time period before the i-th training fraud sample in the current month; represents the total number of neighboring fraud samples in the (m + 1)-th historical time period before the i-th training fraud sample in the current month; represents the total number of neighboring fraud samples in the m-th historical time period before the i-th training fraud sample in the current month; represents the number of historical time periods before the i-th training fraud sample in the current month; represents a piecewise linear function; norm() represents a normalization function.

[0014] The greater the growth degree of the training fraud sample, the higher the frequency of occurrence of the training fraud sample.

[0015] Preferably, the obtaining of the total number of neighboring fraud samples of the i-th fraudulent training sample in the current month includes: Preset a distance threshold R, obtain the Manhattan distance between the feature vector of the i-th training fraud sample in the current month and the feature vectors of each other training fraud sample in the current month, and record the training fraud samples in the current month corresponding to when the Manhattan distance is less than or equal to R as the neighboring fraud samples of the i-th fraudulent training sample in the current month.

[0016] Preferably, the obtaining of the neighboring historical time period before the i-th training fraud sample in the current month includes: Obtain the trading time interval between the i-th training fraud sample in the current month and the last historical fraud sample in each previous historical time period of the i-th training fraud sample in the current month as the time interval between the i-th training fraud sample in the current month and each of its historical time periods; record the historical time period corresponding to the minimum value of the time interval as the neighboring historical time period before the i-th training fraud sample in the current month.

[0017] Preferably, the obtaining of the trained LightGBM model includes: Input the training fraud samples and normal transaction samples of the current month into the LightGBM model for training to obtain a trained LightGBM model. Among them, set the number of iteration rounds of the LightGBM model to 500, the maximum tree depth to 5, and select the binary cross-entropy loss function as the loss function of the LightGBM model. For the training fraud samples of the current month, assign corresponding sample weights, while the sample weights of the normal transaction samples of the current month are all set to 1.

[0018] Setting higher sample weights for new and increasingly frequent training fraud samples and then performing model training improves the accuracy of model training.

[0019] Preferably, identifying fraud behaviors in the latest transaction records based on the trained LightGBM model includes: Obtain the feature vector composed of the latest transaction records and input it into the trained LightGBM model to obtain the fraud confidence of the latest transaction records. Set a fraud confidence threshold T. When the fraud confidence of the latest transaction records is greater than or equal to T, mark the latest transaction records as fraud behaviors and freeze the accounts in real time to prevent these accounts from continuing to conduct fraud transactions.

[0020] The present invention has the following beneficial effects: The purpose of the present invention is to obtain the evolution degree of each training fraud sample of the current month based on the difference between the feature vectors of each training fraud sample of the current month and the historical fraud samples in its previous historical time period, which reflects whether each training fraud sample of the current month is a new type of fraud behavior. Then, obtain the neighboring fraud samples in each historical time period before each training fraud sample of the current month, and obtain the growth degree of each training fraud sample of the current month according to the growth situation of the total number of the neighboring fraud samples, which reflects the frequency of occurrence of each training fraud sample of the current month. Finally, obtain the sample weights of each training fraud sample of the current month according to the growth degree and the evolution degree, so that when the LightGBM model is trained based on each training fraud sample of the current month, it can focus on new and growing fraud behaviors, improve the accuracy of model training, and further improve the accuracy of identifying fraud behaviors in the latest transaction records using the trained LightGBM model. Description of the Drawings

[0021] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present invention will become easily understandable. In the drawings, several embodiments of the present invention are shown in an exemplary rather than restrictive manner, and the same or corresponding reference numerals represent the same or corresponding parts, where: Figure 1 is the step flow chart of a transaction security authentication method based on multi-source information fusion according to an embodiment of the present invention. Detailed implementation manners

[0022] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments.

[0023] Please refer to Figure 1 , which shows a step flow chart of a transaction security authentication method based on multi-source information fusion provided by an embodiment of the present invention. The method includes the following steps: S001. Collect the training fraud samples of the current month, the normal transaction samples of the current month, and the historical fraud samples of each historical time period before each training fraud sample of the current month.

[0024] In the embodiments of the present invention, all transaction records of the current month in the transaction system are obtained. For any transaction record of the current month, the transaction amount of the transaction record, the number of transactions of the affiliated user in the recent 7 days, the time interval from the last transaction of the affiliated user (numerical feature, which needs to be normalized / standardized), the payment tool of the affiliated user (categorical feature, which needs to be label-encoded), the maximum transaction amount of the affiliated user, the transaction timestamp, the transaction IP place of origin, whether the transaction IP is a known risk (boolean feature, processed with 0 / 1), and the transaction label (fraud / normal, represented by 1 / 0) confirmed by manual review are used as the feature vector of the transaction record. The feature vector of the transaction record includes several feature components; the feature vectors of all transaction records of the current month are obtained; A preset sample quantity H = 500. In other embodiments, the value of H is preset according to specific implementation situations. Among all transaction records of the current month, H transaction records with a transaction label of fraud are selected as the training fraud samples of the current month, and H transaction records with a transaction label of normal are selected as the normal transaction samples of the current month. The training fraud samples of the current month and the normal transaction samples of the current month are used as the training samples of the current month.

[0025] The preset historical time period unit is one month, and the preset number of historical time periods is M = 6; according to the acquisition method of the training fraud samples of the current month, the training fraud samples in each historical time period before the transaction time of the i-th training fraud sample of the current month are obtained, and are denoted as the historical fraud samples of each historical time period before the i-th training fraud sample of the current month. S002. Obtain the evolution degree of each training fraud sample of the current month according to the differences between the feature vectors of the training fraud samples of the current month and the historical fraud samples of each historical time period before the training fraud samples of the current month.

[0026] It should be noted that since the LightGBM model's identification of fraud behavior only relies on individual training fraud samples, and the traditional LightGBM model assigns the same weight to each training fraud sample. However, the manifestation forms of fraud behavior in the transaction security authentication scenario are not static but are in a continuous evolution process. Therefore, the traditional LightGBM model does not take into account the emergence of new types of fraud behavior and the growth trend of fraud behavior, making it difficult for the model to accurately identify when faced with different growth trend changes and new types of fraud behavior. Therefore, each training fraud sample in the current month is analyzed to obtain the sample weight of each training fraud sample, improving the LightGBM model's identification ability for evolving fraud behavior.

[0027] It should be further noted that if the difference between any fraud training sample in the current month and the feature vectors of each historical fraud sample in each historical time period before this training fraud sample in the current month is relatively large, it indicates that the evolution degree of this fraud training sample is higher. At this time, the possibility that this fraud training sample is a new type of fraud behavior sample is higher. Therefore, first, the difference between any fraud training sample in the current month and the feature vectors of each historical fraud sample in each historical time period before this training fraud sample in the current month is combined to obtain the evolution degree of this fraud training sample in the current month.

[0028] In the embodiment of the present invention, the evolution degree of each training fraud sample in the current month is obtained: ; In the formula, represents the evolution degree of the i-th training fraud sample in the current month; represents the trading time interval between the last historical fraud sample in the m-th historical time period before the i-th training fraud sample in the current month and the i-th training fraud sample; represents the value of the j-th feature component of the i-th fraud training sample in the current month; represents the value of the j-th feature component of the k-th historical fraud sample in the m-th historical time period before the i-th training fraud sample in the current month; represents the number of historical time periods before the i-th training fraud sample in the current month; represents the number of historical fraud samples in each historical time period before the i-th training fraud sample in the current month; represents the variance of the j-th feature component of all training fraud samples in the current month; represents the number of feature components of the i-th training fraud sample; the number of features in the i-th training fraud sample is the same as the number of feature components of the k-th fraud sample in the m-th historical time period before the i-th training fraud sample; || represents the absolute value symbol; exp() represents the exponential function with the natural constant as the base. represents the difference degree between the i-th fraud training sample of the current month and the k-th historical fraud sample in the m-th historical time period before the i-th training fraud sample of the current month. The higher its value, the higher the evolution degree of the i-th fraud training sample of the current month, indicating a higher possibility that the i-th fraud sample of the current month is a new type of fraud behavior; The larger the value of, it indicates that the trading time interval between the last historical fraud sample in the m-th historical time period before the i-th training fraud sample of the current month and the i-th training fraud sample of the current month is shorter. At this time, it indicates that the influence weight of the difference degree between the i-th fraud training sample of the current month and the k-th historical fraud sample in the m-th historical time period before the i-th training fraud sample of the current month is greater.

[0029] S003. Obtain the neighboring fraud samples in each historical time period before each training fraud sample of the current month, and obtain the growth degree of each training fraud sample of the current month according to the growth trend of the number of neighboring fraud samples in the historical time period before each training fraud sample of the current month; according to the growth degree of each training fraud sample of the current month and the evolution degree of each training fraud sample of the current month, obtain the sample weight of each training fraud sample of the current month.

[0030] It should be noted that the evolution degree of each training fraud sample of the current month only represents the difference between each fraud sample of the current month and each historical fraud sample in each previous historical time period, and cannot effectively identify the growth trend of the increasing frequency of the fraud behavior represented by the fraud sample; In the present invention, first, according to the distance between the feature vector of any training fraud sample of the current month and the feature vectors of each historical fraud sample in each historical time period before the training fraud sample, obtain the neighboring fraud samples in each historical time period before each training fraud sample of the current month. Therefore, when the number of neighboring fraud samples in each historical time period before the training fraud sample of the current month increases in chronological order and the greater the increase amplitude of each word, it indicates a higher possible degree that the fraud behavior represented by the training fraud sample of the current month has an increasing frequency. Then, the growth degree of the training fraud sample is higher.

[0031] In the embodiment of the present invention, a preset distance threshold R is set, and the Manhattan distance between the feature vector of the i-th training fraud sample of the current month and the feature vectors of each historical fraud sample in each historical time period before the i-th training fraud sample of the current month is obtained. The historical fraud samples in each historical time period corresponding to when the Manhattan distance is less than or equal to R are recorded as the neighboring fraud samples in each historical time period before the i-th training fraud sample of the current month; Obtain the Manhattan distance between the feature vector of the i-th training fraud sample in the current month and the feature vectors of each other training fraud sample in the current month. Denote the training fraud samples in the current month for which the Manhattan distance is less than or equal to R as the neighboring fraud samples of the i-th fraud training sample in the current month; Obtain the trading time interval between the i-th training fraud sample in the current month and the last historical fraud sample in each previous historical time period of the i-th training fraud sample in the current month, as the time interval of the i-th training fraud sample in the current month with each of its historical time periods; Denote the historical time period corresponding to the minimum value of the time interval as the neighboring historical time period before the i-th training fraud sample in the current month; Obtain the growth degree of each training fraud sample in the current month:

[0032] In the formula, represents the growth degree of the i-th training fraud sample in the current month; represents the total number of neighboring fraud samples of the i-th fraud training sample in the current month; represents the total number of neighboring fraud samples in the neighboring historical time period before the i-th training fraud sample in the current month; represents the total number of neighboring fraud samples in the (m + 1)-th historical time period before the i-th training fraud sample in the current month; represents the total number of neighboring fraud samples in the m-th historical time period before the i-th training fraud sample in the current month; represents the number of historical time periods before the i-th training fraud sample in the current month; represents a piecewise linear function for retaining positive values. When is less than or equal to 0, is 0. When is greater than 0, is the difference; norm() represents a normalization function; represents the growth amount of the total number of neighboring fraud samples of the i-th fraud training sample in the current month compared to the total number of neighboring fraud samples in its most recent historical time period; represents the sum of the growth amounts of the neighboring fraud samples in each historical time period before the i-th training fraud sample in the current month compared to the neighboring fraud samples in the previous historical time period; The larger the value of

[0033] It should be noted that when the evolution degree and growth degree of any training fraud sample in a month are higher, it indicates that the active training fraud sample is a new type of fraud sample with a growth trend, and a relatively large sample weight needs to be set for key identification.

[0034] In the embodiment of the present invention, the sample weight of each training fraud sample in the current month is obtained: ; In the formula, represents the sample weight of the i-th training fraud sample in the current month; represents the evolution degree of the i-th training fraud sample in the current month; represents the growth degree of the i-th training fraud sample in the current month; represents a preset coefficient. In the embodiment of the present invention, the preset coefficient , multiply the value of by so that the value result is between 0 and 2; the training fraud sample with a higher evolution degree and a relatively higher growth degree is a new type of fraud sample with a growth trend. Therefore, the sample weight of this training fraud sample is higher and needs to be key identified.

[0035] S004. Assign corresponding sample weights to the training fraud samples in the current month, and input them together with the normal transaction samples in the current month into the LightGBM model for training to obtain a trained LightGBM model. According to the trained LightGBM model, identify the fraud behavior in the latest transaction record.

[0036] In the embodiment of the present invention, the training fraud samples in the current month and the normal transaction samples in the current month are input into the LightGBM model for training to obtain a trained LightGBM model; among them, the number of iterations of the LightGBM model is set to 500, the maximum tree depth is 5, and the binary cross-entropy loss function is selected as the loss function of the LightGBM model; for the training fraud samples in the current month, corresponding sample weights are assigned; while the sample weights of the normal transaction samples in the current month are all set to 1; the specific training process is well-known content of the LightGBM model, and the specific training process will not be elaborated in this embodiment.

[0037] Obtain the feature vector composed of the latest transaction record, and input the obtained feature vector into the trained LightGBM model to obtain the fraud confidence of the latest transaction record; Set the fraud confidence threshold T = 0.7. In other embodiments, the implementer can preset the value of the fraud confidence threshold according to the specific implementation method. When the fraud confidence of the latest transaction record is greater than or equal to the fraud confidence threshold, mark the latest transaction record as a fraud behavior and freeze the account in real time to prevent the account from continuing to conduct fraud transactions.

[0038] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A transaction security authentication method based on multi-source information fusion, characterized in that, Including: Collecting the feature vectors of the training fraud samples and the normal transaction samples in the current month; Obtaining the historical fraud samples in each historical time period before each training fraud sample in the current month; Obtaining the evolution degree of each training fraud sample in the current month according to the differences between the feature vectors of each training fraud sample in the current month and the historical fraud samples in each previous historical time period; Obtaining the neighboring fraud samples in each historical time period before each training fraud sample in the current month according to the Manhattan distances between each training fraud sample in the current month and the historical fraud samples in each previous historical time period; obtaining the growth degree of each training fraud sample in the current month according to the growth of the total number of neighboring fraud samples in all historical time periods before each training fraud sample in the current month; Obtain the sample weights of each training fraud sample for the current month , represents the sample weight of the i-th training fraud sample for the current month; represents the evolution degree of the i-th training fraud sample for the current month; represents the growth degree of the i-th training fraud sample for the current month; represents a preset coefficient; Training the LightGBM model based on the sample weights of each training fraud sample in the current month, the training fraud samples in the current month, and the normal transaction samples in the current month to obtain the trained LightGBM model; identifying the fraud behavior of the latest transaction records according to the trained LightGBM model.

2. The transaction security authentication method based on multi-source information fusion according to claim 1, wherein, The collecting the feature vectors of the training fraud samples and the normal transaction samples in the current month includes: Obtaining the feature vectors of all transaction records in the current month in the trading system, where the feature vectors of the transaction records include several feature components; Obtaining the transaction labels of each transaction record in the current month confirmed by manual review; presetting the sample quantity H, and selecting H transaction records with the transaction label of fraud from all transaction records in the current month as the training fraud samples in the current month, and selecting H transaction records with the transaction label of normal from all transaction records in the current month as the normal transaction samples in the current month.

3. The transaction security authentication method based on multi-source information fusion according to claim 2, characterized in that The obtaining the historical fraud samples in each historical time period before each training fraud sample in the current month includes: Presetting the historical time period unit as one month, and obtaining the training fraud samples in each historical time period before the transaction time of the i-th training fraud sample in the current month according to the collection method of the training fraud samples in the current month, and recording them as the historical fraud samples in each historical time period before the i-th training fraud sample in the current month.

4. A transaction security authentication method based on multi-source information fusion according to claim 1, characterized in that, The obtaining the evolution degree of each training fraud sample in the current month includes: ; Wherein, represents the evolution degree of the i-th training fraud sample in the current month; represents the trading time interval between the i-th training fraud sample in the current month and the last historical fraud sample in the m-th historical time period before it; and represents the value of the j-th feature component of the i-th fraud training sample in the current month and the value of the j-th feature component of the k-th historical fraud sample in the m-th historical time period before it; and represents the number of historical time periods before the i-th training fraud sample in the current month and the number of historical fraud samples in each historical time period; represents the variance of the j-th feature component of all training fraud samples in the current month; represents the number of feature components of the i-th training fraud sample in the current month; || is the absolute value symbol; exp() is the exponential function with the natural constant as the base.

5. A transaction security authentication method based on multi-source information fusion according to claim 1, characterized in that The obtaining the neighboring fraud samples in each historical time period before each training fraud sample in the current month includes: Presetting the distance threshold R, obtaining the Manhattan distances between the feature vector of the i-th training fraud sample in the current month and the feature vectors of each historical fraud sample in each historical time period before the i-th training fraud sample in the current month, and recording the historical fraud samples in each historical time period corresponding to when the Manhattan distance is less than or equal to R as the neighboring fraud samples in each historical time period before the i-th training fraud sample in the current month.

6. The transaction security authentication method based on multi-source information fusion according to claim 1, wherein The obtaining the growth degree of each training fraud sample in the current month includes: In the formula, represents the growth rate of the i-th training fraud sample in the current month; represents the total number of neighboring fraud samples of the i-th fraud training sample in the current month; represents the total number of neighboring fraud samples in the neighboring historical time period before the i-th training fraud sample in the current month; represents the total number of neighboring fraud samples in the (m + 1)-th historical time period before the i-th training fraud sample in the current month; represents the total number of neighboring fraud samples in the m-th historical time period before the i-th training fraud sample in the current month; represents the number of historical time periods before the i-th training fraud sample in the current month; represents a piecewise linear function; norm() represents a normalization function.

7. A transaction security authentication method based on multi-source information fusion according to claim 6, characterized in that, The obtaining of the total number of the neighboring fraud samples of the i-th fraud training sample in the current month includes: Preset a distance threshold R, obtain the Manhattan distances between the feature vector of the i-th training fraud sample in the current month and the feature vectors of each other training fraud sample in the current month, and record the training fraud samples in the current month corresponding to the Manhattan distances less than or equal to R as the neighboring fraud samples of the i-th fraud training sample in the current month.

8. The transaction security authentication method based on multi-source information fusion according to claim 6, wherein, The obtaining of the neighboring historical time period before the i-th training fraud sample in the current month includes: Obtain the trading time interval between the i-th training fraud sample in the current month and the last historical fraud sample in each previous historical time period of the i-th training fraud sample in the current month as the time interval between the i-th training fraud sample in the current month and each of its historical time periods; record the historical time period corresponding to the minimum value of the time interval as the neighboring historical time period before the i-th training fraud sample in the current month.

9. A transaction security authentication method based on multi-source information fusion according to claim 1, characterized in that, The obtaining of the trained LightGBM model includes: Input the training fraud samples and normal transaction samples in the current month into the LightGBM model for training to obtain a trained LightGBM model; wherein, set the number of iteration rounds of the LightGBM model to 500, the maximum tree depth to 5, select the binary cross-entropy loss function as the loss function of the LightGBM model; for the training fraud samples in the current month, assign corresponding sample weights; and set the sample weights of the normal transaction samples in the current month to 1.

10. A transaction security authentication method based on multi-source information fusion according to claim 1, characterized in that, The identifying of the fraud behavior of the latest transaction record according to the trained LightGBM model includes: Obtain the feature vector composed of the latest transaction record and input it into the trained LightGBM model to obtain the fraud confidence of the latest transaction record; set a fraud confidence threshold T, when the fraud confidence of the latest transaction record is greater than or equal to T, mark the latest transaction record as a fraud behavior and freeze the account in real time to prevent the account from continuing to conduct fraud transactions.

Citation Information

Patent Citations

  • A method for short-term state prediction of a large-scale road network base on an improved k-nearest neighbor

    CN109190797A

  • Topic popularity prediction method and device, equipment and computer storage medium

    CN116187298A

  • Anti-fraud behavior detection system supported by graph database

    CN118733834A

  • Fraud prediction model training method and device, equipment, storage medium and product

    CN118797338A

  • Anti-fraud model training method based on deep learning

    CN119599675A

Cited By

  • An intelligent model building method and system for electricity fraud prevention

    CN122508383A