Continuous abnormal transaction detection method and system for dynamic data changes
By assigning pseudo-labels to the data samples to be exited and adding extended classification headers to the detection model for fine-tuning training, the problem of data deletion when old customers exit in the existing technology is solved, and abnormal transaction detection of dynamic data exit is realized, ensuring the accuracy and stability of the detection.
Patent Information
- Application Number
- CN202510308072.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2045-03-17
AI Technical Summary
The existing abnormal transaction detection system is difficult to meet the data deletion needs of old customers when they exit. Direct deletion of old customer information will affect the detection accuracy and stability of the model.
By assigning pseudo-labels to the data samples to be exited, a new training pair is formed, and an extended classification header is added to the output layer of the detection model, fine-tuning training and cropping is performed, abnormal transaction detection of dynamic data exit is realized.
Effectively erase the contribution of the data to be exited in the model, ensure that the model's judgment of normal transaction samples is not affected, and the accuracy and stability of detection are maintained.
Smart Images

Figure CN119848629B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of abnormal transaction detection, and particularly to a continuous abnormal transaction detection method and system for dynamically changing data. Background Art
[0002] In the field of fintech, abnormal transaction detection is one of the important pillars for financial institutions to prevent risks and operate in compliance. It can identify various potential risky transactions. By real-time monitoring transaction data and promptly discovering possible fraud behaviors, financial institutions can effectively protect customer assets and maintain the stability of the financial market. Existing abnormal transaction detection systems usually only train on existing customer data once and then go live. It is difficult to meet the withdrawal requirements of old customers. Specifically, to meet privacy protection or compliance needs, old customers often hope to completely delete their transaction records in the model when withdrawing, which further increases the requirements for the flexibility and scalability of the model. For users who submit withdrawal requests, simply deleting their corresponding data in the database is not feasible because the abstract knowledge contributed by these data has been accumulated in the model after the detection model training is completed and has become part of the learned feature space. If the old customer information is directly removed, it will damage the existing detection ability of the model, thereby affecting the accuracy and stability of detection. Summary of the Invention
[0003] The purpose of the present invention is to overcome the problems of the prior art and provide a continuous abnormal transaction detection method and system for dynamically changing data.
[0004] The purpose of the present invention is achieved through the following technical solutions: for the continuous abnormal transaction detection method for dynamically changing data, the method includes the abnormal transaction detection step of dynamic data withdrawal:
[0005] Assign a pseudo-label to each data sample to be withdrawn, and form a new training pair with the corresponding data sample to be withdrawn; the data sample to be withdrawn is a historical transaction record, including the transaction time, transaction location, transaction amount, transaction frequency, and transaction account of each user's each financial payment tool;
[0006] Add an extended classification head to the output layer of the detection model to process the pseudo-label, thereby expanding the decision space of the detection model;
[0007] Use the new training pair to fine-tune the detection model. The original classification head of the detection model remains unchanged, and only backpropagation is performed according to the extended classification head, thereby updating the parameters of the detection model;
[0008] Prune the extended classification head of the detection model to obtain the first optimized detection model;
[0009] Execute the abnormal transaction detection task using the first optimized detection model and output the transaction detection result.
[0010] In one example, before executing the abnormal transaction detection task using the first optimized detection model, it further includes:
[0011] Perform a linear operation on the parameters of the first optimized detection model to obtain a second optimized detection model with updated parameters;
[0012] Execute the abnormal transaction detection task using the second optimized detection model and output the transaction detection result.
[0013] In one example, the calculation expression for fine-tuning the detection model using the new training pair is:
[0014] ;
[0015] Among them, represents the parameters of the detection model after fine-tuning training; represents the parameters of the detection model before fine-tuning training; represents the parameters of the model used for optimization during the fine-tuning training process; represents the cross-entropy loss function; represents the feature vector of the data sample to be exited; represents the pseudo-label; represents the data label; represents the set of data samples to be exited.
[0016] In one example, the fine-tuning training of the detection model using the new training pair and the pruning of the extended classification head of the detection model constitute the data forgetting process of the detection model. The constraint condition of the data forgetting process is:
[0017] ;
[0018] Among them, represents the parameters of the detection model after fine-tuning training; represents the feature vector of the data sample to be exited; represents the data label; represents the prediction output of the detection model for the input data after fine-tuning training and pruning of the extended classification head.
[0019] In one example, the method further includes the abnormal transaction detection step of dynamically increasing data:
[0020] Train the detection model for the old task using the old data samples;
[0021] Adopt new data samples, and perform new task training on the detection model trained for the old task based on a regularization method or a module optimization method or a sample replay method, and then update the parameters of the detection model to obtain a third optimized detection model;
[0022] Use the third optimized detection model to perform the abnormal transaction detection task and output the transaction detection result.
[0023] In one example, when training the detection model based on the continuous learning method of the regularization method, it includes:
[0024] During the old task training process, calculate the Fisher information matrix of each parameter of the detection model;
[0025] During the new task training process, adopt the elastic weight consolidation mechanism to introduce a regularization term in the loss function to constrain the important parameters of the detection model in the old task, and thus update the parameters of the detection model.
[0026] In one example, during the new task training process, the expression of the loss function is:
[0027] ;
[0028] Among them, represents the loss function for training the new task by introducing the elastic weight consolidation mechanism; represents the loss function for directly training the old task without introducing the elastic weight consolidation mechanism; is the regularization coefficient; represents the Fisher information matrix; represents the data label; represents the network parameters during the training of the new task; represents the optimal parameters obtained after completing the old task.
[0029] In one example, before using the first optimized detection model or the third optimized detection model to perform the abnormal transaction detection task, it further includes:
[0030] Perform a linear operation on the parameters of the first optimized detection model and the parameters of the third optimized detection model to obtain the final detection model;
[0031] Use the final detection model to perform the abnormal transaction detection task and output the transaction detection result.
[0032] It should be further noted that the technical features corresponding to the above examples can be combined or replaced with each other to form a new technical solution.
[0033] The present invention also includes a continuous abnormal transaction detection system for dynamically changing data. The system includes an interconnected terminal and a server. The server includes a storage server and a training server;
[0034] The terminal inputs the original historical transaction records into the storage server by accessing the storage server; the terminal inputs the data increase and decrease identifier and the training data set into the training server by accessing the training server, and the training data set includes data samples to be exited, old data samples, and new data samples;
[0035] The training server is used to perform the step of training the detection model in the continuous abnormal transaction detection method for dynamically changing data formed by any one of the above examples or a combination of multiple examples;
[0036] The storage server is used to save the second optimized detection model or the third optimized detection model or the final detection model that has completed training;
[0037] The terminal and / or the training server and / or the storage server uses the second optimized detection model or the third optimized detection model or the final detection model to perform the abnormal transaction detection task and outputs the transaction detection result.
[0038] Compared with the prior art, the beneficial effects of the present invention are:
[0039] 1. In one example, by assigning pseudo-labels to the data to be exited and performing fine-tuning training, in the fine-tuning training, only backpropagation is performed according to the extended classification head, so that the decision space corresponding to the pseudo-class is moved to a position different from the original normal detection decision boundary of the model, and by pruning the extended classification head, the decision space corresponding to the pseudo-class is inactivated, thereby effectively erasing the contribution of the data to be exited in the model and not affecting the model's judgment of normal transaction samples, that is, it can avoid damaging the existing detection ability of the model, thus ensuring the accuracy and stability of detection.
[0040] 2. In one example, by performing linear operations on the parameters of the first optimized detection model, the possibility of changing the optimal parameter space of the model caused by model fine-tuning can be minimized, thereby ensuring the detection performance of the model for normal transaction samples (remaining data) that have not been exited.
[0041] 3. In one example, by designing the constraint conditions of the data forgetting process, the model can only forget the prediction of the data samples to be exited, thereby effectively deleting the performance of the data samples to be exited on the model and retaining the detection performance of the model for other data samples.
[0042] 4. In one example, by introducing regularization methods, module optimization methods, and sample replay methods to train the detection model for new tasks, the model can avoid "catastrophic forgetting" of the original detection ability during the dynamic update process, and can effectively adapt to the environmental changes caused by the dynamic change of data, which better meets the requirements of real-world continuous abnormal transaction detection.
[0043] 5. In one example, the present invention measures the importance of each parameter in the old task through the Fisher information matrix, and then ensures that important parameters are constrained during the training of the new task (the task of newly adding data samples), avoiding forgetting old knowledge, so as to ensure the detection performance in the case of adding new data samples.
[0044] 6. In one example, by performing linear operations on the parameters of the first optimized detection model and the parameters of the third optimized detection model, the parameter ratios before and after the model's active forgetting can be flexibly adjusted, that is, choosing to pay more attention to the anomaly detection performance of other non-withdrawn transaction samples, or choosing to pay more attention to the deletion degree of the data samples to be withdrawn in the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The following further describes in detail the specific implementation manners of the present invention with reference to the accompanying drawings. The accompanying drawings provided herein are used to provide a further understanding of the present application and form a part of the present application. The same reference numerals are used to represent the same or similar parts in these drawings. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application.
[0046] Figure 1 It is a flowchart of data dynamic withdrawal processing provided for an example of the present invention;
[0047] Figure 2 It is a block diagram of data dynamic withdrawal processing provided for an example of the present invention;
[0048] Figure 3 It is a CNN model architecture diagram provided for an example of the present invention;
[0049] Figure 4 It is a schematic diagram of the elastic weight consolidation mechanism provided for an example of the present invention;
[0050] Figure 5 It is a schematic diagram of data dynamic change provided for an example of the present invention;
[0051] Figure 6 It is a system framework diagram provided for an example of the present invention.
[0052] In the figure: 101 - the first terminal; 102 - the second terminal; 103 - the server. DETAILED DESCRIPTION OF THE INVENTION
[0053] The following clearly and completely describes the technical solutions of the present invention with reference to the accompanying drawings. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.
[0054] In the description of the present invention, ordinal numbers (e.g., "first and second", etc.) are used to distinguish objects, not limited to this order, and should not be construed as indicating or implying relative importance. Unless otherwise clearly specified and limited, terms such as "interconnected" and "connected" should be understood in a broad sense. For example, they can be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0055] In addition, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0056] In one example, as Figure 1 - Figure 2 shown, a method for continuously detecting abnormal transactions for dynamically changing data includes a step of detecting abnormal transactions for dynamic data exit:
[0057] S11: Assign a pseudo-label to each data sample to be exited, and form a new training pair in cooperation with the corresponding data sample to be exited.
[0058] Specifically, in the financial field, the data sample to be exited is the data sample of the customer (user) to be exited, which can be the historical transaction records of each customer, such as living consumption records, bank loan records, securities trading records, and other transaction information records, etc. Specifically, it includes the transaction time, transaction location, transaction amount, transaction frequency, transaction account, etc. of each user's various financial payment tools. Here, the number of users can be dozens or hundreds or even more massive, so the number of users is limited.
[0059] Further, in step S11, denote the data sample set to be exited as , assign a pseudo-label to each data sample to be exited in the data sample set to be exited, and form a new training pair , represents the feature vector of the data sample to be exited, and this pseudo-label is not the true class label of the data sample to be exited.
[0060] S12: Add an extended classification head to the output layer of the detection model to process the pseudo-label, thereby expanding the decision space of the detection model.
[0061] Specifically, add a new neuron to the output layer of the original detection model to add an extended classification head, thereby expanding the decision space of the model for the pseudo-label .
[0062] Furthermore, the detection model is preferably a deep learning model, which can discover the intricate structures in the empirical data for learning. By constructing a deep learning model with multiple processing layers, multiple levels of abstract representations of data can be created. Deep learning models include Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN), Long Short-Term Memory (LSTM), Gated Recurrent Unit (GRU), Transformer models, etc. A convolutional neural network is a deep learning model specifically designed to process data with a grid structure (such as images), and its core mechanism is to extract the features of the input data through convolutional operations. The architecture of a convolutional neural network usually consists of multiple layers, including convolutional layers, pooling layers, and fully connected layers. As Figure 3 shown, the convolutional neural network used in this example includes a sequentially connected convolutional layer 1, convolutional layer 2, pooling layer 1, convolutional layer 3, pooling layer 2, fully connected layer, Dropout layer, and Softmax layer. In the convolutional layer, multiple filters (convolution kernels) are used to perform convolutional operations on the input data to generate feature maps; the pooling layer is responsible for reducing the dimension of the feature maps, reducing the computational complexity and preventing overfitting; the fully connected layer is used to perform weighted summation on the input features; the Dropout layer is used to randomly discard some neurons to reduce the overfitting of the model to the training data and improve the generalization ability of the model; the Softmax layer is used to convert the input values into a probability distribution and output the classification results. Taking continuous abnormal transaction detection as an example, the CNN takes the customer's historical transaction records as input, and after multiple convolutional and pooling processes, outputs the prediction results (transaction detection results) regarding abnormal transactions.
[0063] S13: Use a new training pair to fine-tune the detection model. The original classification head (other classification heads outside the extended classification head) of the detection model remains unchanged, and only backpropagation is performed according to the extended classification head, and the model parameters are adjusted.
[0064] Through the active forgetting fine-tuning training in step S13, the decision space of the data samples to be forgotten that need to be exited will be mapped to a specified new label space. It should be noted that the decision space corresponding to the pseudo-label is far from the decision space of the existing normal customer data, that is, the data points represented by the pseudo-label have a large position difference from the normal transaction data points in the feature space and belong to different categories.
[0065] S14: Prune the extended classification head of the detection model to obtain the first optimized detection model.
[0066] By pruning the extended classification head to deactivate it, the purpose of the model's failure to predict the data samples to be exited is achieved. The failure to predict is recognized as the model no longer remembering the information of the data samples to be exited, ensuring that the model can actively forget these unnecessary data while ensuring the ability to detect the remaining user transaction information (normal transaction samples, i.e., other non-exited data samples).
[0067] Through the above steps S11 - S14, the model training under the condition of dynamic exit of data samples is completed, and this task is defined as task B. Specifically, in this example, the decision boundary migration method is specifically adopted to effectively erase the model memory of the data samples to be exited. In this process, the data samples to be exited are first processed into a pseudo-class, and then the model is fine-tuned. The decision space is moved to a position different from the original normal detection decision boundary of the model, and the decision space corresponding to this pseudo-class is deactivated. In this way, the contribution of the specified data (data samples to be exited) in the model can be effectively erased without affecting the model's judgment of normal transaction samples, ensuring continuous and effective anomaly detection.
[0068] S15: Use the first optimized detection model to perform the abnormal transaction detection task and output the transaction detection result.
[0069] For the case of dynamic exit of data samples, based on the training of the detection model in the above steps S11 - S14, the trained optimized detection model is used to continuously detect abnormal transactions in the input data samples, and then the (abnormal) transaction detection result is output, such as repeated transactions, market manipulation, etc.
[0070] In an example, before using the first optimized detection model to perform the abnormal transaction detection task, it further includes:
[0071] Perform a linear operation on the parameters of the first optimized detection model to obtain the second optimized detection model with updated parameters;
[0072] Use the second optimized detection model to perform the abnormal transaction detection task and output the transaction detection result.
[0073] After fine-tuning the detection model, it may cause a change in the optimal parameter space of the model for other non-exited data samples, thereby reducing the model performance. Therefore, the present invention performs a linear operation on the parameters of the model after fine-tuning training to ensure the performance of the model for other non-exited data samples and the effective deletion of the data samples to be exited to the greatest extent possible.
[0074] In an example, use new training for The forgetting process of fine-tuning the detection model can be described by the following expression:
[0075] ;
[0076] where, represents the parameters of the detection model after fine-tuning training; represents the parameters of the detection model before fine-tuning training; represents the parameters of the model used for optimization during the fine-tuning training process; the parameters of the model include weights, biases, etc. represents the cross-entropy loss function; represents the feature vector of the data sample to be withdrawn; represents the pseudo-label; represents the data label; represents the set of data samples to be withdrawn. Through the above fine-tuning training process, the contributions of all data samples to be withdrawn can be effectively removed from the model, so that the model no longer remembers how to correctly predict this part of the data samples, achieving the purpose of withdrawing user information protection.
[0077] In one example, fine-tuning the detection model using a new training and pruning the extended classification head of the detection model constitute the data forgetting process of the detection model. The constraint conditions of the data forgetting process are:
[0078] ;
[0079] where, represents the parameters of the detection model after fine-tuning training; represents the feature vector of the data sample to be withdrawn; represents the data label; represents the predicted output of the detection model for the input data after fine-tuning training and pruning the extended classification head. It should be noted that the forgetting process is only applied to the set of customer samples to be withdrawn , and does not change the prediction layer of the model for the original abnormal types.
[0080] In one example, a method for continuously detecting abnormal transactions for dynamic data changes further includes steps for detecting abnormal transactions with dynamically increasing data, including:
[0081] S01: Training the detection model for an old task using old data samples.
[0082] Among them, the old data samples are the collected customer data samples. At this time, there is no customer data withdrawal or addition, and they are the historical transaction records of each customer (transaction records of old customers), which can be living consumption records, bank loan records, securities trading records, and other transaction information records, etc. Specifically, they include the transaction time, transaction location, transaction amount, transaction frequency, transaction account, etc. of each user's various financial payment tools. Here, the old task training is to input the old data samples into the detection model, calculate the model output, use the loss function to calculate the difference between the model prediction value and the true label, and update the model parameters through the backpropagation algorithm to optimize the model performance until convergence.
[0083] S02: Adopt new data samples, and perform new task training on the detection model that has completed the old task training based on the regularization method or the module optimization method or the sample replay method, and then update the parameters of the detection model to obtain the third optimized detection model.
[0084] Among them, the new data samples are the newly added data samples, that is, the newly added customer (user) data samples. At this time, there is an increase in new customer data. The new data samples are the historical transaction records of each customer, which can be living consumption records, bank loan records, securities trading records, and other transaction information records, etc. Specifically, they include the transaction time, transaction location, transaction amount, transaction frequency, transaction account, etc. of each user's various financial payment tools.
[0085] Furthermore, the regularization method, the module optimization method, and the sample replay method are all continual learning methods. Continual Learning (CL) refers to training a model on a data stream of a sequence of tasks, with the training objective being that the trained model can achieve good performance on all the tasks it has learned. Each task has its own separate training set, validation set, and test set. When training the model, it can only access the data of the current training task. The main challenge is to avoid catastrophic forgetting during learning, which means that as new tasks or domains are added, the performance of previously learned tasks or domains should not (significantly) degrade over time. At this time, if the model is directly trained on the newly added data without taking constraints, the existing model will train a new parameter space on the newly added data samples to ensure the anomaly detection performance on the newly added data samples. However, due to the change in the parameter space, the optimal parameter space learned on the previous old data samples will be destroyed, and thus the model detection performance on the historical data will also drop sharply, resulting in catastrophic forgetting of the old knowledge in the model. To solve this technical problem, the present invention uses the above three continual learning methods to train the detection model that has completed the old task training for the new task. Among them, the regularization method restricts the change of the model parameters by adding a regularization term to the loss function, thereby avoiding forgetting the old tasks; the module optimization method reduces the interference between tasks by constructing task-specific parameters or dynamically adjusting the model structure, and thus avoids catastrophic forgetting; the sample replay method alleviates catastrophic forgetting by storing or generating the data of the old tasks. In this way, through the above continual learning methods, during the dynamic update process of the detection model, it can not only avoid "catastrophic forgetting" of the original detection ability of the old tasks, but also effectively adapt to the environmental changes caused by the customer flow, that is, it can continuously learn the information of the newly added data samples and maintain the monitoring ability of the latest transaction risks.
[0086] S03: Use the third optimized detection model to perform the abnormal transaction detection task and output the transaction detection result.
[0087] In this example, the regularization method, the module optimization method, and the sample replay method are introduced to train the detection model for the new task, so as to update the model parameters, achieve the optimization of the model, and thus during the dynamic update process of the model, it can not only avoid "catastrophic forgetting" of the original detection ability, but also effectively adapt to the environmental changes caused by the dynamic change of the data, which better meets the requirements of the realistic continual abnormal transaction detection.
[0088] In one example, when training the detection model based on the continual learning method of the regularization method, it includes:
[0089] During the old task training process, calculate the Fisher information matrix of each parameter of the detection model.
[0090] The Fisher information matrix is a metric for measuring information content. Assuming that according to the distribution of a random variable, the parameter used for modeling is , then the Fisher information represents the amount of information carried by the random variable about . Therefore, when fixing the value, with the random variable as the independent variable, the Fisher information will indicate how much information this random variable value can contribute to . The larger the value of the Fisher information, the greater the amount of information the random variable has about . Conversely, it means that the amount of information the random variable has about is smaller. Starting from this intuitive definition, the Elastic Weight Consolidation (EWC) mechanism uses the Fisher information matrix to calculate important parameters at the end of model training.
[0091] During the training process of a new task, the EWC mechanism is adopted to introduce a regularization term in the loss function to constrain the important parameters of the detection model in the old task, thereby updating the parameters of the detection model.
[0092] Specifically, in the regularization method, the EWC mechanism is widely used because it can effectively constrain new tasks by remembering important parameter information. As Figure 4 shows, the important parameters corresponding to different tasks are also different, and EWC can remember the important parameters of the old task and constrain the training of the new task on the premise of unchanged memory overhead. Therefore, the present invention preferably adopts EWC as the constraint method for continuous modeling of the detection model.
[0093] Preferably, considering the task A of customer addition (dynamic data increase) as a dual-task process, it is specifically divided into the old task A1 and the new task A2. A convolutional neural network is used to train this process to optimize the detection of customer transactions. Before starting the training, for the transaction records of old customers, the existing transaction data is collected and data cleaning means are used to improve the generalization ability of the model.
[0094] In the construction of the CNN model, the network usually consists of multiple convolutional layers and pooling layers, and these layers are responsible for extracting features from the input data. For task A1, when training, the transaction records of old customers are used as the input of the model, and the forward propagation formula of the CNN can be expressed as:
[0095] ;
[0096] where, is the output of the network; is the activation function (such as ReLU); is the weight matrix; is the feature of the input data, i.e., the transaction records of multiple old customers; is the bias term. In the process of updating the weights through the backpropagation algorithm, the cross-entropy loss is used to evaluate the model performance, and the loss function can be expressed as:
[0097] ;
[0098] where, represents the loss function for directly training the old task A1; represents the parameters of the model; represents the total number of samples; represents the true label of the model; represents the predicted label of the model.
[0099] After completing the training of task A1, the training of the new task A2 will be entered. During the training process of the new task A2, the parameters of the detection model need to be fine-tuned to adapt to the characteristics of the newly added data samples. The present invention introduces an elastic weight consolidation mechanism, aiming to ensure the constraint of the important parameters of the old task A1 during the training of the new task A2 by remembering the important parameters of the old task A1, and avoid forgetting the old knowledge.
[0100] Specifically, the present invention uses the Fisher information matrix to measure the importance of each parameter in the old task A1, and combines the L2 penalty as the penalty term. When using EWC for training the new customer task A2, the loss function can be expressed as:
[0101] ;
[0102] where, where, represents the loss function for training the new task by introducing the elastic weight consolidation mechanism; represents the loss function for directly training the old task without introducing the elastic weight consolidation mechanism; is the regularization coefficient; represents the Fisher information matrix; represents the data label; represents the network parameters when training the new task; represents the optimal parameters obtained after completing the old task. In this way, the CNN model can still maintain the ability to recognize the old customer patterns when learning the new transaction features, thereby effectively improving the accuracy and stability of the system. For the sake of clearer subsequent technical description, after task A is completed, the parameters of the current model are denoted as .
[0103] In an example, before performing the abnormal transaction detection task using the first optimized detection model or the third optimized detection model, it further includes:
[0104] Perform a linear operation on the parameters of the first optimized detection model and the parameters of the third optimized detection model to obtain the final detection model; use the final detection model to perform the abnormal transaction detection task and output the transaction detection result.
[0105] Specifically, the calculation expressions for the parameters of the first optimized detection model and the parameters of the third optimized detection model are:
[0106] ;
[0107] where is the optimized model parameter under the condition of dynamic data increase, that is, the parameter of the third optimized detection model; is the optimized model parameter under the condition of dynamic data exit, that is, the parameter of the first optimized detection model; is an integer, and the parameter ratio before and after the active forgetting of the detection model can be adjusted according to the actual situation and evaluation indicators. When is larger, the model pays more attention to the abnormal transaction detection performance of other non-exited transaction samples. When is smaller, the model pays more attention to the deletion cleanliness of the exited data samples in the model, so as to improve the flexibility and adjustability of the detection model.
[0108] Through the above abnormal transaction detection steps of dynamic data exit and linear operation steps, it can be ensured that during the customer exit process of abnormal transaction detection, the model can not only successfully remove the memory of the data samples to be exited, but also continuously maintain high sensitivity and accuracy to the remaining transaction data.
[0109] It should be noted that regarding customer addition (dynamic data increase) and customer exit (dynamic data exit) as two tasks that the detection model needs to train, they are respectively defined as task A (steps S01 - S02) and task B (steps S11 - S14), and task A or task B can be executed separately according to specific scenario requirements. In the actual scenario, as Figure 5 shown, customer addition and customer exit continue, so it is necessary for the detection model to perform real-time detection for new user data addition and user request to exit, that is, at this time, it is preferred to execute task A and task B in parallel. The method includes the following steps:
[0110] S100: Read the historical transaction records in the private dataset provided by the new customer and submit them to the global database;
[0111] S200: Based on the globally updated database at the current moment, perform model training for abnormal transaction detection to obtain the CNN model at the latest moment;
[0112] S300: While the CNN model training is completed, calculate the Fisher information matrix of the model parameters, which is used to characterize the different importance of the model parameters;
[0113] S400: When new period user data increases, design a new training loss function based on the Fisher information matrix to update the previous CNN model with the new data samples. At this time, the model parameters are denoted as .
[0114] S500: When a historical user submits an exit application, first expand the decision space of the CNN model (away from the normal decision space), and then fine-tune the model using the data of these exiting users, and directionally transfer the data contribution to the additional decision space;
[0115] S600: After the CNN model is fine-tuned, perform a pruning process on the model, that is, prune the additional decision space part of it, so that the CNN model is reduced to its original size, thereby achieving the purpose of forgetting the specified user data in the model; At this time, the parameters of the model are denoted as .
[0116] S700: Perform a linear operation on and to optimize the model parameters and obtain an object detection model to achieve the accurate deletion of exiting users and the correct retention of the remaining data;
[0117] S800: While processing the customer's application to exit, completely delete the private data of the exiting customer from the global database and build a new database;
[0118] S900: Use the object detection model to detect the input existing customer private transaction data and output the abnormal transaction detection result.
[0119] The present invention proposes a continuous abnormal transaction detection method for data dynamic increase and exit, which can take into account the abnormal transaction detection of new customer addition and old customer exit at the same time. On the one hand, the model can continuously absorb new customer information and maintain the monitoring ability of the latest transaction risks. On the other hand, it can actively forget the exiting customers and can effectively delete the samples of the exited customers and their influence on the model from the model. This improvement enables the model to avoid "catastrophic forgetting" of the original detection ability and effectively adapt to the environmental changes brought by customer flow during the dynamic update process, and can actively forget some users to meet the needs of users who want to exit. Therefore, the present invention realizes the flexible management of dynamic customer increase and decrease and the data usage security while ensuring the abnormal transaction detection accuracy, and reasonably utilizes the memory and computing power resources.
[0120] The present invention further includes a continuous abnormal transaction detection system for dynamically changing data, such as Figure 6 As shown, the system includes a terminal and a server 103. The terminal includes a first terminal 101 and a second terminal 102. The server 103 includes a storage server and a training server. The first terminal 101 and the second terminal 102 are both connected to the storage server and the training server.
[0121] In this example, the first terminal is an application exit terminal, and the second terminal is an application join terminal, which can be a desktop computer, a smart phone, a tablet computer, a laptop computer, etc., but is not limited thereto. The first terminal and the second terminal transmit basic training information to the training server through the SSH protocol, including data increase and decrease identifiers and training data sets. The training data set is a historical transaction sample after data preprocessing, including data samples to be exited, old data samples, and new data samples. The first terminal and the second terminal directly or indirectly access the storage server through a wired connection or a wireless network and input the original historical transaction records to the storage server.
[0122] Furthermore, the storage server and the training server can be independent physical servers or server clusters. The training server is used to execute the training steps of the continuous abnormal transaction detection model for data increase and data exit of the present invention, and trains the target model based on the basic information provided by the terminal. The storage server is used to save the trained detection model, which is accessible to the first terminal and the second terminal. At this time, the terminal and / or the training server and / or the storage server execute the abnormal transaction detection task using the second optimized detection model or the third optimized detection model or the final detection model, and output the transaction detection result. Optionally, the server can be equipped with a Linux operating system and GPU computing resources.
[0123] The above specific embodiments are detailed descriptions of the present invention. It cannot be determined that the specific embodiments of the present invention are only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention belongs, without departing from the concept of the present invention, several simple deductions and substitutions can still be made, which should all be regarded as belonging to the protection scope of the present invention.
Claims
1. A method for detecting continuous abnormal transactions with dynamic data changes, characterized in that: The method includes the abnormal transaction detection step of data dynamic exit: Assign a pseudo label to each data sample to be withdrawn, and form a new training pair with the corresponding data sample to be withdrawn; the data sample to be withdrawn is a historical transaction record, including the transaction time, transaction location, transaction amount, transaction frequency, and transaction account of each user's financial payment tool; Add an extended classification head to the output layer of the detection model to process pseudo labels, thereby expanding the decision space of the detection model; The new training pairs are used to fine-tune the detection model. The original classification head of the detection model remains unchanged, and only backpropagation is performed based on the extended classification head to update the parameters of the detection model. Crop the extended classification head of the detection model to obtain a first optimized detection model; The first optimized detection model is used to perform abnormal transaction detection tasks and output transaction detection results.
2. The method for detecting continuous abnormal transactions according to claim 1, characterized in that: Before the first optimized detection model is used to perform the abnormal transaction detection task, the method further includes: Performing linear operations on the parameters of the first optimized detection model to obtain a second optimized detection model with updated parameters; The second optimized detection model is used to perform abnormal transaction detection tasks and output transaction detection results.
3. The method for detecting continuous abnormal transactions according to claim 1, characterized in that: The calculation expression for fine-tuning the detection model using the new training pair is: ; in, represents the parameters of the detection model after fine-tuning training; represents the parameters of the detection model before fine-tuning training; Represents the parameters of the model used for optimization during fine-tuning training; represents the cross entropy loss function; A feature vector representing the data sample to be exited; represents a pseudo label; Indicates data label; Indicates the data sample set to be exited.
4. The method for detecting continuous abnormal transactions according to claim 1, characterized in that: The use of new training pairs to fine-tune the detection model and the trimming of the extended classification head of the detection model constitute the data forgetting process of the detection model. The constraints of the data forgetting process are: ; in, represents the parameters of the detection model after fine-tuning training; A feature vector representing the data sample to be exited; Indicates data label; It means that after fine-tuning training and extended classification head pruning, the detection model is effective for input data The predicted output.
5. The method for detecting continuous abnormal transactions according to claim 1, characterized in that: The method also includes the step of detecting abnormal transactions with dynamically increasing data: Use old data samples to train the detection model on old tasks; Using new data samples, and based on a regularization method, a module optimization method, or a sample replay method, performing new task training on the detection model that has completed the old task training, and then updating the parameters of the detection model to obtain a third optimized detection model; The third optimized detection model is used to perform abnormal transaction detection tasks and output transaction detection results.
6. The method for detecting continuous abnormal transactions according to claim 5, characterized in that: When training the detection model based on the continuous learning method in a regularized manner, it includes: During the training of the old task, the Fisher information matrix of each parameter of the detection model is calculated; During the training process of the new task, an elastic weight consolidation mechanism is used to introduce regularization terms in the loss function to constrain the important parameters of the detection model in the old task, so as to update the parameters of the detection model.
7. The method for detecting continuous abnormal transactions according to claim 6, characterized in that: During the training of a new task, the loss function is expressed as: ; in, Represents the loss function for training new tasks by introducing an elastic weight consolidation mechanism; Represents the loss function of directly training the old task without introducing the elastic weight consolidation mechanism; is the regularization coefficient; represents the Fisher information matrix; Indicates data label; Represents the network parameters when training new tasks; Represents the optimal parameters obtained by completing the old task.
8. The method for detecting continuous abnormal transactions according to claim 5, characterized in that: Before using the first optimized detection model or the third optimized detection model to perform the abnormal transaction detection task, the method further includes: Performing linear operations on the parameters of the first optimized detection model and the parameters of the third optimized detection model to obtain a final detection model; The final detection model is used to perform abnormal transaction detection tasks and output transaction detection results.
9. A continuous abnormal transaction detection system for dynamic data changes, characterized in that: The system includes interconnected terminals and servers, and the servers include storage servers and training servers; The terminal accesses the storage server and inputs the original historical transaction records to the storage server; the terminal accesses the training server and inputs the data increase and decrease identifier and the training data set to the training server, and the training data set includes the data samples to be withdrawn, the old data samples, and the new data samples; The training server is used to perform the step of training the detection model in the method for detecting continuous abnormal transactions with dynamic changes in data as described in any one of claims 1 to 8; The storage server is used to store the trained second optimized detection model, the third optimized detection model, or the final detection model; The terminal and / or the training server and / or the storage server uses the second optimized detection model or the third optimized detection model or the final detection model to perform the abnormal transaction detection task and output the transaction detection result.
Citation Information
Patent Citations
Incremental semantic segmentation method based on meta-learning and pseudo-label strategy
CN116030254A
Automatic driving target detection method based on incremental small sample learning
CN117612136A