Abnormal transaction behavior identification method, client and server
Through federated learning and adaptive weight adjustment mechanisms, the problem of poor abnormal transaction identification caused by single data set training has been solved, model training and optimization across financial institutions has been achieved, and the accuracy of identifying abnormal transaction behaviors has been improved.
Patent Information
- Application Number
- CN202510527705.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-09-09
AI Technical Summary
In the existing technology, the abnormal transaction behavior identification model trained based on a single data set is difficult to adapt to the transaction data and abnormal transaction behavior of different financial institutions, resulting in poor recognition effect.
A federated learning method is adopted to perform local training on the client, and the global model parameters are updated by aggregating local model parameters through the server. Combined with the cosine penalty regularization term and the adaptive weight coefficient adjustment mechanism, the model training process is optimized, information barriers are broken, and the local data and computing resources of each client are fully utilized.
It has improved the robustness and accuracy of the abnormal transaction identification model, effectively improved the identification accuracy of abnormal transaction behavior, and promoted model training and optimization among different financial institutions.
Smart Images

Figure CN120612084A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method, client, and server for identifying abnormal transaction behavior. Background Art
[0002] Abnormal trading behavior refers to behavior that deviates from normal financial transaction patterns. Related technologies typically train a neural network model using a pre-set transaction dataset to generate a trained neural network model. This trained neural network model is then deployed across various financial institutions to identify and address abnormal trading behavior. However, these technologies often train models for identifying abnormal trading behavior based on a single dataset, resulting in poor performance in this task. Summary of the Invention
[0003] The embodiments of the present application provide a method, client, and server for identifying abnormal transaction behavior, which are used to improve the accuracy of identifying abnormal transaction behavior.
[0004] On the one hand, an embodiment of the present application provides a method for identifying abnormal transaction behavior, which is applied to a client and includes:
[0005] In the current communication round, the global model parameters sent by the server are obtained and a local training operation is performed to obtain the local model parameters and feed them back to the server, so that the server updates the global model parameters based on the local model parameters, and the updated global model parameters are obtained as the new global model parameters. If a preset first condition is not met, the next communication round is skipped; otherwise, the global model parameters are sent to each client.
[0006] Based on the global model parameters, an abnormal transaction identification model is obtained;
[0007] The abnormal transaction identification model is used to identify the attribute data of the transaction behavior to be tested to obtain an abnormal transaction identification result; wherein the abnormal transaction identification result is used to indicate whether the transaction behavior to be tested is abnormal.
[0008] Furthermore, in some embodiments, the operation of obtaining the global model parameters sent by the server and performing local training to obtain the local model parameters and feed them back to the server includes: determining the current global model parameters as the local model parameters of the first training round; wherein the current global model parameters are the global model parameters sent by the server in the current communication round; in the current training round, performing the operation of local training based on the local model parameters of the current training round and the transaction data subset, and updating the local model parameters of the current training round to obtain the updated local model parameters as the new local model parameters; if the preset second condition is not met, jumping to the next training round, otherwise feeding back the local model parameters to the server.
[0009] Furthermore, in some embodiments, updating the local model parameters of the current training round includes: obtaining a cosine penalty regularization term of the current training round based on the local model parameters of the current training round, the current global model parameters, and historical global model parameters; wherein the historical global model parameters are the global model parameters sent by the server in the previous communication round; obtaining an initial loss of the current training round based on a subset of transaction data of the current training round; obtaining a target loss of the current training round based on the cosine penalty regularization term and the initial loss of the current training round; and updating the local model parameters of the current training round using the target loss of the current training round.
[0010] On the other hand, an embodiment of the present application provides a method for identifying abnormal transaction behavior, which is applied to a server and includes:
[0011] In a current communication round, sending global model parameters to at least one client participating in the current communication round, so that the at least one client participating in the current communication round obtains the global model parameters and performs a local training operation to obtain local model parameters and feeds them back to the server;
[0012] Updating the global model parameters based on the local model parameters to obtain the updated global model parameters as new global model parameters;
[0013] If the preset first condition is not met, jump to the next communication round; otherwise, send the global model parameters to each client, so that each client obtains an abnormal transaction identification model based on the global model parameters, and uses the abnormal transaction identification model to identify the attribute data of the transaction behavior to be tested to obtain an abnormal transaction identification result; wherein, the abnormal transaction identification result is used to indicate whether the transaction behavior to be tested is abnormal.
[0014] Furthermore, in some embodiments, updating the global model parameters based on the local model parameters includes: performing a full-level update or a shallow-level update on the global model parameters according to the current communication round, a preset shallow update frequency, and a preset deep update frequency, combined with the local model parameters of all the clients participating in the current communication round.
[0015] Further, in some embodiments, the global model parameters include shallow global model parameters; the local model parameters include shallow local model parameters; and according to the current communication round, the preset shallow update frequency and the preset deep update frequency, combined with the local model parameters of all the clients participating in the current communication round, the global model parameters are updated at all levels or shallow levels, including: if the target remainder is greater than a preset value and less than the deep update frequency, the global model parameters are updated based on the local model parameters of all the clients participating in the current communication round; wherein the target remainder is the remainder of the ratio of the current communication round to the preset shallow update frequency; or, if the target remainder is greater than the deep update frequency, the shallow global model parameters are updated based on the shallow local model parameters of all the clients participating in the current communication round.
[0016] Furthermore, in some embodiments, the method also includes: using a central kernel alignment method to divide the global model into a shallow part and a deep part, so that the global model parameters are divided into shallow global model parameters and deep global model parameters, and the local model parameters are divided into shallow local model parameters and deep local model parameters.
[0017] In another aspect, an embodiment of the present application provides a method for identifying abnormal transaction behavior, which is applied to a server and a client, and the method includes:
[0018] In a current communication round, sending, through the server, global model parameters to at least one of the clients participating in the current communication round;
[0019] Obtaining the global model parameters and performing local training operations through at least one of the clients participating in the current communication round, obtaining local model parameters and feeding them back to the server;
[0020] Updating the global model parameters based on the local model parameters by the server to obtain the updated global model parameters as new global model parameters;
[0021] If the preset first condition is not met, jumping to the next communication round through the server, otherwise sending the global model parameters to each of the clients through the server;
[0022] Obtaining an abnormal transaction identification model based on the global model parameters by each of the clients;
[0023] Each client uses the abnormal transaction identification model to identify the attribute data of the transaction behavior to be tested, thereby obtaining an abnormal transaction identification result; wherein the abnormal transaction identification result is used to indicate whether the transaction behavior to be tested is abnormal.
[0024] In another aspect, an embodiment of the present application provides a client, including:
[0025] a local training module, configured to obtain, in a current communication round, global model parameters sent by the server and perform local training operations, obtain local model parameters and feed them back to the server, so that the server updates the global model parameters based on the local model parameters, obtain the updated global model parameters as new global model parameters, and jump to the next communication round if a preset first condition is not met; otherwise, send the global model parameters to each client;
[0026] A model building module, configured to obtain an abnormal transaction identification model based on the global model parameters;
[0027] The abnormal transaction identification module is used to use the abnormal transaction identification model to identify the attribute data of the transaction behavior to be tested and obtain an abnormal transaction identification result; wherein the abnormal transaction identification result is used to indicate whether the transaction behavior to be tested is abnormal.
[0028] In another aspect, an embodiment of the present application provides a server, comprising:
[0029] a sending module, configured to send, in a current communication round, global model parameters to at least one client participating in the current communication round, so that the at least one client participating in the current communication round obtains the global model parameters and performs a local training operation to obtain local model parameters and feed them back to the server;
[0030] A model updating module, configured to update the global model parameters based on the local model parameters, and obtain the updated global model parameters as new global model parameters;
[0031] A logic processing module is configured to jump to the next communication round if a preset first condition is not met; otherwise, send the global model parameters to each of the clients so that each of the clients obtains an abnormal transaction identification model based on the global model parameters, and utilize the abnormal transaction identification model to identify attribute data of the transaction behavior to be tested to obtain an abnormal transaction identification result; wherein the abnormal transaction identification result is used to indicate whether the transaction behavior to be tested is abnormal.
[0032] According to an abnormal transaction behavior identification method, client and server provided in an embodiment of the present application, in a current communication round, the server sends global model parameters to at least one client participating in the current communication round; at least one client participating in the current communication round receives the global model parameters and performs a local training operation to obtain local model parameters and feeds them back to the server; the server updates the global model parameters based on the received local model parameters, and if a preset first condition is not met, it jumps to the next communication round, otherwise it sends the global model parameters to all clients; each client obtains an abnormal transaction identification model based on the global model parameters, and uses the model to identify attribute data of the transaction behavior to be tested to obtain an abnormal transaction identification result; wherein the abnormal transaction identification result is used to indicate whether the transaction behavior to be tested is abnormal, thereby breaking down the information barriers between financial institutions, enabling different financial institutions to jointly train and optimize models without leaking sensitive information, prompting model training to no longer rely on a single data set and fully utilize the local data and computing resources of each client, thereby effectively improving the robustness and accuracy of the abnormal transaction identification model, and thereby improving the identification accuracy of abnormal transaction behavior.
[0033] Other features and advantages of the present application will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present application. The purposes and other advantages of the present application can be achieved and obtained through the structures particularly pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 A schematic diagram of an implementation environment for a method for identifying abnormal transaction behavior provided in an embodiment of the present application;
[0035] Figure 2 A schematic diagram of the processing process of a method for identifying abnormal transaction behavior provided in an embodiment of the present application;
[0036] Figure 3 A flowchart of a method for identifying abnormal trading behavior provided in this application;
[0037] Figure 4 Example diagram of model parameter divergence provided for this application;
[0038] Figure 5 An example diagram of the center core alignment method provided in this application;
[0039] Figure 6 Another example diagram of the center core alignment method provided in this application;
[0040] Figure 7 A schematic diagram of the structure of a client provided for this application;
[0041] Figure 8 A schematic diagram of the structure of a server provided for this application. DETAILED DESCRIPTION
[0042] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0043] The present application is further described below in conjunction with the accompanying drawings and specific embodiments. The described embodiments should not be considered as limiting the present application. All other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0044] In the following description, reference is made to “some embodiments,” which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0045] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0046] Abnormal trading behavior refers to any deviation from normal financial trading patterns, including but not limited to abnormal trading frequency, size, or timing, trading between related accounts, self-trading, or cross-trading. Identifying and addressing abnormal trading behavior is crucial to maintaining the healthy operation of financial markets. It not only helps prevent financial fraud and market manipulation, but also protects the legitimate rights and interests of investors and maintains the fairness and transparency of financial markets.
[0047] In related technologies, a preset neural network model is usually trained using a preset transaction data set to obtain a trained neural network model. The trained neural network model is then deployed in different financial institutions to identify and process abnormal transaction behaviors. At present, the data between financial institutions are not interoperable, and the transaction data is diverse. Abnormal transaction behaviors tend to be complex. For example, the transaction data of different financial institutions vary significantly in terms of transaction scale, transaction frequency, transaction time, etc., and the abnormal transaction behaviors of different financial institutions also vary significantly. However, the transaction data set used by related technologies for model training often uses the transaction data of a certain financial institution. That is, related technologies often train a model for identifying abnormal transaction behaviors based on a single data set. This makes it difficult for related technologies to adapt to the transaction data and abnormal transaction behaviors of different financial institutions, resulting in poor results in the task of identifying abnormal transaction behaviors.
[0048] Therefore, the present application provides a method, client, and server for identifying abnormal transaction behavior, aiming to effectively improve the accuracy of identifying abnormal transaction behavior.
[0049] The specific implementation of the embodiment of the present application is described in detail below with reference to the accompanying drawings.
[0050] Reference Figure 1 , Figure 1 A schematic diagram of the implementation environment for an abnormal transaction behavior identification method provided in an embodiment of the present application. In this implementation environment, the primary hardware and software components involved include a client and a server. The client refers to a client deployed at a financial institution, and there is at least one client.
[0051] Specifically, the client may be installed with a relevant application that can be used to execute the abnormal transaction behavior identification method provided in the embodiments of this application, and the server may be a server for this application. The client and server are connected in communication. The abnormal transaction behavior identification method provided in the embodiments of this application can be executed solely on the client side, solely on the server side, or based on data exchange between the client and server.
[0052] For example, taking the abnormal transaction behavior identification method provided in the embodiment of the present application as an example, based on the data interaction between the client and the server, Figure 2 , Figure 2 A schematic diagram of the processing process of a method for identifying abnormal transaction behavior provided in an embodiment of the present application is shown as follows: Figure 2As shown, in the current communication round, the server sends global model parameters to at least one client participating in the current communication round; at least one client participating in the current communication round obtains the global model parameters and performs local training operations to obtain local model parameters and feeds them back to the server; the server updates the global model parameters based on the local model parameters fed back by all clients participating in the current communication round, and uses the updated global model parameters as new global model parameters. If the preset first condition is not met, the server jumps to the next communication round to implement iterative training, otherwise the server sends the global model parameters to all clients; each client constructs an abnormal transaction identification model based on the global model parameters, and uses the abnormal transaction identification model to identify the attribute data of the transaction behavior to be tested to obtain an abnormal transaction identification result. The abnormal transaction identification result is used to indicate whether the transaction behavior to be tested is abnormal, thereby realizing the identification and processing of abnormal transaction behavior.
[0053] Clients may include, but are not limited to, mobile phones, computers, intelligent voice interaction devices, smart home appliances, in-vehicle terminals, and aircraft. A server can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Furthermore, a server can be a node server in a blockchain network.
[0054] In addition, a communication connection can be established between the client and the server via a wireless network or a wired network. The wireless network or wired network uses standard communication technologies and / or protocols. The network can be set to the Internet or any other network, such as, but not limited to, a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network, or any combination of a virtual private network. Furthermore, the aforementioned software and hardware entities can use the same communication connection method or different communication connection methods, which is not specifically limited in this embodiment.
[0055] In combination with the above description of the implementation environment, a method for identifying abnormal transaction behavior provided in an embodiment of the present application is introduced and explained in detail.
[0056] Reference Figure 3 , Figure 3This is a flowchart of a method for identifying abnormal transaction behavior provided by this application. The method is applied to Figure 1 The client shown in FIG. 1 may include the following steps S101 - S103 , and the method applied to the client may include the following steps S101 - S103 .
[0057] S101, in the current communication round, obtain the global model parameters sent by the server and perform local training operations, obtain the local model parameters and feed them back to the server, so that the server updates the global model parameters based on the local model parameters, and obtains the updated global model parameters as the new global model parameters. If the preset first condition is not met, jump to the next communication round, otherwise send the global model parameters to each client.
[0058] It should be noted that global model parameters refer to the model parameters of the global model deployed on the server, while local model parameters refer to the model parameters of the local model deployed on the client.
[0059] It can be understood that the operation of local training refers to the operation of building a local model based on the global model parameters, and training and updating the parameters of the local model based on the local data set. The operation of local training will be explained in detail in the following embodiments.
[0060] In this step, for the rth communication round, the server sends global model parameters to the clients participating in the rth communication round. The global model parameters here refer to the global model parameters updated in the r-1th communication round and applicable to the rth communication round. The clients participating in the rth communication round receive the global model parameters and perform local training operations. After completing the local training operations, the clients participating in the rth communication round feed back the local model parameters to the server. It is understandable that only the clients selected to participate in the rth communication round can receive the global model parameters sent by the server. The server updates the global model parameters based on the received local model parameters, obtains the updated global model parameters as the new global model parameters, and then determines whether the first condition is met; if not, it means that the federated learning training has not yet been completed. At this time, the server uses the new global model parameters as the global model parameters used in the r+1th communication round and jumps to the r+1th communication round to achieve iterative training; if so, it means that the federated learning training can be completed. At this time, the server sends the global model parameters to each client to assist each client in building an abnormal transaction identification model.
[0061] Optionally, there is at least one client participating in the rth communication round.
[0062] Optionally, the first condition can be set based on actual circumstances, and this embodiment of the present application does not specifically limit this. For example, the first condition can be r=R, where R refers to the communication count threshold, which can be flexibly set based on actual circumstances. For another example, the first condition can be that the current communication duration reaches the communication duration threshold, which can also be flexibly set based on actual circumstances.
[0063] S102: Obtain an abnormal transaction identification model based on the global model parameters.
[0064] It should be noted that the abnormal transaction identification model is used to identify the attribute data of the transaction behavior to be tested and obtain the abnormal transaction identification result.
[0065] In this step, after receiving the global model parameters sent by the server, each client constructs an abnormal transaction identification model locally based on the global model parameters to implement identification and processing of abnormal transaction behaviors.
[0066] Optionally, the type of abnormal transaction identification model can be flexibly set according to actual conditions. For example, the abnormal transaction identification model can be a machine learning model such as a support vector machine, a random forest, or a neural network model such as a convolutional neural network, but is not limited thereto.
[0067] S103, using the abnormal transaction identification model to identify the attribute data of the transaction behavior to be tested, and obtaining an abnormal transaction identification result; wherein the abnormal transaction identification result is used to indicate whether the transaction behavior to be tested is abnormal.
[0068] It should be noted that the transaction behavior to be tested refers to the transaction behavior that occurs at a financial institution and is under test. The attribute data of the transaction behavior to be tested may include but is not limited to the transaction time, transaction amount, transaction type (such as transfer, payment, withdrawal, etc.), the identity information of the account initiating the transaction, the identity information of the account being traded, the balance of the account initiating the transaction, and the geographical location of the transaction. In addition, the abnormal transaction identification result is used to indicate whether the transaction behavior to be tested is abnormal.
[0069] It should be emphasized that in each specific implementation of this application, when it comes to the need to perform relevant processing based on data related to user identity or characteristics such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of such data will comply with the relevant laws, regulations, and standards of the relevant countries and regions. In addition, when the embodiment of this application needs to obtain the user's sensitive personal information, the user's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or consent, the necessary user-related data for the normal operation of the embodiment of this application will be obtained.
[0070] In this step, after each client completes the deployment of the abnormal transaction identification model, each client can input the attribute data of the transaction behavior to be tested into the abnormal transaction identification model, identify the abnormal transaction behavior through the abnormal transaction identification model, and obtain an abnormal transaction identification result indicating whether the transaction behavior to be tested is abnormal, thereby realizing the identification and processing of abnormal transaction behavior.
[0071] It can be seen that the embodiment of the present application introduces federated learning from the training process of the abnormal transaction identification model, performs local training on each client, and updates the global model parameters on the server by aggregating the local model parameters fed back by each client. The server and each client obtain the final global model parameters through multiple rounds of communication and iterative training. Each client will build an abnormal transaction identification model based on the global model parameters to realize the identification and processing of abnormal transaction behavior. In this way, the embodiment of the present application can break the information barriers between financial institutions, allowing different financial institutions to jointly train and optimize the model without leaking sensitive information, so that model training no longer relies on a single data set and fully utilizes the local data and computing resources of each client. This can effectively improve the robustness and accuracy of the abnormal transaction identification model, thereby improving the identification accuracy of abnormal transaction behavior.
[0072] The above steps will be further described below.
[0073] In some embodiments, in the above step S101, the process of obtaining the global model parameters sent by the server and performing local training operations, obtaining the local model parameters and feeding them back to the server may include the following steps S01-S02.
[0074] S01, determining the current global model parameters as the local model parameters of the first training batch.
[0075] It should be noted that the current global model parameters refer to the global model parameters sent by the server in the current communication round.
[0076] In this step, each client participating in the rth communication round receives the global model parameters sent by the server in the rth communication round, and determines the global model parameters as the local model parameters of the first training batch, thereby initializing the local model. For ease of understanding, the initialization process can be expressed as w r Refers to the global model parameters sent by the server in the rth communication round, It refers to the local model parameters of the i-th client participating in the r-th communication round in the first training batch.
[0077] S02: In the current training batch, local training is performed based on the local model parameters and transaction data subset of the current training batch, and the local model parameters of the current training batch are updated to obtain the updated local model parameters as the new local model parameters. If the preset second condition is not met, jump to the next training batch; otherwise, the local model parameters are fed back to the server.
[0078] It should be noted that the transaction data subset is a data subset obtained by partitioning the client's local transaction data set. The local transaction data set may include, but is not limited to, attribute data samples and label information for several transaction behaviors within a historical time period. The attribute data samples for transaction behaviors may include, but are not limited to, transaction time, transaction amount, transaction type (e.g., transfer, payment, withdrawal, etc.), transaction initiator account identity information, transaction target account identity information, transaction initiator account balance, transaction location, etc., while the transaction behavior label information is used to indicate whether the transaction behavior is abnormal.
[0079] It is understandable that the local model parameters fed back by the client to the server are the updated local model parameters in the last training batch.
[0080] In this step, after completing the initialization of the local model, each client participating in the rth communication round will perform local training based on the initialized local model parameters and the local transaction dataset, and iteratively update the local model parameters during the local training process. Specifically, for each client participating in the rth communication round, the local training process can be divided into multiple training batches, and accordingly, the client pre-divides the local transaction data set into transaction data subsets corresponding to multiple training batches before performing local training, and builds a local model based on the local model parameters of the first training round; for the mth training batch, the client performs local training based on the local model parameters and transaction data subset of the mth training batch, and updates the local model parameters of the mth training batch, obtaining the updated local model parameters in the mth training batch as the new local model parameters in the mth training batch, and then determines whether the second condition is met; if not, it means that the local training has not yet been completed. At this time, the client uses the new local model parameters under the mth training batch as the local model parameters used in the m+1th training batch, and jumps to the m+1th training batch to achieve iterative training; if so, it means that the local training can be completed. At this time, the client feeds back the new local model parameters under the mth training batch to the server to assist the server in updating the global model parameters, thereby achieving local training.
[0081] Optionally, the client's local model type can be flexibly configured based on actual circumstances. For example, the local model can be a machine learning model such as a support vector machine or random forest, or a neural network model such as a convolutional neural network, but is not limited thereto. However, it should be noted that the local model type must be consistent with the abnormal transaction identification model type.
[0082] Optionally, the second condition can be set according to actual conditions, and this embodiment does not specifically limit this. For example, the second condition can be that the current training duration reaches the training duration threshold, and the training duration threshold can be flexibly set according to actual conditions. For another example, the second condition can be m=M, where M represents the local training times threshold, and the local training times threshold can be flexibly set according to actual conditions. For example, the local training times threshold can be a pre-calibrated value, and for another example, the local training times threshold can be obtained based on the total number of samples of the client's local transaction data set, the capacity of the training batch, and the number of local training rounds. This ensures that local training is fully carried out and improves the accuracy of local model parameters. Specifically, the local training times threshold can be expressed as the following formula (1):
[0083]
[0084] In formula (1), M i represents the local training times threshold of the i-th client participating in the r-th communication round; |Di | represents the total number of samples in the local transaction dataset of the i-th client participating in the r-th communication round; |B| represents the capacity of each training batch, which is a preset value; E represents the number of local training rounds, which is also a preset value.
[0085] It can be seen that this embodiment can effectively improve the security of training data, reduce the risk of data leakage, and reduce data transmission costs and server training pressure by conducting local training on each client, which is conducive to model training to fully utilize the local data and computing resources of each client, thereby improving the performance of the abnormal transaction identification model and effectively improving the identification accuracy of abnormal transaction behavior.
[0086] In some implementations, in step S02 above, the implementation process of updating the local model parameters of the current training batch may include:
[0087] Obtain the cosine penalty regularization term for the current training batch based on the local model parameters of the current training batch, the current global model parameters, and the historical global model parameters; where the historical global model parameters are the global model parameters sent by the server in the previous communication round;
[0088] Get the initial loss of the current training batch based on the transaction data subset of the current training batch;
[0089] According to the cosine penalty regularization term and initial loss of the current training batch, the target loss of the current training batch is obtained;
[0090] Update the local model parameters of the current training batch using the target loss of the current training batch.
[0091] In this embodiment, although federated learning offers significant advantages for training abnormal transaction identification models, it still presents some key challenges, particularly when dealing with non-independent and identically distributed (Non-IID) data. Specifically, local transaction data varies significantly between financial institutions. This data distribution discrepancy leads to significant weight divergence between the local and global models, which not only slows model convergence but can also degrade model performance. For example, the distribution of local transaction datasets is inconsistent with the global data distribution, making it difficult for the local transaction dataset of any one financial institution's clients to reflect the data distribution of all financial institutions' clients. Another example is that the local transaction datasets of clients at different financial institutions may contain different data volumes and uneven label distributions. Furthermore, the number of clients at a financial institution may far exceed the average number of samples held by its clients. In this case, the contributions of the clients from each financial institution to model training cannot be effectively integrated, resulting in slower model convergence and degraded performance. This performance degradation stems primarily from a failure to fully account for the data characteristics and distribution differences between the clients, which severely impacts the generalization and stability of the global model. In the aforementioned non-IID data environment, existing federated learning methods often lack attention to the consistency of the model's update direction, resulting in the inability of global model parameters to effectively approach the optimal solution. Specifically, when training the abnormal transaction identification model, existing federated learning methods often adopt a simple averaging strategy when updating local model parameters, while ignoring the difference in update direction between the local model and the global model. The inconsistency of update direction will hinder the effective convergence of the model, causing the model weight differences to accumulate continuously, resulting in serious divergence of model parameters, and thus causing the global model parameters to move away from the optimal parameters. For example, referring to Figure 4 , Figure 4 It shows the divergence of model parameters when training independent and identically distributed (IID) data and non-IID data. Figure 4 It can be seen that model parameter divergence is more serious in non-IID data environments. This shows that existing federated learning methods cannot effectively solve the problem of consistency in model update direction when dealing with data heterogeneity, thus limiting their reliability and effectiveness in practical applications.
[0092] Therefore, this embodiment introduces cosine similarity as a metric in the update process of the local model to evaluate the consistency between the update direction of the global model and the update direction of the local model. In order to solve the problem of model weight divergence in a non-independent and identically distributed data environment, this embodiment reduces the weight difference between models by maximizing the cosine similarity between the update direction of the global model and the update direction of the local model. At the same time, considering that the update direction of the global model conveys different amounts of information in different training stages, in order to effectively adjust the model update weight under different data distributions, this embodiment provides an adaptive weight coefficient adjustment mechanism, which uses an adaptive weight coefficient determined by the update distance of the local model, thereby enhancing the flexibility and stability of the model update weight. Subsequently, this embodiment introduces the above-mentioned cosine similarity and the above-mentioned adaptive weight coefficient in the loss function to improve the update accuracy of the local model parameters, thereby improving the performance of the abnormal transaction identification model. Specifically, for each client participating in the rth communication round, in the mth training batch, there are:
[0093] First, the client obtains the update direction similarity of the mth training batch based on the local model parameters of the mth training batch and the global model parameters sent by the server in the r-1th communication round and the rth communication round. The update direction similarity of the mth training batch is used to indicate the cosine similarity between the update direction of the global model in the rth communication round and the update direction of the local model in the mth training batch, which satisfies the following formula (2):
[0094]
[0095] In formula (2), cosθ i (m) represents the update direction similarity of the i-th client participating in the r-th communication round in the m-th training batch; represents the local model parameters of the i-th client participating in the r-th communication round in the m-th training batch; w r represents the global model parameters sent by the server in the rth communication round; w r-1 represents the global model parameters sent by the server in the r-1th communication round; represents the update direction of the local model of the i-th client participating in the r-th communication round in the m-th training batch; w r -w r-1 represents the update direction of the global model in the rth communication round.
[0096] Optionally, to reduce the fluctuation caused by random changes in instantaneous angles in each iteration of local training, this embodiment introduces a smoothing angle in the updated direction similarity. The smoothing angle is determined by averaging the angles observed in the previous training batch. Specifically, in the first training batch, the updated direction similarity is not smoothed; in the training batches other than the first training batch, the updated direction similarity is smoothed, which can be expressed as the following formula (3):
[0097]
[0098] In formula (3), represents the smoothed angle of the i-th client participating in the r-th communication round in the m-th training batch. By performing a cosine operation on the smoothed angle, the final updated angle similarity of the i-th client participating in the r-th communication round in the m-th training batch can be obtained; represents the smoothed angle of the i-th client participating in the r-th communication round in the m-1-th training batch.
[0099] In this way, by calculating the cosine similarity between the update direction of the local model and the update direction of the global model, the update direction of the local model can be flexibly adjusted according to the characteristics of different clients, thereby realizing dynamic adjustment of the update strategy of the local model, making the update direction of the local model closer to the update direction of the global model, ensuring that the local model and the global model maintain a high degree of consistency during the model training process, and thus effectively improving the performance of the abnormal transaction identification model.
[0100] In practical applications, when the weight adjustment parameter is set to a constant, the adaptive weight coefficient will gradually decrease as the update amplitude increases. However, literature research shows that as the number of communication rounds increases, the update amplitude of the local model will also gradually decrease, which means that when fixed parameters are used, the adaptive weight coefficient is smaller in the early stage of training and larger in the later stage. This result is surprising because a large number of experimental results show that the update direction of the global model in the early stage of training provides more valuable information, while in the later stage close to convergence, its change direction contributes less to the information. Accordingly, in this embodiment, the client will obtain the adaptive weight coefficient of the mth training batch based on the local model parameters of the mth training batch and the global model parameters sent by the server in the rth communication round. Among them, the adaptive weight coefficient satisfies the following formula (4):
[0101]
[0102] In formula (4), λ i(m) represents the adaptive weight coefficient of the i-th client participating in the r-th communication round in the m-th training batch; μ represents a hyperparameter, which is a preset value. For example, the hyperparameter may be 0.05, but is not limited thereto.
[0103] Thus, after introducing the adaptive weight coefficient, the weight of the regularization term that affects the local update amplitude will depend on the update distance of the most recent local iteration. The weight associated with the model's cosine loss is expected to decrease with the increase in communication rounds. This effectively ensures that the regularization term can exert a greater influence in the early stages of training, when the update direction of the global model is more important; while in the later stages of training, when the amount of information in the global model update decreases, the influence of the regularization term is relatively reduced. Moreover, the introduction of the adaptive weight coefficient can effectively expand the range of model parameter adjustment, improve the model's adaptability to different training stages, and more effectively utilize the information of the global model, thereby helping to improve the performance of the abnormal transaction identification model.
[0104] After obtaining the adaptive weight coefficient and update direction similarity of the mth training batch, the client obtains a difference in update direction similarity relative to the mth training batch, and obtains the product of the difference and the adaptive weight coefficient of the mth training batch as the cosine penalty regularization term of the mth training batch to regulate the target loss of the mth training batch.
[0105] In addition, the client obtains the initial loss of the mth training batch based on the transaction data subset of the mth training batch, combined with a preset loss function type. Specifically, the client can use the predicted label set and label information set of the mth training batch, combined with loss function types such as the cross entropy loss function and the mean square error loss function, to calculate the initial loss of the mth training batch. The predicted label set of the mth training batch includes the prediction results of several transaction behaviors in the transaction data subset of the mth training batch, and the label information set of the mth training batch includes the label information of several transaction behaviors in the transaction data subset of the mth training batch.
[0106] Afterwards, the client obtains the sum of the initial loss of the mth training batch and the cosine penalty regularization term of the mth training batch as the target loss of the mth training batch. The target loss satisfies the following formula (5):
[0107] L i,m (x i,m ; w r , d r )=l m (x i,m )+λ i (m)(1-cosθ i (m)) (5);
[0108] In formula (5), l i,m (x i,m ;w r ,d r ) represents the target loss of the i-th client participating in the r-th communication round in the m-th training batch, d r =w r -w r-1 represents the update direction of the global model in the rth communication round, x i,m represents the transaction data subset used by the i-th client participating in the r-th communication round in the m-th training batch; l m (x i,m ) represents the initial loss of the i-th client participating in the r-th communication round in the m-th training batch; λ i (m)(1-*cosθ i (m)) represents the cosine penalty regularization term of the i-th client participating in the r-th communication round in the m-th training batch.
[0109] Finally, the client calculates the gradient of the target loss of the mth training batch and updates the local model parameters of the mth training batch using methods such as gradient descent, thereby completing the local update. The gradient of the target loss satisfies the following formula (6):
[0110]
[0111] In formula (6), represents the updated local model parameters of the i-th client participating in the r-th communication round in the m-th training batch; represents the local model parameters of the i-th client participating in the r-th communication round in the m-1-th training batch; η represents the preset learning rate, which can be flexibly set according to actual conditions.
[0112] In addition, refer to Figure 2 The present application also provides a method for identifying abnormal transaction behavior, which can be applied to Figure 1 The server shown in FIG. 1 may include the following steps S201 to S203:
[0113] S201, in a current communication round, sending global model parameters to at least one client participating in the current communication round, so that the at least one client participating in the current communication round obtains the global model parameters and performs a local training operation, obtains local model parameters and feeds them back to the server;
[0114] S202, updating the global model parameters based on the local model parameters, obtaining the updated global model parameters as new global model parameters;
[0115] S203: If the preset first condition is not met, jump to the next communication round; otherwise, send the global model parameters to each client so that each client can obtain an abnormal transaction identification model based on the global model parameters, and use the abnormal transaction identification model to identify the attribute data of the transaction behavior to be tested to obtain an abnormal transaction identification result; wherein, the abnormal transaction identification result is used to indicate whether the transaction behavior to be tested is abnormal.
[0116] It should be noted that the technical solution of the embodiment of the present application can refer to the description of steps S101 to S103. The only difference is that the execution subject of the embodiment of the present application is the server, which will not be repeated here.
[0117] In some implementations, in step S202 above, the process of updating the global model parameters based on the local model parameters may include the following steps:
[0118] According to the current communication round, the preset shallow update frequency and the preset deep update frequency, the global model parameters are updated at all levels or shallowly in combination with the local model parameters of all clients participating in the current communication round.
[0119] It should be noted that the shallow update frequency refers to the update frequency of the shallow part of the global model, while the deep update frequency refers to the update frequency of the deep part of the global model.
[0120] In this implementation, the server's global update methods can be divided into two types: the first is a full-level update method, in which the server updates both shallow and deep parameters of the global model; the second is a shallow update method, in which the server only updates shallow parameters of the global model, without updating deep parameters. Accordingly, when updating global model parameters, the server compares the current communication round, the preset shallow update frequency, and the preset deep update frequency. Based on the comparison results, the server selects a corresponding global update method. The server then updates the global model parameters using the local model parameters of all clients participating in the current communication round in combination with the selected global update method.
[0121] Thus, this implementation fully considers the functions of different layers of the global model and their impact on weight divergence. It selects an appropriate global update method based on the relationship between the current communication round, the preset shallow update frequency, and the preset deep update frequency, effectively improving the accuracy of global model parameter updates. Furthermore, the server implements global updates by aggregating local model parameters fed back by clients participating in the current communication round. This fully utilizes each client's local data and computing resources, effectively improving the robustness and accuracy of the abnormal transaction identification model, and thus improving the accuracy of identifying abnormal trading behavior.
[0122] In some embodiments, the global model parameters may include shallow global model parameters, and the local model parameters may include shallow local model parameters. The process of performing a full-level update or a shallow-level update on the global model parameters based on the current communication round, the preset shallow update frequency, and the preset deep update frequency, in combination with the local model parameters of all clients participating in the current communication round, may include the following steps:
[0123] If the target remainder is greater than a preset value and less than the deep update frequency, the global model parameters are updated based on the local model parameters of all clients participating in the current communication round; where the target remainder is the remainder of the ratio of the current communication round to the preset shallow update frequency;
[0124] Alternatively, if the target remainder is greater than or equal to the deep update frequency, the shallow global model parameters are updated based on the shallow local model parameters of all clients participating in the current communication round.
[0125] In this implementation, although federated learning offers significant advantages in training the abnormal transaction identification model, it still faces some key challenges, particularly when dealing with non-independent and identically distributed (Non-IID) data. Specifically, existing methods lack effective inter-layer similarity measurement tools, and their assessment of the similarity between different layers of the neural network is insufficient. This results in the inability to effectively evaluate and rationally utilize the role of different layers in feature extraction and information transfer in the NID data environment. The learning capabilities of each layer are not effectively integrated, making it difficult to fully reflect the contribution of each layer to model training. Consequently, the model is unable to fully utilize the functional characteristics of each layer, seriously affecting model performance and generalization capabilities.
[0126] To better understand the functions of different neural network layers and their impact on weight divergence, this implementation combines a hierarchical structure with a similarity module to implement global model parameter updates. It should be understood that the neural network here refers to the global model deployed on the server and the local model deployed on the client.
[0127] Specifically, the neural network can be divided into a shallow layer and a deep layer. For the shallow layer, since the features it extracts are more universal and less affected by non-IID data, the shallow layer can be updated at the same frequency as in traditional federated learning methods. Conversely, for the deep layer, since its features are more specific and more susceptible to non-IID data, this embodiment reduces the deep layer's update frequency. In other words, the deep layer's update frequency is lower than the shallow layer's.
[0128] Optionally, both the shallow update frequency and the deep update frequency can be flexibly calibrated according to actual conditions. This embodiment does not specifically limit this. However, it should be noted that the deep update frequency and the shallow update frequency must be greater than zero. For example, the frequency ratio shown in the following formula (7) is used to measure the shallow update frequency and the deep update frequency:
[0129] freq=f d / f s (7);
[0130] In formula (7), freq represents the frequency ratio; f d Indicates the frequency of deep update; f s Indicates the shallow update frequency. For example, if the deep update frequency is set to 5 and the shallow update frequency is set to 15, then f req =5 / 15, which means that in every 15 communication rounds, the parameters of the deep part of the global model are only uploaded and downloaded in the last 5 communication rounds.
[0131] Based on this, in this implementation, the server collects gradient updates, i.e., local model parameters, from each client participating in the current communication round. Then, based on the deep update frequency and the shallow update frequency, it determines whether the level of the global model participating in the aggregation in the current communication round is limited to the shallow part or all layers including the shallow part and the deep part. Finally, the global model parameters applicable to the next communication round are calculated by weighted averaging the local model parameters of all clients participating in the current communication round.
[0132] Specifically, for the rth communication round, the server first obtains the remainder of the ratio of r to the shallow update frequency as the target remainder; then, the server determines whether the target remainder is greater than a preset value and less than the deep update frequency. It is understood that the preset value is less than the deep update frequency.
[0133] If so, the update mode for all levels is determined as the global update mode, that is, based on the local model parameters (including shallow local model parameters and deep local model parameters) of all clients participating in the r-th communication round, the global model parameters (including shallow global model parameters and deep global model parameters) are updated, and the updated global model parameters are determined as the new global model parameters for the r+1-th communication round. The updated global model parameters satisfy the following formula (8):
[0134]
[0135] In formula (8), w r+1 represents the global model parameters updated in the rth communication round, which is used in the r+1th communication round; represents the total number of samples in the local transaction dataset of all clients participating in the r-th communication round; represents the local model parameters of the i-th client participating in the r-th communication round; i∈S r Indicates that the i-th client is selected to participate in the r-th communication round.
[0136] If not, the shallow update mode is determined as the global update mode, that is, the shallow global model parameters of the global model parameters are updated based on the shallow local model parameters of all clients participating in the r-th communication round, and the updated shallow global model parameters and the unupdated deep global model parameters are determined as the new global model parameters for the r+1-th communication round.
[0137] Among them, the updated shallow global model parameters satisfy the following formula (9):
[0138]
[0139] In formula (9), represents the shallow global model parameters updated in the rth communication round; represents the shallow local model parameters of the i-th client participating in the r-th communication round.
[0140] Optionally, the preset value can be set according to actual conditions, which is not specifically limited in this embodiment. For example, the preset value can be zero, but is not limited thereto.
[0141] It can be seen that this embodiment can implement differentiated weight update strategies based on the characteristics of different layers by evaluating the similarities between different layers of the neural network, realize differentiated processing of hierarchical features, and further optimize model performance. This approach not only deepens the understanding of the internal structure of the neural network, but also promotes the improvement of model performance in non-independent and identically distributed data environments. In resource-constrained scenarios, optimization of different layers helps to improve the adaptability and efficiency of the global model, thereby enabling the abnormal transaction identification model to perform better in practical applications.
[0142] In some embodiments, the above method further comprises:
[0143] The global model is divided into a shallow part and a deep part using a central kernel alignment method, so that the global model parameters are divided into shallow global model parameters and deep global model parameters, and the local model parameters are divided into shallow local model parameters and deep local model parameters.
[0144] In this embodiment, it is proposed to use the centered kernel alignment (CKA) method to evaluate the similarity of different layers of the neural network in order to deeply understand the functional characteristics of different layers. It can be understood that the centered kernel alignment method belongs to the prior art. Simply put, in the centered kernel alignment method, for a given representation matrix x and representation matrix y, the representation matrix x can be understood as the feature representation of the current layer of the neural network, and the representation matrix y can be understood as the feature representation of the previous layer or the next layer of the current layer of the neural network. First, the two Gram matrices of the two representation matrices are calculated, and then the two Gram matrices are centered. After that, the two centered Gram matrices are processed using vectorization operations, and the Hilbert-Schmidt independence criterion value is calculated. Finally, the similarity between the two representation matrices is calculated by standardizing the Hilbert-Schmidt independence criterion value to obtain the centered kernel alignment value of the current layer of the neural network. The centered kernel alignment value can measure the similarity between the current layer and the previous layer. The centered kernel alignment values of all layers of the neural network can be obtained by the above method. Due to different data distributions, the center kernel alignment value of the shallow part of the neural network is often higher than that of the deep part. Therefore, for the current layer of the neural network, if the center kernel alignment value of the current layer is greater than the preset layer threshold, the current layer is determined to be a shallow layer, otherwise the current layer is determined to be a deep layer. In this way, the neural network can be divided into shallow and deep parts, thereby accurately evaluating the similarity between different layers of the neural network, which is conducive to improving the update accuracy of the global model parameters.
[0145] In this embodiment, the global model deployed on the server and the local model deployed on each client both follow the above-mentioned shallow and deep division method, that is, the global model deployed on the server is divided into shallow and deep parts through the above-mentioned method, so that the global model parameters are correspondingly divided into shallow global model parameters and deep global model parameters. The server or client can divide the shallow and deep parts of the local model according to the division of the global model, so that the local model parameters are also divided into shallow local model parameters and deep local model parameters.
[0146] For example, Figure 5 As shown, Figure 5 This is an example diagram of the center kernel alignment method provided in this application. By evaluating the similarity between different layers in the SimpleCNN model using the center kernel alignment method, the following can be obtained: Figure 5 The heat map shown in Figure 2 shows the SimpleCNN model divided into shallow and deep layers.
[0147] For example, Figure 6 As shown, Figure 6 This is another example diagram of the center kernel alignment method provided in this application. The center kernel alignment method can be used to evaluate the similarity between different layers in the ResNet34 model, and the following can be obtained: Figure 6 The heat map shown in the figure can be used to divide the ResNet34 model into shallow and deep parts.
[0148] In addition, the present invention also provides a method for identifying abnormal transaction behavior, which can be applied to Figure 1 The server and client shown in FIG. 4 may include the following steps S301 to S306:
[0149] S301, in a current communication round, sending global model parameters to at least one client participating in the current communication round through a server;
[0150] S302, obtaining global model parameters and performing local training operations through at least one client participating in the current communication round, obtaining local model parameters and feeding them back to the server;
[0151] S303, updating the global model parameters based on the local model parameters through the server, and obtaining the updated global model parameters as new global model parameters;
[0152] S304: If the preset first condition is not met, jump to the next communication round through the server; otherwise, send the global model parameters to each client through the server;
[0153] S305, obtaining an abnormal transaction identification model based on the global model parameters by each client;
[0154] S306 , each client uses an abnormal transaction identification model to identify the attribute data of the transaction behavior to be tested, and obtains an abnormal transaction identification result; wherein the abnormal transaction identification result is used to indicate whether the transaction behavior to be tested is abnormal.
[0155] It should be noted that the technical solution of the embodiment of the present application can refer to the description of steps S101 to S103 and the description of steps S201 to S203. The only difference is that the execution subjects of the embodiment of the present application are the server and the client, which will not be repeated here.
[0156] To better illustrate the technical solution of this embodiment, a specific example is presented below. This example is applied to several financial clients and a central server. Financial clients refer to clients deployed in financial institutions, such as computers, tablets, etc. A financial institution can deploy one or more financial clients, but is not limited to this. The central server divides the global model into shallow and deep parts based on the central core alignment method. At the same time, the central server or each financial client divides the local model into shallow and deep parts based on the division of the global model. In addition, the central server sets the deep update frequency and the shallow update frequency.
[0157] In this example, the specific process for identifying and handling abnormal transaction behavior is as follows:
[0158] S401: In the rth communication round, the central server selects all financial clients or randomly selects some financial clients as clients participating in the rth communication round, and sends global model parameters to the selected financial clients.
[0159] S402: Each financial client receives the global model parameters sent by the central server and performs local training operations, and then feeds back the local model parameters to the server when all training batches are completed.
[0160] Specifically, first, the financial client determines the global model parameters as the local model parameters of the first training batch; then, for the mth training batch, is the local training times threshold. The financial client performs local training based on the local model parameters and transaction data subset of the mth training batch, and updates the local model parameters of the mth training batch. The updated local model parameters in the mth training batch are obtained as the new local model parameters in the mth training batch. Then, it is determined whether m is equal to If so, the financial client will feed back the new local model parameters under the mth training batch to the server. Otherwise, the financial client will use the new local model parameters under the mth training batch as the local model parameters used in the m+1th training batch and jump to the m+1th training batch to achieve iterative training.
[0161] The transaction data subset is a data subset obtained by dividing the client's local transaction data set. The local transaction data set may include, but is not limited to, attribute data samples and label information of several transaction behaviors within a historical time period. The attribute data samples of transaction behaviors may include, but are not limited to, the transaction time, transaction amount, transaction type (such as transfer, payment, withdrawal, etc.), the identity information of the account initiating the transaction, the identity information of the account being traded, the balance of the account initiating the transaction, the geographical location of the transaction, etc., while the label information of the transaction behavior is used to indicate whether the transaction behavior is abnormal. In addition, the update process of the local model parameters follows the above formulas (2)-(6).
[0162] S403: The central server selects a global update method, and updates the global model parameters based on the global update method and the local model parameters of all clients participating in the rth communication round.
[0163] Specifically, the central server determines whether the remainder of the ratio of r to the shallow update frequency is greater than zero and less than the deep update frequency. If so, the central server updates the global model parameters (including shallow global model parameters and deep global model parameters) based on the local model parameters (including shallow local model parameters and deep local model parameters) of all clients participating in the r-th communication round, and determines the updated global model parameters as new global model parameters for the r+1-th communication round, which follows the above formula (8). If not, the central server updates the shallow global model parameters of the global model parameters based on the shallow local model parameters of all clients participating in the r-th communication round, and determines the updated shallow global model parameters and the unupdated deep global model parameters as new global model parameters for the r+1-th communication round, which follows the above formula (9).
[0164] S404, the central server determines whether r=R, where R refers to the communication times threshold; if not, proceeds to step S405; if so, proceeds to step S406.
[0165] S405: The central server jumps to the r+1th communication round, ie, sets r=r+1 and returns to step S401.
[0166] S406: The central server completes model training and sends global model parameters to all financial clients.
[0167] S407: One or more financial clients deploy an abnormal transaction identification model locally based on the global model parameters.
[0168] At step S408, one or more financial clients input the attribute data of the transaction to be tested into the abnormal transaction identification model. The abnormal transaction identification model then identifies the abnormal transaction and obtains an abnormal transaction identification result indicating whether the transaction to be tested is abnormal, thereby achieving identification and processing of the abnormal transaction. The attribute data of the transaction to be tested may include, but is not limited to, the transaction time, transaction amount, transaction type (e.g., transfer, payment, withdrawal, etc.), the account identity information of the initiating transaction, the account identity information of the transaction being initiated, the balance of the account initiating the transaction, and the geographic location of the transaction.
[0169] Corresponding to the above method embodiment, the present application also provides an embodiment of a server, Figure 7 A schematic diagram of the structure of a client according to an embodiment of the present application is shown in FIG. Figure 7 As shown, the client may include:
[0170] The local training module 501 is configured to obtain the global model parameters sent by the server and perform local training operations in the current communication round, obtain the local model parameters and feed them back to the server so that the server updates the global model parameters based on the local model parameters, and obtain the updated global model parameters as the new global model parameters. If the preset first condition is not met, the process jumps to the next communication round; otherwise, the global model parameters are sent to each client.
[0171] A model building module 502 is used to obtain an abnormal transaction identification model based on global model parameters;
[0172] The abnormal transaction identification module 503 is used to identify the attribute data of the transaction behavior to be tested using the abnormal transaction identification model to obtain an abnormal transaction identification result; wherein the abnormal transaction identification result is used to indicate whether the transaction behavior to be tested is abnormal.
[0173] The above is a schematic diagram of a client solution of this embodiment. It should be noted that the technical solution of this client solution is based on the same concept as the technical solution of the method for identifying abnormal transaction behavior applied to the client described above. For details not described in detail in the technical solution of the client solution, please refer to the description of the technical solution of the method for identifying abnormal transaction behavior applied to the client described above.
[0174] Corresponding to the above method embodiment, the present application also provides an embodiment of a server, Figure 8 A schematic diagram showing the structure of a server according to an embodiment of the present application is shown in FIG. Figure 8 As shown, the server may include:
[0175] A sending module 601 is configured to send global model parameters to at least one client participating in the current communication round in the current communication round, so that the at least one client participating in the current communication round obtains the global model parameters and performs a local training operation to obtain local model parameters and feed them back to the server;
[0176] A model updating module 602 is configured to update the global model parameters based on the local model parameters, and obtain the updated global model parameters as new global model parameters;
[0177] Logic processing module 603 is used to jump to the next communication round if the preset first condition is not met; otherwise, it sends the global model parameters to each client so that each client can obtain an abnormal transaction identification model based on the global model parameters, and use the abnormal transaction identification model to identify the attribute data of the transaction behavior to be tested to obtain an abnormal transaction identification result; wherein the abnormal transaction identification result is used to indicate whether the transaction behavior to be tested is abnormal.
[0178] The above is a schematic diagram of a server solution for this embodiment. It should be noted that the technical solution of this server is based on the same concept as the technical solution of the aforementioned method for identifying abnormal transaction behavior applied to a server. For details not described in detail in the technical solution of the server, please refer to the description of the technical solution of the aforementioned method for identifying abnormal transaction behavior applied to a server.
[0179] In summary, while protecting data privacy, this application utilizes private datasets at each financial institution for model training. Each financial institution can dynamically adjust the aggregation interval based on its own data volume and training stage, and conduct gradient transmission and model updates with the central server node. The application focuses on using global updates to guide local updates, improving the consistency of model update direction. Furthermore, the application can determine the update frequency of different layers based on their functional characteristics, thereby improving the generalization capability of the abnormal transaction identification model and reducing communication losses. Accordingly, this application has at least one of the following technical effects:
[0180] First, this application significantly reduces model weight divergence caused by non-independent and identically distributed data by maximizing the cosine similarity between the global model update direction and the local model update direction. This consistency is particularly important in resource-constrained scenarios, where computing and communication resources are limited and inconsistency in update direction may lead to more severe performance degradation. This application effectively suppresses directional inconsistency during local model updates, thereby improving the overall consistency of the model, so that the final global model performs better than traditional federated learning algorithms on multiple datasets, meeting the requirements of efficient distributed training.
[0181] Secondly, this application eliminates the need for additional hyperparameter selection by dynamically adjusting adaptive weights. It also effectively leverages global model information at different training stages, significantly reducing communication overhead during training while maintaining model performance. This significantly outperforms other federated learning algorithms. This is crucial for achieving efficient distributed training in resource-constrained scenarios.
[0182] Third, this application uses central kernel alignment to evaluate the similarity between different layers, enabling differentiated weight update strategies tailored to the characteristics of different layers, thereby further optimizing model performance. This approach not only deepens understanding of the internal structure of neural networks but also promotes performance improvements in non-independent and identically distributed data environments. In resource-constrained scenarios, optimizing different layers helps improve the adaptability and efficiency of the model, enabling it to perform better in practical applications.
[0183] In summary, by training an abnormal transaction identification model and deploying it on the client side of various financial institutions through the embodiments of this application, the performance of the abnormal transaction identification model can be effectively improved, thereby effectively increasing the accuracy of identifying abnormal transaction behavior, assisting financial institutions in effectively identifying potential abnormal transaction behavior, improving financial risk control capabilities, and reducing fraud risks. At the same time, the privacy of user data can be protected, avoiding the direct sharing of raw data during model training, and improving data security during model training.
[0184] In some optional embodiments, the function / operation mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the function / operation involved, the two boxes shown in succession can actually be executed substantially simultaneously or the boxes can sometimes be executed in reverse order. In addition, the embodiments presented and described in the flow chart of the present application are provided in an exemplary manner for the purpose of providing a more comprehensive understanding of the technology. The disclosed method is not limited to the operations and logical flows presented herein. Optional embodiments are contemplated in which the order of the various operations is changed and the sub-operations described as a part of a larger operation are performed independently.
[0185] In addition, although the present application is described in the context of functional modules, it should be understood that, unless otherwise stated, one or more of the functions and / or features may be integrated into a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It is also understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present application. More specifically, given the properties, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the module will be understood within the routine skills of an engineer. Therefore, a person skilled in the art can implement the present application as set forth in the claims using ordinary techniques without undue experimentation. It is also understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present application, which is determined by the full scope of the appended claims and their equivalents.
[0186] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several programs for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0187] The logic and / or steps represented in a flowchart or otherwise described herein, for example, may be considered as an ordered list of executable programs for implementing the logical functions, and may be embodied in any computer-readable medium for use by, or in conjunction with, a program execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can retrieve and execute a program from a program execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" may be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, a program execution system, apparatus, or device.
[0188] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering, or processing in another suitable manner as necessary, and then stored in a computer memory.
[0189] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable program execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0190] In the above description of this specification, reference to the terms "one embodiment / example," "another embodiment / example," or "certain embodiments / examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples.
[0191] Although the embodiments of the present application have been shown and described, those skilled in the art will appreciate that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and intent of the present application, and that the scope of the present application is defined by the claims and their equivalents.
[0192] The above is a specific description of the preferred implementation of the present application, but the present application is not limited to the embodiments. Those skilled in the art may make various equivalent modifications or substitutions without violating the spirit of the present application. These equivalent modifications or substitutions are all included in the scope defined by the claims of the present application.
Claims
1. A method for identifying abnormal transaction behavior, characterized in that: Applied to a client, the method includes: In the current communication round, the global model parameters sent by the server are obtained and a local training operation is performed to obtain the local model parameters and feed them back to the server, so that the server updates the global model parameters based on the local model parameters, and the updated global model parameters are obtained as the new global model parameters. If a preset first condition is not met, the next communication round is skipped; otherwise, the global model parameters are sent to each client. Based on the global model parameters, an abnormal transaction identification model is obtained; The abnormal transaction identification model is used to identify the attribute data of the transaction behavior to be tested to obtain an abnormal transaction identification result; wherein the abnormal transaction identification result is used to indicate whether the transaction behavior to be tested is abnormal.
2. The method according to claim 1, characterized in that The operation of obtaining the global model parameters sent by the server and performing local training to obtain the local model parameters and feeding them back to the server includes: Determining current global model parameters as local model parameters for a first training round; wherein the current global model parameters are the global model parameters sent by the server in the current communication round; In the current training round, local training operations are performed based on the local model parameters and transaction data subset of the current training round, and the local model parameters of the current training round are updated to obtain the updated local model parameters as the new local model parameters. If the preset second condition is not met, jump to the next training round; otherwise, the local model parameters are fed back to the server.
3. The method according to claim 2, characterized in that The updating of the local model parameters of the current training round includes: Obtaining a cosine penalty regularization term for the current training round based on the local model parameters of the current training round, the current global model parameters, and historical global model parameters; wherein the historical global model parameters are the global model parameters sent by the server in a previous communication round; Obtaining an initial loss for the current training round based on a subset of transaction data for the current training round; Obtaining a target loss for the current training round based on the cosine penalty regularization term and the initial loss for the current training round; The local model parameters of the current training round are updated using the target loss of the current training round.
4. A method for identifying abnormal transaction behavior, characterized in that: Applied to a server, the method includes: In a current communication round, sending global model parameters to at least one client participating in the current communication round, so that the at least one client participating in the current communication round obtains the global model parameters and performs a local training operation to obtain local model parameters and feeds them back to the server; Updating the global model parameters based on the local model parameters to obtain the updated global model parameters as new global model parameters; If the preset first condition is not met, jump to the next communication round; otherwise, send the global model parameters to each client, so that each client obtains an abnormal transaction identification model based on the global model parameters, and uses the abnormal transaction identification model to identify the attribute data of the transaction behavior to be tested to obtain an abnormal transaction identification result; wherein, the abnormal transaction identification result is used to indicate whether the transaction behavior to be tested is abnormal.
5. The method according to claim 4, characterized in that The updating of the global model parameters based on the local model parameters includes: According to the current communication round, the preset shallow update frequency and the preset deep update frequency, the global model parameters are updated in all levels or shallowly in combination with the local model parameters of all the clients participating in the current communication round.
6. The method according to claim 5, characterized in that The global model parameters include shallow global model parameters; the local model parameters include shallow local model parameters; and the updating of the global model parameters in all layers or shallow layers according to the current communication round, a preset shallow update frequency, and a preset deep update frequency, in combination with the local model parameters of all the clients participating in the current communication round, comprises: If the target remainder is greater than a preset value and less than the deep update frequency, updating the global model parameters based on the local model parameters of all the clients participating in the current communication round; wherein the target remainder is the remainder of the ratio of the current communication round to the preset shallow update frequency; Alternatively, if the target remainder is greater than or equal to the deep update frequency, the shallow global model parameters are updated based on the shallow local model parameters of all the clients participating in the current communication round.
7. The method according to any one of claims 4 to 6, characterized in that The method further comprises: The global model is divided into a shallow part and a deep part using a central kernel alignment method, so that the global model parameters are divided into shallow global model parameters and deep global model parameters, and the local model parameters are divided into shallow local model parameters and deep local model parameters.
8. A method for identifying abnormal transaction behavior, characterized in that: Applied to a server and a client, the method includes: In a current communication round, sending, through the server, global model parameters to at least one of the clients participating in the current communication round; Obtaining the global model parameters and performing local training operations through at least one of the clients participating in the current communication round, obtaining local model parameters and feeding them back to the server; Updating the global model parameters based on the local model parameters by the server to obtain the updated global model parameters as new global model parameters; If the preset first condition is not met, jumping to the next communication round through the server, otherwise sending the global model parameters to each of the clients through the server; Obtaining an abnormal transaction identification model based on the global model parameters by each of the clients; Each client uses the abnormal transaction identification model to identify the attribute data of the transaction behavior to be tested, thereby obtaining an abnormal transaction identification result; wherein the abnormal transaction identification result is used to indicate whether the transaction behavior to be tested is abnormal.
9. A client, characterized in that: include: a local training module, configured to obtain, in a current communication round, global model parameters sent by the server and perform local training operations, obtain local model parameters and feed them back to the server, so that the server updates the global model parameters based on the local model parameters, obtain the updated global model parameters as new global model parameters, and jump to the next communication round if a preset first condition is not met; otherwise, send the global model parameters to each client; A model building module, configured to obtain an abnormal transaction identification model based on the global model parameters; The abnormal transaction identification module is used to use the abnormal transaction identification model to identify the attribute data of the transaction behavior to be tested and obtain an abnormal transaction identification result; wherein the abnormal transaction identification result is used to indicate whether the transaction behavior to be tested is abnormal.
10. A server, characterized in that: include: a sending module, configured to send, in a current communication round, global model parameters to at least one client participating in the current communication round, so that the at least one client participating in the current communication round obtains the global model parameters and performs a local training operation to obtain local model parameters and feed them back to the server; A model updating module, configured to update the global model parameters based on the local model parameters, and obtain the updated global model parameters as new global model parameters; A logic processing module is configured to jump to the next communication round if a preset first condition is not met; otherwise, send the global model parameters to each of the clients so that each of the clients obtains an abnormal transaction identification model based on the global model parameters, and utilize the abnormal transaction identification model to identify attribute data of the transaction behavior to be tested to obtain an abnormal transaction identification result; wherein the abnormal transaction identification result is used to indicate whether the transaction behavior to be tested is abnormal.