Bank account risk identification method, system and device based on subject-level loss
By calculating the loss function at the subject level, the problem of inconsistency between model training and evaluation in bank account risk identification is solved, thereby improving the accuracy and recall of risky account identification.
Patent Information
- Application Number
- CN202511122241.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-11-21
AI Technical Summary
In existing technologies for identifying bank account risks, inconsistencies exist in model training and evaluation based on sample dimensions, causing model performance to be affected by noise labels and making it difficult to effectively identify risky accounts.
A subject-level loss function is adopted, which is calculated by calculating the predicted value of the sample with the highest prediction probability for each bank account subject. This loss function is then used for model training to ensure that the model reaches its optimum in the subject dimension and to reduce the impact of noise.
It improved the recall and accuracy of the model in identifying risky accounts, especially in predicting high-risk accounts, with a significant improvement in recall rate of 20-25%.
Smart Images

Figure CN120996921A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of financial technology, account risk control, account identification, etc., and in particular to a bank account risk identification method, system and device based on subject level loss. BACKGROUND
[0002] With the continuous development and progress of network technology and the Internet era, bank accounts are usually used as a channel for transferring funds, which has caused serious damage to the normal operation and reputation of banks. Therefore, banks should also upgrade the identification and control technology of suspicious accounts involved in fraud and gambling to reduce the risk brought by fraudulent accounts.
[0003] Various machine learning and deep learning play an important role in the risk control scene in the financial field. They build features by analyzing the historical behavior and basic information of accounts or other subjects, and then use a binary classification method to model. The model training basically uses the cross-entropy loss function of the sample dimension.
[0004] In many scenarios, a subject may correspond to multiple samples, and its abnormal behavior may only account for a small part of the historical behavior. If you want to label the abnormal samples corresponding to each subject, you need a lot of manpower and material resources. Therefore, when using machine learning modeling, the sample labels are usually confirmed based on the subject dimension (i.e. the labels of all samples of a subject are the same), and the quality of the sample labels will be lower. At this time, if the model is trained based on the sample dimension, the model will learn the noise in part of the labels, affecting the model effect. At the same time, the effect of the machine learning model obtained is usually evaluated from the perspective of the subject, while the model is based on sample convergence, so there is a natural inconsistency between model training and effect evaluation. SUMMARY
[0005] The present application aims to overcome the shortcomings of the prior art and provides a bank account risk identification method, system and device based on subject level loss.
[0006] The purpose of the present application is achieved by the following technical solutions: in the first aspect, the present application provides a bank account risk identification method based on subject level loss, which comprises the following steps: (1) Obtain bank account sample data, each bank account has multiple samples, and each sample corresponds to a label indicating whether there is a risk; (2) Build a neural network model, calculate the loss function based on the predicted value of the sample with the maximum predicted probability of each bank account subject based on the bank account sample data and the label, and train the model; (3) Based on the trained neural network model, the real-time obtained bank account is identified for risk.
[0007] Further, in step (1), after obtaining the bank account sample data, data cleaning is needed, specifically: 1) Integrate the basic information and transaction information of the bank account, and obtain the label of the bank account (whether it is a risk account) through the public security report, freezing and other data.
[0008] 2) According to expert experience, data analysis and the criminal method of risk accounts, process risk indicators to increase the data representation ability of the account risk.
[0009] Further, in step (2), when calculating the loss function, first calculate the position of the maximum value of the predicted value relative to each bank account subject }, i represents the bank account subject number, j represents the sample number corresponding to the bank account subject, if the maximum value of the predicted probability of the bank account subject corresponding to all samples is =1, otherwise =0.
[0010] Further, in step (2), the loss function is as follows: Wherein, represents the loss function of the subject level, N represents the number of bank account subjects, represents the most corresponding sample of the i-th subject; represents the bank account sample label, 1 with risk, 0 without risk; is the probability that the model predicts the sample is 1, and the value is a real number between 0 and 1.
[0011] Secondly, the present application also provides a bank account risk identification system based on subject level loss, which comprises: A data acquisition module is used to acquire bank account sample data, each bank account has multiple samples, and each sample corresponds to a label indicating whether it is risky or not; A model training module is used to build a neural network model, and according to the bank account sample data and the label, the loss function is calculated based on the predicted value of the sample with the maximum predicted probability of each bank account subject to train the model; A risk identification module is used to identify the risk of the bank account in real time based on the trained neural network model.
[0012] In a third aspect, the present application also provides a bank account risk identification device based on subject-level loss, comprising a memory and one or more processors, the memory stores executable code, and the processor executes the executable code to implement the bank account risk identification method based on subject-level loss.
[0013] In a fourth aspect, the present application also provides a computer readable storage medium, which stores a program, and the program is executed by a processor to implement the bank account risk identification method based on subject-level loss.
[0014] In a fifth aspect, the present application also provides a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the bank account risk identification method based on subject-level loss.
[0015] The present application has the beneficial effects that the present application proposes a subject-level loss function for the bank account risk identification process, which can reduce the influence of noise in the risk label on the model training process, make the model converge based on the subject account, and ensure the consistency of model training and evaluation. BRIEF DESCRIPTION OF DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiment or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.
[0017] Figure 1 It is a schematic diagram of the traditional machine learning training process.
[0018] Figure 2 It is a schematic diagram of the machine learning model training process using the subject-level loss function.
[0019] Figure 3 It is a structural diagram of the bank account risk identification device based on subject-level loss. DETAILED DESCRIPTION
[0020] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described below in combination with the drawings and implementation examples. It should be understood that the specific implementation examples described herein are only used to explain the present application, and are not used to limit the present application.
[0021] The present application provides a bank account risk identification method based on subject-level loss, which comprises the following steps: (1) Obtain sample data of bank accounts. Each bank account has multiple samples, and each sample corresponds to a label indicating whether there is a risk. (2) Construct a neural network model. Based on the bank account sample data and labels, calculate the loss function based on the predicted value of the sample with the highest prediction probability for each bank account subject and train the model. (3) Risk identification of real-time bank accounts based on the trained neural network model.
[0022] like Figure 1 As shown, the basic workflow of existing machine learning modeling solutions is as follows: data acquisition, data cleaning, defining sample labels, feature processing, model training, and model evaluation. Data cleaning specifically involves: 1) Integrate basic and transaction information of bank accounts, and obtain whether a bank account is a risk account by means of public security notices and frozen data.
[0023] 2) Based on expert experience, data analysis, and the modus operandi of risky accounts, process the acquired data into risk indicators to enhance the data's ability to represent account risks.
[0024] The model training phase uses the sample-based cross-entropy loss function or its variants, and its calculation process is as follows: Suppose we have a dataset D, where the samples are... Where i represents the subject number (i=1,2,...,N, there are N subjects in total), and j represents the sample number corresponding to the subject (j=1,2,...,...). The i-th subject corresponds to at most (samples); labeled as (Values can be 0 or 1), for each subject i, there is = =......= The data example is shown in Table 1: Table 1 During model training, the cross-entropy loss function is calculated using the following formula: in For model prediction The probability of being 1 is a real number between 0 and 1.
[0025] like Figure 2 As shown, the present invention proposes a bank account risk identification method based on subject-level loss, and the modeling process is the same as that of existing machine learning modeling schemes.
[0026] The difference lies in that different loss functions are used in the model training process: Compared with the traditional cross-entropy loss function, the loss function is calculated based on the predicted value of the sample with the maximum predicted probability of each subject, so that the model training can be optimized in the subject dimension, ensuring that the model training and evaluation use the same dimension, reducing the problem caused by the inconsistency of the training and evaluation dimensions; at the same time, since each subject only calculates the loss function for one sample, this can greatly reduce the influence of noise caused by the same value of the label corresponding to each subject.
[0027] In order to realize the calculation of the loss function based on the predicted value of the sample with the maximum predicted probability of each subject, the position of the maximum predicted value relative to each subject is first calculated, that is, by calculation , if is the maximum value of the predicted probability of subject i for all samples, then = 1, otherwise = 0. The step of grouping maximum value can continue the gradient back propagation, so that the subject level loss function has the possibility of practical application; at the same time, the above form adds weight to each sample on the traditional loss function, so the traditional loss function can be regarded as a special case of the subject level loss function, that is only for 1.
[0028] In addition to the above two points, since the present application is an improvement of the loss function, the present application can be applied to different scenarios and models.
[0029] The present application takes the identification of a certain bank's gambling and fraud account as an example, and compares and analyzes the performance of the original loss function and the subject level loss function in practical application. The experimental data set covers the training data from March 23 to March 24 and the test data from April 24 to June 24. The model is trained using two kinds of loss functions on the same training set, and the model performance is evaluated on the same test set. The performance of the two models is compared by comparing the recall rate of the top 1000, 5000 and 10000 accounts in the model prediction value to the gambling and fraud accounts. The experimental results are shown in Table 2. Under each evaluation threshold, the subject level loss function improves the recall rate by an average of 20%, especially in the top 1000 accounts with the highest prediction value, the recall rate is improved by 25%, which shows that the subject level loss function is significantly better than the original loss function in terms of recall rate. The subject level loss function proposed in the present application can effectively improve the accuracy of bank account gambling and fraud identification, and has important practical application value.
[0030] Table 2 Corresponding to the foregoing embodiment of the bank account risk identification method based on subject-level loss, the application also provides an embodiment of a bank account risk identification system based on subject-level loss. The system comprises a data acquisition module, a model training module, and a risk identification module. The specific implementation process of each module can refer to the specific steps of the foregoing embodiment of the bank account risk identification method based on subject-level loss. The data acquisition module is configured to acquire bank account sample data. Each bank account has multiple samples, and each sample corresponds to a label indicating whether there is a risk. The model training module is configured to construct a neural network model. According to the bank account sample data and the label, the model training is performed based on the predicted value of the sample with the maximum predicted probability of each bank account subject to calculate a loss function. The risk identification module is configured to identify the risk of a real-time acquired bank account based on the trained neural network model.
[0031] Corresponding to the foregoing embodiment of the bank account risk identification method based on subject-level loss, the application also provides an embodiment of a bank account risk identification device based on subject-level loss.
[0032] Referring to Figure 3 , the embodiment of the application provides a bank account risk identification device based on subject-level loss, which comprises a memory and one or more processors. The memory stores executable code. When the processor executes the executable code, the bank account risk identification method based on subject-level loss in the foregoing embodiment is implemented.
[0033] The embodiment of the bank account risk identification device based on subject-level loss provided by the application can be applied to any device with data processing capability, which can be a device or apparatus such as a computer. The device embodiment can be realized by software, hardware, or a combination of software and hardware. Taking software realization as an example, as a logical device, it is formed by reading the corresponding computer program instructions in the non-volatile memory into the memory and running by the processor of the device with data processing capability. From the hardware level, as shown in Figure 3 , it is a hardware structure diagram of the device with data processing capability of the bank account risk identification device based on subject-level loss provided by the application. In addition to the processor, memory, network interface, and non-volatile memory shown in Figure 3 , the device with data processing capability in the embodiment usually comprises other hardware according to the actual functions of the device with data processing capability, and details are not repeated.
[0034] The implementation process of the functions and roles of each unit in the above device is specifically described in the implementation process of the corresponding steps in the above method, which will not be repeated here.
[0035] For the device embodiment, since it basically corresponds to the method embodiment, the relevant part can be referred to the part of the method embodiment. The device embodiment described above is only illustrative, and the units described as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., they can be located in one place or distributed on multiple network units. Some or all of the modules can be selected to achieve the purpose of the present application according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0036] The embodiment of the application also provides a computer readable storage medium, which stores a program, and the program is executed by a processor to realize the bank account risk identification method based on subject level loss in the above embodiment.
[0037] The computer readable storage medium can be an internal storage unit of any data processing capable device, such as a hard disk or a memory. The computer readable storage medium can also be an external storage device of any data processing capable device, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. Further, the computer readable storage medium can include both the internal storage unit and the external storage device of any data processing capable device. The computer readable storage medium is used to store the computer program and other programs and data required by the data processing capable device, and can also be used to temporarily store data that has been output or will be output.
[0038] The application also provides a computer program product, which includes a computer program, and the computer program is executed by a processor to realize the bank account risk identification method based on subject level loss.
[0039] The above embodiments are used to explain and illustrate the application, but not to limit the application. Any modification and change made to the application within the spirit and protection scope of the claims falls within the protection scope of the application.
Claims
1. A method for identifying bank account risk based on principal-level loss, characterized in that, The method includes the following steps: (1) Obtain sample data of bank accounts. Each bank account has multiple samples, and each sample corresponds to a label indicating whether there is a risk. (2) Construct a neural network model. Based on the bank account sample data and labels, calculate the loss function based on the predicted value of the sample with the highest prediction probability for each bank account subject and train the model. (3) Risk identification of real-time bank accounts based on the trained neural network model.
2. The method for identifying bank account risk based on principal-level loss according to claim 1, characterized in that, In step (1), after obtaining the bank account sample data, data cleaning is required, specifically: 1) Integrate basic and transaction information of bank accounts, and obtain whether a bank account is a risk account by means of public security notices and frozen data. 3.2) Based on expert experience, data analysis and the modus operandi of risky accounts, process the acquired data into risk indicators to increase the data’s ability to represent account risk.
4. The method for identifying bank account risk based on principal-level loss according to claim 1, characterized in that, In step (2), when calculating the loss function, the position of the maximum predicted value relative to each bank account entity is first calculated. }, where i represents the bank account holder number and j represents the sample number corresponding to the bank account holder. If the maximum predicted probability of the bank account holder for all samples is the highest, then =1, otherwise =0.
5. The method for identifying bank account risk based on principal-level loss according to claim 3, characterized in that, In step (2), the loss function is as follows: ; in, This represents the loss function at the subject level, where N represents the number of bank account subjects. This indicates that the i-th subject corresponds to at most One sample; This represents a sample label for a bank account; a value of 1 indicates risk, while a value of 0 indicates no risk. Predict samples for the model The probability of being 1 is a real number between 0 and 1.
6. A bank account risk identification system that implements the bank account risk identification method based on subject-level loss as described in any one of claims 1-4, characterized in that, The system includes: The data acquisition module is used to acquire sample data of bank accounts. Each bank account has multiple samples, and each sample corresponds to a label indicating whether there is a risk. The model training module is used to build neural network models. Based on bank account sample data and labels, it calculates the loss function for model training by calculating the predicted value of the sample with the highest prediction probability for each bank account subject. The risk identification module is used to identify risks in real-time bank accounts based on a trained neural network model.
7. A bank account risk identification device based on subject-level loss, comprising a memory and one or more processors, wherein the memory stores executable code, characterized in that, When the processor executes the executable code, it implements a bank account risk identification method based on subject-level loss as described in any one of claims 1-4.
8. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it implements a bank account risk identification method based on subject-level loss as described in any one of claims 1-4.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements a bank account risk identification method based on subject-level loss as described in any one of claims 1-4.