Bank risk detection method, device, equipment and product based on machine learning
Patent Information
- Application Number
- CN202510141372.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-05-30
Smart Images

Figure CN120070035A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and particularly relates to a bank risk detection method, device, equipment and product based on machine learning. Background Art
[0002] In recent years, the Internet finance field has risen rapidly and also stepped into a situation where risks are becoming increasingly complex. These risks cover multiple aspects of the traditional financial field, such as credit risk, market risk and operational risk, etc. Among them, Internet finance is characterized by high-frequency trading and a large amount of user data, which brings greater challenges to the identification and control of risks.
[0003] Currently, the situations faced by Internet finance risk control mainly include: the business complexity increases, the risk control elements involved are diverse, the data sources are extensive and the scale is large, resulting in low processing efficiency of traditional risk control methods; the risk control models need to be continuously updated and optimized to adapt to the rapid changes in the market; there is a time lag in risk early warning and control measures and it is difficult to respond to risk events in a timely manner, etc. Thus, these problems have prompted Internet finance enterprises to explore new risk control ways. Among them, traditional risk control means (such as risk audits of banking operations) mainly rely on manual work, with fixed rules and limited processing speed. Therefore, it is difficult to effectively cope with complex and changeable risk situations. Based on this, how to provide a bank risk detection method with high efficiency and strong adaptability has become an urgent problem to be solved. Summary of the Invention
[0004] The purpose of the present invention is to provide a bank risk detection method, device, equipment and product based on machine learning to solve the problems of low efficiency and inability to adapt to the rapid changes in the market existing in the prior art.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions:
[0006] In the first aspect, a bank risk detection method based on machine learning is provided, including:
[0007] Obtain a training data set, wherein the training data set includes historical financial credit feature data of several sample users, and the historical financial credit feature data of any sample user is constructed based on the historical financial data, historical social network data and personal data of the any sample user;
[0008] Use the historical financial credit feature data of each sample user in the training data set as input and the overdue probability of each sample user as output to train a risk detection model to obtain a trained risk detection model;
[0009] Obtain the financial data, social network data, and personal data of the target user, and use the financial data, social network data, and personal data of the target user to construct the financial credit feature data of the target user;
[0010] Input the financial credit feature data into the trained risk detection model to obtain the overdue probability of the target user;
[0011] Generate a risk warning message based on the overdue probability of the target user, and send the risk warning message to the bank monitoring terminal.
[0012] Based on the above disclosed content, the present invention first obtains a plurality of historical financial credit feature data constructed from the historical financial data, historical social network data, and personal data of sample users. Then, using the foregoing historical financial credit feature data as training data, a neural network model is trained, that is, using the historical financial credit feature data of each sample user as input and the overdue probability of each sample user as output to train the neural network model, thereby obtaining the trained model; thereafter, the financial data, social network data, and personal data of the target user are obtained, and based on the foregoing data, the financial credit feature data of the target user is constructed; then, the financial credit feature data of the target user is input into the foregoing trained model, and the overdue probability of the target user can be obtained; finally, based on this overdue probability, a corresponding risk warning message can be sent to the bank monitoring terminal, thereby realizing the real-time detection and warning of banking business risks.
[0013] Through the above design, the present invention uses the rich financial historical data, social network data, and personal data of different users to train the neural network model, thereby obtaining the trained risk detection model; then, based on this trained model, the real-time detection of banking business risks can be realized; based on this, using machine learning methods for risk detection, compared with traditional technologies, not only can improve data processing efficiency, but also the neural network model has the ability of self-learning and adaptation. Therefore, it can adapt to the rapid changes in the market, and thus can effectively identify the potential risks of the business. Therefore, the present invention effectively improves the efficiency, accuracy, and adaptability of risk control, and thus is very suitable for large-scale application and promotion.
[0014] In a possible design, obtaining the training data set includes:
[0015] Obtain an initial data set, where the initial data set includes the credit risk associated data of several sample users and the label data of each sample user. The credit risk associated data of any sample user includes the historical financial data, historical social network data, and personal data of the any sample user, and the label data of the any sample user is used to represent whether the any sample user is an overdue user;
[0016] Perform data preprocessing on each credit risk associated data in the initial dataset to obtain a preprocessed dataset;
[0017] Perform feature screening on the preprocessed dataset to obtain multiple credit-granting attribute information for each sample user;
[0018] Perform feature encoding on each credit-granting attribute information of each sample user, so that after the feature encoding, multiple credit-granting feature attributes corresponding to each sample user are obtained;
[0019] Perform feature concatenation on the multiple credit-granting feature attributes corresponding to each sample user to obtain historical financial credit-granting feature data corresponding to each sample user;
[0020] Use the label data and historical financial credit-granting feature data of each sample user to form the training dataset.
[0021] In a possible design, performing data preprocessing on each credit risk associated data in the initial dataset to obtain a preprocessed dataset includes:
[0022] Perform data cleaning on each credit risk associated data to obtain each credit risk associated data after cleaning;
[0023] Perform missing value filling on each credit risk associated data after cleaning to obtain each corrected credit risk associated data, and perform standardization processing on each corrected credit risk associated data, so that after the standardization processing, the preprocessed dataset is obtained.
[0024] In a possible design, after obtaining multiple credit-granting feature attributes corresponding to each sample user, the method further includes:
[0025] Perform normalization on each credit-granting feature attribute corresponding to each sample user to obtain multiple normalized feature attributes corresponding to each sample user, so as to use the multiple normalized feature attributes corresponding to each sample user to concatenate and obtain historical financial credit-granting feature data corresponding to each sample user.
[0026] In a possible design, multiple credit-granting feature attributes corresponding to each sample user are all numerically encoded, where performing normalization on each credit-granting feature attribute corresponding to each sample user to obtain multiple normalized feature attributes corresponding to each sample user includes:
[0027] For the j-th credit-granting feature attribute of any sample user, select the maximum value and the minimum value of the j-th credit-granting feature attribute from the multiple credit-granting feature attributes corresponding to all sample users;
[0028] Normalize the j-th credit feature attribute according to the maximum value and the minimum value of the j-th credit feature attribute, so as to obtain the normalized feature attribute corresponding to the j-th credit feature attribute;
[0029] Increment j by 1, and re-screen the maximum value and the minimum value of the j-th credit feature attribute from the multiple credit feature attributes corresponding to all sample users until j is equal to M, so as to obtain the multiple normalized feature attributes corresponding to any sample user, where the initial value of j is 1, and M is the total number of credit feature attributes of any sample user.
[0030] In a possible design, normalizing the j-th credit feature attribute according to the maximum value and the minimum value of the j-th credit feature attribute to obtain the normalized feature attribute corresponding to the j-th credit feature attribute includes:
[0031] Calculate a first difference between the maximum value and the minimum value, and calculate a second difference between the j-th credit feature attribute and the minimum value;
[0032] Use the ratio between the second difference and the first difference as the normalized feature attribute corresponding to the j-th credit feature attribute.
[0033] In a possible design, the loss function of the risk detection model is:
[0034]
[0035] In formula (1), Loss represents the loss function, f i (X i ) represents the overdue probability output by the risk detection model when taking the historical financial credit feature data of the i-th sample user as input, y i represents the label data of the i-th sample user, and N represents the total number of sample users.
[0036] In a second aspect, a bank risk detection device based on machine learning is provided, including:
[0037] An acquisition unit, configured to acquire a training data set, where the training data set includes historical financial credit feature data of a number of sample users, and the historical financial credit feature data of any sample user is constructed based on the historical financial data, historical social network data and personal data of the any sample user;
[0038] A training unit, configured to train a risk detection model by using the historical financial credit feature data of each sample user in the training data set as input and the overdue probability of each sample user as output, so as to obtain a trained risk detection model;
[0039] An acquisition unit, configured to acquire the financial data, social network data, and personal data of a target user, and construct the financial credit feature data of the target user by using the financial data, social network data, and personal data of the target user;
[0040] A risk detection unit, configured to input the financial credit feature data into a trained risk detection model to obtain the overdue probability of the target user;
[0041] An early warning unit, configured to generate a risk early warning message according to the overdue probability of the target user, and send the risk early warning message to the bank monitoring end.
[0042] In a third aspect, another bank risk detection device based on machine learning is provided. Taking the device as an example of an electronic device, it includes a memory, a processor, and a transceiver that are communicatively connected in sequence. Among them, the memory is used to store a computer program, the transceiver is used to send and receive messages, and the processor is used to read the computer program and execute the bank risk detection method based on machine learning as described in the first aspect or any possible design in the first aspect.
[0043] In a fourth aspect, a storage medium is provided, on which instructions are stored. When the instructions run on a computer, the bank risk detection method based on machine learning as described in the first aspect or any possible design in the first aspect is executed.
[0044] In a fifth aspect, a computer program product containing instructions is provided. When the instructions run on a computer, the computer is made to execute the bank risk detection method based on machine learning as described in the first aspect or any possible design in the first aspect.
[0045] Beneficial effects:
[0046] (1) The present invention uses the rich financial historical data, social network data, and personal data of different users to train a neural network model, thereby obtaining a trained risk detection model; then, based on this trained model, real-time detection of bank business risks can be achieved; based on this, machine learning is used for risk detection. Compared with traditional technologies, it can not only improve data processing efficiency, but also the neural network model has the ability of self-learning and adaptation. Therefore, it can adapt to the rapid changes in the market, and thus can effectively identify potential risks of business. Therefore, the present invention effectively improves the efficiency, accuracy, and adaptability of risk control, and is thus very suitable for large-scale application and promotion. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 It is a schematic flowchart of the steps of the bank risk detection method based on machine learning provided by an embodiment of the present invention;
[0048] Figure 2 The structural schematic diagram of the bank risk detection device based on machine learning provided by the embodiment of the present invention;
[0049] Figure 3 The structural schematic diagram of the electronic device provided by the embodiment of the present invention. Detailed implementation manners
[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the present invention will be briefly introduced below with reference to the accompanying drawings and the descriptions of the embodiments or the prior art. Obviously, the following descriptions of the structures of the drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. It should be noted here that the descriptions of these embodiment modes are used to help understand the present invention, but do not constitute a limitation to the present invention.
[0051] It should be understood that although terms such as first and second may be used herein to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, the first unit may be referred to as the second unit, and similarly, the second unit may be referred to as the first unit, without departing from the scope of the exemplary embodiments of the present invention.
[0052] It should be understood that for the term "and / or" that may appear in this article, it is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, B exists alone, and A and B exist simultaneously; for the term " / and" that may appear in this article, it is a description of another association object relationship, indicating that two relationships may exist. For example, A / and B may represent: A exists alone, and A and B exist alone; in addition, for the character " / " that may appear in this article, generally, the front and rear associated objects represent an "or" relationship.
[0053] Embodiment:
[0054] See Figure 1As shown in the figure, the bank risk detection method based on machine learning provided in this embodiment performs risk detection of banking business based on machine learning. In this way, compared with traditional technologies, not only can the data processing efficiency be improved, but the neural network model has the ability of self-learning and adaptation. Therefore, it can adapt to the rapid changes in the market, and thus can effectively identify the potential risks of the business. Therefore, this method effectively improves the efficiency, accuracy and adaptability of risk control, and is thus very suitable for large-scale application and promotion. Among them, for example, this method can be run on the risk detection side, and optionally, the risk detection side can be, but is not limited to, a personal computer (PC) and / or a server (preferably a server in this embodiment). It can be understood that the foregoing execution subject does not constitute a limitation on the embodiments of the present application. Correspondingly, the running steps of this method can be, but are not limited to, as shown in the following steps S1 to S5.
[0055] S1. Obtain a training data set. Among them, the training data set includes historical financial credit feature data of several sample users, and the historical financial credit feature data of any sample user is constructed based on the historical financial data, historical social network data and personal data of the sample user. In this embodiment, it is equivalent to integrating data sources from multiple channels to perform feature extraction and data storage, so as to construct the historical financial credit feature data of different sample users. Then, use the historical financial credit feature data of each sample user to form a training data set.
[0056] In specific applications, one of the construction methods of the foregoing training data set is disclosed below, which can be, but is not limited to, the following steps S11 to S16.
[0057] S11. Obtain an initial data set. Among them, the initial data set includes credit risk associated data of several sample users and label data of each sample user. The credit risk associated data of any sample user includes the historical financial data, historical social network data and personal data of the sample user, and the label data of the sample user is used to represent whether the sample user is an overdue user. In specific implementation, for example, the historical financial data of any sample user can be, but is not limited to: the transaction history, behavior log, and credit assessment report of the sample user, and the personal data can be, but is not limited to, the personal information of the sample user, such as name, age, gender, education level, etc. At the same time, for example, the label data of any sample user is represented by a digital label. For example, label 1 indicates that the sample user is an overdue user, and label 0 indicates that the sample user is not an overdue user.
[0058] Thus, obtaining diverse data sources can assist the model in comprehensively capturing potential risk factors, thereby improving the accuracy of prediction. After collecting data from multiple channels, data preprocessing can be carried out to reduce noise interference and ensure the reliability of the data. Among them, the data preprocessing process can include but is not limited to the steps shown in the following step S12.
[0059] S12. Perform data preprocessing on each credit risk-related data in the initial data set to obtain a preprocessed data set. In this embodiment, due to possible problems such as missing values, noise, and outliers in the original data, data preprocessing is to perform data cleaning, filling in missing values, and standardization processing, that is: first perform data cleaning on each credit risk-related data to obtain each cleaned credit risk-related data; then, perform missing value filling on each cleaned credit risk-related data to obtain each corrected credit risk-related data; finally, perform standardization processing on each corrected credit risk-related data. Thus, after standardization processing, the aforementioned preprocessed data set can be obtained.
[0060] After completing the data preprocessing, feature screening processing can be carried out, and the process is as shown in the following step S13.
[0061] S13. Perform feature screening on the preprocessed data set to obtain multiple credit attribute information of each sample user. In specific implementation, feature screening is a key step in improving the model's performance. Its purpose is to transform the original data into feature variables that can reflect the user's risk level. Among them, for example, but not limited to, the recursive feature elimination method can be used. By repeatedly constructing models, evaluating the impact of features on user delinquency, and gradually removing the least important features until the specified number of features or the ideal model performance level is reached. Thus, features irrelevant to whether the user is delinquent can be eliminated, and key features, that is, the credit attribute features of the credit users, can be retained as the model input. Of course, the aforementioned recursive feature elimination method is a commonly used method for feature importance evaluation, and its principle will not be elaborated here.
[0062] Thus, based on the aforementioned step S13, after completing feature screening and obtaining the credit attribute features of each sample user, feature encoding processing can be carried out to digitize the text features, thereby facilitating model processing. Among them, the feature encoding process can include but is not limited to the steps shown in the following step S14.
[0063] S14. Perform feature encoding processing on the credit granting attribute information of each sample user, so as to obtain multiple credit granting feature attributes corresponding to each sample user after the feature encoding processing; in specific implementation, it is equivalent to converting each credit granting attribute information into digital encoding, such as using one-hot encoding or other encodings, etc.; of course, the specific encoding method can be specifically set according to actual use and is not specifically limited here.
[0064] Meanwhile, after completing the feature encoding, since each credit granting feature attribute has been converted into a number, therefore, in order to unify data with different dimensions to the same scale for subsequent comparison and analysis, this embodiment also performs normalization processing on each credit granting feature attribute corresponding to each sample user, so as to obtain multiple normalized feature attributes corresponding to each sample user.
[0065] Among them, the following takes the multiple credit granting feature attributes of any sample user as an example to elaborate. The specific process of normalization can be but is not limited to the following first to third steps.
[0066] First step: For the j-th credit granting feature attribute of any sample user, screen out the maximum value and the minimum value of the j-th credit granting feature attribute from the multiple credit granting feature attributes corresponding to all sample users; in this embodiment, an example is used to elaborate this step. Assume that the j-th credit granting feature attribute is educational background, and the educational backgrounds of all sample users are mainly divided into "HighSchool (high school)", "Bachelor (bachelor)", "Master (master)", and "PhD (doctor)". At the same time, the digital encodings corresponding to the foregoing educational backgrounds are 0, 1, 2, and 3 in sequence. Thus, the maximum value and the minimum value of the credit granting feature attribute of educational background are 3 and 0 respectively; of course, the screening methods for the maximum value and the minimum value of other different credit granting feature attributes are the same as the foregoing example and will not be elaborated here.
[0067] After screening out the maximum value and the minimum value of the j-th credit granting feature attribute, the normalization processing can be performed, and the process is as shown in the following second step.
[0068] Second step: According to the maximum value and the minimum value of the j-th credit granting feature attribute, perform normalization processing on the j-th credit granting feature attribute to obtain the normalized feature attribute corresponding to the j-th credit granting feature attribute; in this embodiment, the example can be but is not limited to first calculating the first difference between the maximum value and the minimum value, and calculating the second difference between the j-th credit granting feature attribute and the minimum value; then, taking the ratio between the second difference and the first difference as the normalized feature attribute corresponding to the j-th credit granting feature attribute.
[0069] In this way, through the foregoing method, the normalization process of the j-th credit feature attribute can be completed; then, in the same way as above, the normalization process of the remaining credit feature attributes can be completed; among them, the loop process is shown in the following third step.
[0070] Third step: increment j by 1, and then re-screen the maximum and minimum values of the j-th credit feature attribute from the multiple credit feature attributes corresponding to all sample users until j is equal to M, where the initial value of j is 1, and M is the total number of credit feature attributes of any sample user.
[0071] Thus, through the foregoing first to third steps, the normalization process of the credit feature attributes of each sample user can be completed; then, the historical financial credit feature data corresponding to each sample user can be obtained by splicing the multiple normalized feature attributes corresponding to each sample user, and the process is shown in the following step S15.
[0072] S15. Perform feature splicing on the multiple credit feature attributes corresponding to each sample user to obtain the historical financial credit feature data corresponding to each sample user; in this embodiment, it is equivalent to splicing the multiple credit feature data of each sample user (of course, referring to the features after normalization) into a vector, so as to form the historical financial credit feature data corresponding to each sample user; in this way, the training data set can be formed by combining the foregoing label data, and the process is shown in the following step S16.
[0073] S16. Use the label data and historical financial credit feature data of each sample user to form the training data set.
[0074] In this way, through the foregoing steps S11 to S16, the construction of the training data set can be completed. For example, the training data can be stored in a suitable database (such as a relational database or a non-relational database) to ensure fast access and efficient storage of the data, so as to facilitate subsequent model training and testing; of course, database design also needs to consider data permission management and privacy protection to ensure the security of user data.
[0075] After the training data set is constructed, the training of the neural network model can be carried out, and the process can be but is not limited to the following step S2.
[0076] S2. Use the historical financial credit feature data of each sample user in the training dataset as the input, and the overdue probability of each sample user as the output to train the risk detection model, so as to obtain the trained risk detection model. In specific implementation, for example, the aforementioned risk detection model can be, but is not limited to, machine learning models such as logistic regression, decision tree, or support vector machine. At the same time, for example, the loss function of the aforementioned risk detection model can be, but is not limited to, as shown in the following formula (1).
[0077]
[0078] In formula (1), Loss represents the loss function, and f i (X i ) represents the overdue probability output by the risk detection model when using the historical financial credit feature data of the i-th sample user as the input, and y i represents the label data of the i-th sample user, and N represents the total number of sample users.
[0079] In this way, the model training can be carried out in combination with the aforementioned loss function.
[0080] Furthermore, for example, but not limited to, through K-fold cross-validation and grid search methods, the model training and optimization can be carried out, that is, the generalization performance of the machine learning model can be tested by using methods such as cross-validation and stratified sampling, so as to ensure that the model has strong generalization ability. At the same time, the prediction performance of the model can also be comprehensively analyzed through evaluation indicators such as ROC curve and AUC value to further optimize the model and improve the prediction accuracy.
[0081] Even further, through grid search and cross-validation techniques, the optimal parameter combination can be found to ensure that the model performs excellently in both the training and validation stages. In addition, ensemble learning methods such as random forest and XGBoost can further enhance the prediction performance, and by integrating the advantages of multiple basic models, the error of a single model can be effectively reduced, thereby improving the stability of the system.
[0082] In this way, through the aforementioned step S2, the model training can be completed. Then, the trained model can be embedded in the system to achieve the real-time risk warning function. Among them, for example, but not limited to, a distributed computing framework (such as Hadoop and Spark) can be used to achieve parallel processing of large-scale data to speed up the system response speed, and stream data processing technology (such as Apache Flink) can be used to achieve real-time risk warning, so as to ensure that the system can quickly respond to emergencies.
[0083] After the model deployment is completed, the risk warning of real-time business can be carried out, and the process is as shown in the following steps S3 - S5.
[0084] S3. Obtain the financial data, social network data, and personal data of the target user, and use the financial data, social network data, and personal data of the target user to construct the financial credit feature data of the target user; in this embodiment, for the content included in the financial data, social network data, and personal data of the target user, reference can be made to the foregoing sample user, and the construction process of the corresponding financial credit feature data can also be referred to the sample user, which will not be elaborated here.
[0085] After constructing the financial credit feature data of the target user based on step S3, it can be input into the trained risk detection model to obtain the overdue probability of the target user; among them, the risk detection process is as shown in step S4 below.
[0086] S4. Input the financial credit feature data into the trained risk detection model to obtain the overdue probability of the target user; in this embodiment, after predicting the overdue probability of the target user based on the foregoing trained risk detection model, corresponding warning information can be generated to prompt the bank staff; among them, the warning process is as shown in step S5 below.
[0087] S5. Generate risk warning information according to the overdue probability of the target user and send the risk warning information to the bank monitoring end; in specific implementation, for example, when the overdue probability exceeds the preset threshold, corresponding risk warning information can be generated, and the warning information can include the personal information of the target user and the business handled. Therefore, after sending the risk warning information to the bank monitoring end, the bank staff can be prompted to take corresponding risk control measures.
[0088] Optionally, for example, this embodiment can be applied to the field of bank loan review. That is, when the overdue probability of the target user obtained based on the trained risk detection model exceeds the preset threshold, the target user can be classified as a non-creditworthy customer, and thus, it can be added to the risk warning information and sent to the bank monitoring end; in this way, the bank staff can conduct loan review based on this warning information; of course, when the overdue probability of the target user is obtained to exceed the preset threshold, the result of loan review not passed can also be directly generated and sent to the bank staff, so that the bank staff can use this as the final review result for loan review reply.
[0089] In addition, the foregoing is only one business scenario of this embodiment applied to the bank, and it is not limited to this here.
[0090] Thus, through the method for bank risk detection based on machine learning described in detail in the foregoing steps S1 to S5, the present invention performs risk detection of banking services in a machine learning manner. In this way, compared with the traditional technology, not only can the data processing efficiency be improved, but also the neural network model has the ability of self-learning and adaptation. Therefore, it can adapt to the rapid changes in the market, and thus can effectively identify the potential risks of the services. Therefore, the present invention effectively improves the efficiency, accuracy and adaptability of risk control, and is thus very suitable for large-scale application and promotion.
[0091] As Figure 2 shown, in the second aspect of this embodiment, there is provided a hardware device for implementing the method for bank risk detection based on machine learning described in the first aspect of the embodiment, including:
[0092] An acquisition unit, configured to acquire a training data set, where the training data set includes historical financial credit feature data of a number of sample users, and the historical financial credit feature data of any sample user is constructed based on the historical financial data, historical social network data and personal data of the any sample user.
[0093] A training unit, configured to use the historical financial credit feature data of each sample user in the training data set as input and the overdue probability of each sample user as output to train a risk detection model to obtain a trained risk detection model.
[0094] An acquisition unit, configured to acquire the financial data, social network data and personal data of a target user, and construct the financial credit feature data of the target user by using the financial data, social network data and personal data of the target user.
[0095] A risk detection unit, configured to input the financial credit feature data into the trained risk detection model to obtain the overdue probability of the target user.
[0096] An early warning unit, configured to generate a risk early warning message according to the overdue probability of the target user and send the risk early warning message to the bank monitoring end.
[0097] For the working process, working details and technical effects of the device provided in this embodiment, reference can be made to the first aspect of the embodiment, which will not be elaborated here.
[0098] As Figure 3As shown, the third aspect of this embodiment provides another bank risk detection device based on machine learning. Taking the device as an electronic device as an example, it includes: a memory, a processor, and a transceiver that are communicatively connected in sequence. Among them, the memory is used to store computer programs, the transceiver is used to send and receive messages, and the processor is used to read the computer programs and execute the bank risk detection method based on machine learning described in the first aspect of the embodiment.
[0099] Specifically, the memory may include, but is not limited to, random access memory (RAM), read only memory (ROM), flash memory, first input first output (FIFO), and / or first in last out (FILO), etc.; specifically, the processor may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor can be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). At the same time, the processor can also include a main processor and a coprocessor. The main processor is a processor used to process data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state.
[0100] In some embodiments, the processor may integrate a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. For example, the processor may not be limited to a microprocessor of the STM32F105 series, a reduced instruction set computer (RISC) microprocessor, an X86 architecture processor, or a processor integrated with an embedded neural-network processing unit (NPU); the transceiver may be, but is not limited to, a Wi-Fi wireless transceiver, a Bluetooth wireless transceiver, a General Packet Radio Service (GPRS) wireless transceiver, a ZigBee (low-power local area network protocol based on the IEEE 802.15.4 standard) wireless transceiver, a 3G transceiver, a 4G transceiver, and / or a 5G transceiver, etc. In addition, the device may also include, but is not limited to, a power module, a display screen, and other necessary components.
[0101] For the working process, working details, and technical effects of the electronic device provided in this embodiment, reference may be made to the first aspect of the embodiment, and details are not elaborated herein.
[0102] In the fourth aspect of this embodiment, a storage medium storing instructions for the machine learning-based bank risk detection method described in the first aspect of the embodiment is provided, that is, instructions are stored on the storage medium, and when the instructions run on a computer, the machine learning-based bank risk detection method described in the first aspect of the embodiment is executed.
[0103] Among them, the storage medium refers to a carrier for storing data, and may include, but is not limited to, a floppy disk, an optical disc, a hard disk, a flash memory, a USB flash drive, and / or a Memory Stick, etc. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.
[0104] For the working process, working details, and technical effects of the storage medium provided in this embodiment, reference may be made to the first aspect of the embodiment, and details are not elaborated herein.
[0105] In the fifth aspect of this embodiment, a computer program product containing instructions is provided. When the instructions run on a computer, the computer is caused to execute the machine learning-based bank risk detection method described in the first aspect of the embodiment, where the computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.
[0106] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A bank risk detection method based on machine learning, characterized in that: include: Acquire a training data set, wherein the training data set includes historical financial credit feature data of a number of sample users, and the historical financial credit feature data of any sample user is constructed based on the historical financial data, historical social network data, and personal data of the any sample user; Taking the historical financial credit feature data of each sample user in the training data set as input and the overdue probability of each sample user as output, the risk detection model is trained to obtain a trained risk detection model; Acquire the target user's financial data, social network data, and personal data, and construct the target user's financial credit feature data using the target user's financial data, social network data, and personal data; Inputting the financial credit feature data into the trained risk detection model to obtain the overdue probability of the target user; According to the overdue probability of the target user, risk warning information is generated and sent to the bank monitoring terminal.
2. The method according to claim 1, characterized in that Obtaining the training data set includes: Acquire an initial data set, wherein the initial data set includes credit risk-related data of several sample users and label data of each sample user, the credit risk-related data of any sample user includes the historical financial data, historical social network data and personal data of any sample user, and the label data of any sample user is used to characterize whether any sample user is an overdue user; Performing data preprocessing on each credit risk associated data in the initial data set to obtain a preprocessed data set; Performing feature screening processing on the preprocessed data set to obtain multiple credit attribute information of each sample user; Performing feature coding processing on each credit attribute information of each sample user, so as to obtain multiple credit feature attributes corresponding to each sample user after the feature coding processing; Perform feature concatenation processing on multiple credit feature attributes corresponding to each sample user to obtain historical financial credit feature data corresponding to each sample user; The training data set is composed of the label data and historical financial credit feature data of each sample user.
3. The method according to claim 2, characterized in that Data preprocessing is performed on each credit risk associated data in the initial data set to obtain a preprocessed data set, including: Perform data cleaning on each credit risk-related data to obtain cleaned credit risk-related data; The cleaned credit risk-related data are each subjected to missing value filling processing to obtain corrected credit risk-related data, and the corrected credit risk-related data are each subjected to standardization processing to obtain the pre-processed data set after the standardization processing.
4. The method according to claim 2, characterized in that: After obtaining a plurality of credit feature attributes corresponding to each sample user, the method further includes: Each credit feature attribute corresponding to each sample user is normalized to obtain multiple normalized feature attributes corresponding to each sample user, so as to use the multiple normalized feature attributes corresponding to each sample user to splice and obtain the historical financial credit feature data corresponding to each sample user.
5. The method according to claim 4, characterized in that The multiple credit feature attributes corresponding to each sample user are all digitally encoded, wherein each credit feature attribute corresponding to each sample user is normalized to obtain multiple normalized feature attributes corresponding to each sample user, including: For the j-th credit feature attribute of any sample user, the maximum value and the minimum value of the j-th credit feature attribute are screened out from multiple credit feature attributes corresponding to all sample users; According to the maximum value and the minimum value of the j-th credit feature attribute, normalizing the j-th credit feature attribute to obtain a normalized feature attribute corresponding to the j-th credit feature attribute; Add 1 to j, and re-screen the maximum and minimum values of the j-th credit feature attribute from the multiple credit feature attributes corresponding to all sample users until j is equal to M, and obtain multiple normalized feature attributes corresponding to any sample user, where the initial value of j is 1, and M is the total number of credit feature attributes of any sample user.
6. The method according to claim 5, characterized in that According to the maximum value and the minimum value of the j-th credit feature attribute, the j-th credit feature attribute is normalized to obtain a normalized feature attribute corresponding to the j-th credit feature attribute, including: Calculating a first difference between the maximum value and the minimum value, and calculating a second difference between the j-th credit feature attribute and the minimum value; The ratio of the second difference to the first difference is used as the normalized characteristic attribute corresponding to the j-th credit characteristic attribute.
7. The method according to claim 1, characterized in that The loss function of the risk detection model is: In formula (1), Loss represents the loss function, f i (X i ) represents the overdue probability output by the risk detection model when the historical financial credit feature data of the i-th sample user is used as input, y i represents the label data of the i-th sample user, and N represents the total number of sample users.
8. A bank risk detection device based on machine learning, characterized in that: include: An acquisition unit is used to acquire a training data set, wherein the training data set includes historical financial credit feature data of a number of sample users, and the historical financial credit feature data of any sample user is constructed based on the historical financial data, historical social network data and personal data of the any sample user; A training unit, used to train a risk detection model using the historical financial credit feature data of each sample user in the training data set as input and the overdue probability of each sample user as output, so as to obtain a trained risk detection model; An acquisition unit, configured to acquire financial data, social network data, and personal data of a target user, and construct financial credit feature data of the target user using the financial data, social network data, and personal data of the target user; A risk detection unit, used to input the financial credit feature data into the trained risk detection model to obtain the overdue probability of the target user; The early warning unit is used to generate risk early warning information according to the overdue probability of the target user, and send the risk early warning information to the bank monitoring terminal.
9. An electronic device, characterized in that: include: A memory, a processor and a transceiver that are sequentially communicatively connected, wherein the memory is used to store computer programs, the transceiver is used to send and receive messages, and the processor is used to read the computer program, and execute the bank risk detection method based on machine learning as described in any one of claims 1 to 7.
10. A computer program product comprising instructions, characterized in that When the instructions are executed on a computer, the computer is enabled to execute the bank risk detection method based on machine learning as described in any one of claims 1 to 7.