Deep learning algorithm-based fraud-related user identification method and device

Through the method based on deep learning algorithm, the terminal device model characteristics are used to identify fraud-related users, which solves the problem of low recognition accuracy in the existing technology, and achieves more accurate telecommunications network fraud governance and intuitive credibility scores.

CN120018137APending Publication Date: 2025-05-16广州市申迪计算机系统有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510042560.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the prior art, the identification accuracy of fraud-related users is low, and the terminal device model and its derivative characteristics are not fully considered.

Method used

Using a deep learning algorithm method, the target user's terminal device model is obtained, the corresponding features are extracted, and the pre-trained deep learning model is input to output fraud probability values ​​and credibility scores to determine whether the user is a fraud-related user.

Benefits of technology

It improves the identification accuracy of fraud-related users, provides a new dimension for risk assessment, can more accurately control telecommunications network fraud, and intuitively understand the results through credibility scores.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120018137A_ABST
    Figure CN120018137A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of communication security, and discloses a fraud-related user identification method and device based on a deep learning algorithm, and the method comprises the steps: obtaining a terminal equipment model of a target user; extracting terminal equipment model characteristics corresponding to the target user based on the terminal equipment model; inputting the model features of the terminal equipment into a pre-trained deep learning model in a feature vector format, and outputting whether the model of the terminal equipment of the target user relates to fraud or not and a fraud-related probability value; according to the fraud-related probability value, determining a credibility score of the terminal equipment type of the target user; judging whether the credibility score meets a fraud condition or not; and if the credibility score satisfies a fraud-related condition, determining that the target user is a fraud-related user. According to the method and the device, the fraud-related user is evaluated from a brand new perspective of the terminal brand model, so that the evaluation accuracy of the fraud-related user is improved, and the effect of treating the telecommunication network fraud more accurately is facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of communication security technology, and in particular to a method and device for identifying fraudulent users based on a deep learning algorithm. Background Art

[0002] At present, the offensive and defensive confrontation against telecommunications network fraud is constantly escalating. Fraudsters are constantly using new fraud methods, and telecommunications operators' systems are constantly upgrading their prevention methods. The confrontation has entered a white-hot period.

[0003] In the existing technology, the risk assessment of users involved in fraud is often calculated based on artificial rules constructed based on manual experience, or calculated using machine learning algorithm modeling. There are limitations in the identification of users involved in fraud, and the identification accuracy needs to be improved. There is little practice in the field of deep learning algorithm modeling. The selected features often focus on user behavior characteristics, such as historical calls, traffic, text messages, etc., and user profile information characteristics, such as network access duration, package value, whether the package is integrated, user age, user registered residence, etc., without considering the terminal device model and a series of key features derived from it.

[0004] With respect to the above-mentioned related technologies, the inventors have discovered that the existing fraud-related user identification method has the problem of low identification accuracy. Summary of the invention

[0005] In order to improve the accuracy of identifying fraudulent users, the present application provides a method and device for identifying fraudulent users based on a deep learning algorithm.

[0006] In a first aspect, the present application provides a method for identifying fraudulent users based on a deep learning algorithm.

[0007] This application is achieved through the following technical solutions:

[0008] A method for identifying fraudulent users based on a deep learning algorithm comprises the following steps:

[0009] Obtain the terminal device model of the target user;

[0010] Based on the terminal device model, extracting terminal device model features corresponding to the target user;

[0011] Input the terminal device model features into a pre-trained deep learning model in a feature vector format, and output whether the terminal device model of the target user is involved in fraud and the probability value of fraud;

[0012] Determining the credibility score of the terminal device model of the target user according to the fraud probability value;

[0013] Determining whether the credibility score satisfies the fraud-related condition;

[0014] If the credibility score meets the fraud-related condition, the target user is determined to be a fraud-related user.

[0015] In a preferred example, the present application can be further configured as follows: the steps of constructing the deep learning model include:

[0016] Based on the long short-term memory temporal neural network model, the model's input vector dimension is set to 18, the middle layer is 1, the number of nodes in the middle layer is 1, and the iteration step is 150; and

[0017] The model is set to use a nonlinear activation function and a non-bidirectional neural network model.

[0018] In a preferred example, the present application can be further configured as follows: the training step of the deep learning model includes:

[0019] Extract the corresponding terminal device model historical features from the signaling history data of the fraudulent user and construct a sample data set;

[0020] Dividing the sample data set into a training set and a test set in proportion;

[0021] Based on the training set, combined with a preset excitation function and a loss function, the deep learning model is iteratively trained to output a fraud probability value, wherein the excitation function is a softplus function and the loss function is a MAPE function;

[0022] When the iteration step of the deep learning model reaches a set value, the deviation between the current fraud probability value of the deep learning model and the true value is calculated to determine whether the deviation is reduced to a preset deviation threshold;

[0023] When the deviation is reduced to the deviation threshold, the current deep learning model is output as the trained deep learning model.

[0024] In a preferred example, the present application can be further configured as follows: it also includes the following steps:

[0025] When the deviation is not reduced to the deviation threshold, repeatedly iteratively training the deep learning model, and counting the number of calculations of the deviation;

[0026] If the deviation is not equal to the deviation threshold, determining whether the number of calculations reaches a preset number threshold;

[0027] When the number of calculations reaches the number threshold, the current deep learning model is output as a trained deep learning model.

[0028] In a preferred example, the present application may be further configured as follows: the training step of the deep learning model further includes:

[0029] Setting the optimization function of the deep learning model to a gradient descent function, performing optimization training on the deep learning model, and iteratively increasing the number of layers and the number of nodes in the middle layer;

[0030] Until the loss function reaches a minimum value, the current deep learning model is output as the trained deep learning model. At this time, the middle layer is 3 and the number of nodes in the middle layer is 6.

[0031] In a preferred example, the present application may be further configured as follows: the step of determining the credibility score of the terminal device model of the target user according to the fraud probability value includes:

[0032] Based on the fraud probability value and the normal user probability value, an evaluation model is constructed;

[0033] Wherein, the expression of the evaluation model includes:

[0034] score=A+B×ln(odds)

[0035] Where, odds is the ratio of the fraud probability value to the normal user probability value, p0 is the score corresponding to a certain probability value, and PDO is the score that increases when a certain probability value doubles;

[0036] The fraud probability value is substituted into the evaluation model to obtain the credibility score of the terminal device model of the target user.

[0037] In a second aspect, the present application provides a fraudulent user identification device based on a deep learning algorithm.

[0038] This application is achieved through the following technical solutions:

[0039] A fraud-related user identification device based on a deep learning algorithm, comprising:

[0040] The device model module is used to obtain the terminal device model of the target user;

[0041] A feature module, used to extract a terminal device model feature corresponding to the target user based on the terminal device model;

[0042] A calculation module, used to input the terminal device model features into a pre-trained deep learning model in a feature vector format, and output whether the terminal device model of the target user is fraudulent and the probability value of fraudulent;

[0043] A mapping module, used to determine the credibility score of the terminal device model of the target user according to the fraud probability value;

[0044] A detection module, used to determine whether the credibility score meets the fraud-related condition;

[0045] The identification module is used to determine that the target user is a fraudulent user when the credibility score meets the fraudulent condition.

[0046] In a third aspect, the present application provides a computer device.

[0047] This application is achieved through the following technical solutions:

[0048] A computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of any one of the above-mentioned methods for identifying fraudulent users based on a deep learning algorithm are implemented.

[0049] In a fourth aspect, the present application provides a computer-readable storage medium.

[0050] This application is achieved through the following technical solutions:

[0051] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of any one of the above-mentioned methods for identifying fraudulent users based on a deep learning algorithm.

[0052] In a fifth aspect, the present application provides a computer program product.

[0053] This application is achieved through the following technical solutions:

[0054] A computer program product includes a computer program, which, when executed by a processor, implements the steps of any one of the above-mentioned methods for identifying fraudulent users based on a deep learning algorithm.

[0055] In summary, compared with the prior art, the technical solution provided by this application has at least the following beneficial effects:

[0056] The terminal device model of the target user is obtained, and then the terminal device model features corresponding to the target user are extracted, so as to evaluate the fraud risk of the target user from a new perspective of terminal brand model; the terminal device model features are input into the pre-trained deep learning model in feature vector format, and the target user's terminal device model is output whether it is fraudulent and the probability of fraud. This can fully leverage the advantages of deep learning algorithms to improve the accuracy of assessment of fraudulent users, which is conducive to more accurate governance of telecommunications network fraud; according to the probability of fraud, the credibility score of the target user's terminal device model is determined to determine whether the fraud conditions are met, so as to realize the conversion of the probability value of the fraudulent terminal device model output by the model into an easy-to-understand model credibility score, which is intuitive and clear. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 A schematic diagram of the main process of a method for identifying fraudulent users based on a deep learning algorithm is provided as an exemplary embodiment of the present application.

[0058] Figure 2 A schematic diagram of price characteristics of a terminal device model based on a deep learning algorithm is provided as an exemplary embodiment of the present application.

[0059] Figure 3 A schematic diagram of new and old features of terminal device models based on a deep learning algorithm is provided as an exemplary embodiment of the present application.

[0060] Figure 4 A schematic diagram of the distribution characteristics of Apple and Android terminals in a non-communication work order for terminal device models based on a deep learning algorithm is provided as an exemplary embodiment of the present application.

[0061] Figure 5 A schematic diagram of the distribution characteristics of Apple and Android terminals in a communication work order for a terminal device model based on a deep learning algorithm is provided as an exemplary embodiment of the present application.

[0062] Figure 6 A schematic diagram of the distribution characteristics of dual-SIM dual-standby models and single-SIM single-standby models of terminal equipment models based on a deep learning algorithm for an exemplary embodiment of the present application, including Apple terminals and Android terminals in communication work orders. DETAILED DESCRIPTION

[0063] This specific embodiment is merely an explanation of the present application and is not a limitation of the present application. After reading this specification, those skilled in the art may make modifications to the present embodiment without any creative contribution as needed. However, as long as it is within the scope of the claims of the present application, it shall be protected by the patent law.

[0064] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0065] In addition, the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article, unless otherwise specified, generally means that the associated objects before and after are in an "or" relationship.

[0066] The embodiments of the present application are further described in detail below in conjunction with the drawings in the specification.

[0067] Reference Figure 1 , an embodiment of the present application provides a method for identifying fraudulent users based on a deep learning algorithm, and the main steps of the method are described as follows.

[0068] S1: Obtain the terminal device model of the target user;

[0069] S2: Based on the terminal device model, extracting terminal device model features corresponding to the target user;

[0070] S3: Input the terminal device model features into a pre-trained deep learning model in a feature vector format, and output whether the terminal device model of the target user is involved in fraud and the probability value of fraud;

[0071] S4: Determine the credibility score of the terminal device model of the target user according to the fraud probability value;

[0072] S5: Determine whether the credibility score meets the fraud-related condition;

[0073] S6: If the credibility score meets the fraud-related condition, the target user is determined to be a fraud-related user.

[0074] Specifically, taking a mobile phone as an example, the terminal device collects the target user's signaling data, uses the IMEI body code information based on the operator's number, and matches the brand and model data of the terminal device through the terminal library to obtain the target user's terminal device model.

[0075] The collected signaling data are shown in Tables 1 and 2 below.

[0076] Table 1

[0077]

[0078] The above table shows the key fields of 4G S1MME signaling data, namely the VOLTE registration signaling table.

[0079] Table 2

[0080]

[0081] The above table is the key fields of 5G N1 N2 signaling data, that is, the 5G N1 N2 signaling table.

[0082] The terminal library is pre-built through historical signaling data and corresponding brand model data. Based on the IMEI body code information used by the operator's number, the terminal device model of the target user can be obtained by querying the terminal library.

[0083] Next, based on the terminal device model, the terminal device model features corresponding to the target user are extracted, wherein the terminal device model features are obtained by big data technology analysis.

[0084] In this embodiment, the mining and analysis of typical mobile phone model features of fraud users are as follows:

[0085] The user numbers in the work orders involved in the case were used as the source of fraudulent users, and the internal users of telecom operators were used as non-fraudulent users to construct positive and negative samples for model training. Among them, there were 16,125 non-fraudulent users and 3,280 fraudulent users, totaling 19,405 numbers. A comparative analysis was conducted on the mobile phone model features of the above two types of users.

[0086] like Figure 2 As shown, through statistics, most of the mobile phone terminals used by fraud users are low-priced models, which is in line with the original intention of fraudsters to save costs, that is, to use the least money to purchase mobile phone terminals for fraud. After all, fraud terminals are very easy to be detected by the "black terminal model" of telecom operators, resulting in the number card on the terminal cannot be used normally after the fraud is carried out, and often can only be discarded.

[0087] like Figure 3 As shown, through statistics, a large proportion of fraud terminals are old mobile phones. Old terminal devices are defined as terminal devices that have been on the market for more than 2 years. This is in line with the original intention of fraudsters to save costs, that is, to spend less money to purchase mobile phone terminals from the second-hand mobile phone market for fraud.

[0088] like Figure 4As shown, through statistics, among the work orders involved in no communication (no telephone call signaling or call records occurred between the fraudsters and the victims), iPhone terminals accounted for a large proportion. The reason is that a large proportion of the work orders involved in the case were fraudulent using Facetime calls, and the calling terminal of this type of fraud can only be an Apple iPhone terminal.

[0089] like Figure 5 As shown, through statistics, Android terminals account for a large proportion of the work orders involved in communication (there is a telephone call signaling or call bill between the fraudsters and the victims), which is in line with expectations, because the vast majority of current communication frauds are implemented through simple GOIP networking, which is a method of connecting two mobile phones with an audio cable, and this type of networking is often only available on Android phones, and Apple phones do not support it.

[0090] like Figure 6 As shown, through statistics, the proportion of dual-SIM dual-standby models in the work orders involved in the case is relatively high. This is related to the efficiency considerations of fraudsters. In order to improve the efficiency of fraud, fraudsters tend to use dual-SIM dual-standby mobile phones and insert mobile phone cards of two operators at the same time to commit fraud.

[0091] Furthermore, according to statistics, among the work orders involved in the case, Redmi phones, OPPO phones, VIVO phones, and Xiaomi phones (which can modify IMEI, so they account for a certain proportion) account for a relatively high proportion.

[0092] Therefore, through the above big data analysis and statistics, we can fully explore the terminal device model characteristics of typical fraud-related users, including terminal device price characteristics, terminal device launch time characteristics, whether the terminal device model in the work order without communication is an Apple terminal, whether the terminal device model in the work order with communication is an Android terminal, whether the terminal device in the work order with communication is a dual-SIM dual-standby model, and whether the terminal device is a preset terminal device model (such as Redmi, OPPO, VIVO and Xiaomi La) as the terminal device model characteristics corresponding to the target users.

[0093] The terminal device models of the target users and the corresponding typical terminal device model features are shown in Table 3 below.

[0094] Table 3

[0095]

[0096] Among them, TAC ID is the first eight digits of the IMEI body code used based on the operator's number, which can be obtained through the IMEI information in the collected signaling data.

[0097] Next, based on the mined terminal device model features, a deep learning model is used to assess the fraud risk of target users.

[0098] In one embodiment, the step of constructing the deep learning model includes:

[0099] Based on the long short-term memory temporal neural network model, the model's input vector dimension is set to 18, the middle layer is 1, the number of nodes in the middle layer is 1, and the iteration step is 150; and

[0100] The model is set to use a nonlinear activation function and a non-bidirectional neural network model.

[0101] Since the data content studied in this scheme and the signaling data involved are all time series data, the long short-term memory temporal neural network (LSTM) model is based on bionics theory and improves the linear process of the traditional model. It retains the information of the previous layer and combines it with the new information input by the current layer for joint calculation. The cyclic network can allow some information to persist. Therefore, compared with other neural network models, the LSTM model has a stronger memory ability for time series. Therefore, the LSTM model is selected to assess the fraud risk of target users.

[0102] The LSTM model consists of a network of neurons connected by a weighted matrix. Positive values ​​indicate active connections, while negative values ​​indicate passive connections. All inputs are corrected and calculated through the weighted matrix. The LSTM model controls the information state through three gates: the forget gate, the input gate, and the output gate. Through a set of forget gates, input gates, and output gates, the input, calculation, and output process of a cell structure of the LSTM model is completed. After that, the model will perform iterative calculations in the middle layer, repeating the cycle until the specified result is obtained or the predetermined number of iterations is reached.

[0103] The LSTM model solves the problem of short memory of input vectors. Based on the standard recurrent neural network model, it can more fully consider the influencing factors of historical signaling data and achieve better fitting.

[0104] When applied, with the help of the Python-based Pytorch module (an open source Python software package developed by Facebook for machine learning in the United States), the torch.nnLSTMCell() method is called to build the LSTM model framework and set the parameters used by the model. Among them, the input vector dimension Input_size=18, the number of nodes in the middle layer Hidden_size=1, the number of stacked layers (middle layer) Num_layers=1 and the iteration step length epochs=150. In addition, a nonlinear excitation function is selected, and Nonlinearity=1 is set; a non-bidirectional neural network model is selected, and Bidirectional=0 is set.

[0105] In one embodiment, the training step of the deep learning model includes:

[0106] Extract the corresponding terminal device model historical features from the signaling history data of the fraudulent user and construct a sample data set;

[0107] Dividing the sample data set into a training set and a test set in proportion;

[0108] Based on the training set, combined with a preset excitation function and a loss function, the deep learning model is iteratively trained to output a fraud probability value, wherein the excitation function is a softplus function and the loss function is a MAPE function;

[0109] When the iteration step of the deep learning model reaches a set value, the deviation between the current fraud probability value of the deep learning model and the true value is calculated to determine whether the deviation is reduced to a preset deviation threshold;

[0110] When the deviation is reduced to the deviation threshold, outputting the current deep learning model as a trained deep learning model;

[0111] When the deviation is not reduced to the deviation threshold, repeatedly iteratively training the deep learning model, and counting the number of calculations of the deviation;

[0112] If the deviation is not equal to the deviation threshold, determining whether the number of calculations reaches a preset number threshold;

[0113] When the number of calculations reaches the number threshold, the current deep learning model is output as a trained deep learning model.

[0114] Specifically, the sample data set used for deep learning model training is the historical features of terminal device models extracted from the signaling history data of fraudulent users, which is consistent with the terminal device model features of the target users analyzed and counted above. The sample data set includes positive and negative samples. The sample data set is divided into training sets and test sets in proportion, such as training sets (80%) and test sets (20%). The training set is used for model training and parameter adjustment, and the test set is used for the final model evaluation. Some features of the sample data set are binarized, with 0 representing no and 1 representing yes. An example of a sample data set is shown in Table 4 below:

[0115] Table 4

[0116]

[0117] Label the positive and negative samples, and set the activation function and loss function of the constructed deep learning model.

[0118] The purpose of using the excitation function is to transform the linear input of the model into a nonlinear input. The softplus function solves the problem of the high saturation rate of the sigmoid function; at the same time, it improves the problem of the high probability of gradient vanishing of the tanh function; and the derivative calculation speed is fast, which can improve the computational efficiency of the model and converge faster; the softplus function can also make the output value of some intermediate layers zero, improve the sparsity of the model, and reduce the probability of overfitting; in addition, compared with the sigmoid function and the tanh function, which have the problem of gradient vanishing when calculating on time series data, the softplus function has better support for time series data. Therefore, in this embodiment, the excitation function uses the softplus function, and the expression is as follows:

[0119] Softplusζ(x)=ln(1+exp(x)).

[0120] When applying the activation function, set y_softplus = F.softplus(x).data.numpy().

[0121] The loss function (objective function) is a function that calculates the deviation between the predicted value obtained by the model and the true value. In this embodiment, the loss function uses the MAPE (absolute mean percentage deviation) function. If the calculation result of the loss function is small, it means that the model fits well, the difference between the predicted value and the true value is small, and the model performance is excellent; if the calculation result of the loss function is large, it means that the model needs to be adjusted and the parameters of the model need to be reset.

[0122] When applying the loss function, set mape=sum(np.abs((y_true-y_pred) / y_true)) / n*100.

[0123] The labeled training set is input into the deep learning model for iterative training, and the fraud probability value is output.

[0124] When the iteration step of the deep learning model reaches the set value, the deviation between the current fraud probability value of the deep learning model and the true value is calculated to determine whether the deviation is reduced to the preset deviation threshold.

[0125] When the deviation is reduced to the deviation threshold, the current deep learning model is output as the trained deep learning model.

[0126] When the deviation is not reduced to the deviation threshold, the deep learning model is iteratively trained and the number of deviation calculations is counted.

[0127] If the deviation is not equal to the deviation threshold, continue to determine whether the number of calculations reaches a preset number threshold.

[0128] When the number of calculations reaches the threshold, the current deep learning model is output as the trained deep learning model.

[0129] In one embodiment, the step of training the deep learning model further includes:

[0130] Setting the optimization function of the deep learning model to a gradient descent function, performing optimization training on the deep learning model, and iteratively increasing the number of layers and the number of nodes in the middle layer;

[0131] Until the loss function reaches a minimum value, the current deep learning model is output as the trained deep learning model. At this time, the middle layer is 3 and the number of nodes in the middle layer is 6.

[0132] In this embodiment, the optimization function of the deep learning model uses a gradient descent function. By searching in the direction with the largest function slope, a local extreme value can be obtained. Under the premise that the data to be processed is a convex set, by searching in the direction with the largest function slope, a global extreme value can be obtained.

[0133] Compared with the Newton method and the quasi-Newton method, the gradient descent method has better computational performance for time series data.

[0134] When applying the optimization function, set optimizer = torch.optim.SGD(model.parameters(),lr = learning_rate).

[0135] The middle layer is a key link in neural network calculations. The number of middle layers will have an important impact on the accuracy of model data feature recognition and model prediction, so they need to be tuned separately.

[0136] The parameters of the model's middle layer are optimized by seeking the local extreme value of the loss function. The test set is input into the deep learning model in the form of a feature vector, and the binary prediction probability and result of whether it is a high-risk mobile phone model involved in fraud is output, and the loss value calculated by the loss function is passed back to the initial input layer of the deep learning model to complete the long-term memory transfer of the neural network model, and the number of layers and nodes of the middle layer are iteratively increased. The loss function is used each time to evaluate the performance of the model at this time to determine the number of layers and nodes of the middle layer, as shown in Tables 5 and 6 below.

[0137] Table 5

[0138] Middle Layers MAE MAPE RMSE 1 9.33 0.247 10.17 2 8.48 0.185 9.75 3 8.35 0.174 9.59 4 8.86 0.215 10.18 5 9.18 0.228 10.56 6 9.49 0.258 10.86

[0139] It can be observed in Table 5 that the number of intermediate layers is first set to 1 by default, at which time the MAPE value is 0.247, the MAE value is 9.33, and the RMSE value is 10.17. When the number of intermediate layers is set to 3, the loss function reaches a minimum value, at which time the MAPE is 0.174, the MAE is 8.35, and the RMSE is 9.59. When the number of intermediate layers is increased, the MAPE and other test values ​​begin to increase, proving that when the number of intermediate layers is set to 3, the loss function reaches a minimum value.

[0140] Table 6

[0141] Middle Layers MAE MAPE RMSE 1 8.87 0.156 9.78 2 7.82 0.123 9.16 3 7.67 0.189 8.97 4 7.78 0.095 8.89 5 6.78 0.097 8.54 6 6.47 0.058 8.07 7 7.67 0.067 8.64 8 7.56 0.095 8.92

[0142] It can be observed in Table 6 that the MAPE value is the minimum when the number of nodes is set to 6, and the calculation results of the MAE and RMSE functions are also the minimum values. Therefore, it is judged that when the number of nodes is 6, the model performance is locally optimal.

[0143] Therefore, the deep learning model is optimized and trained through the optimization function, and the number of layers and the number of nodes in the middle layer are iteratively increased until the loss function reaches a local minimum. At this time, the number of middle layers is 3 and the number of nodes in the middle layer is 6. The module parameters Num_layers and Hidden_size are updated, and the current deep learning model is output as the trained deep learning model.

[0144] The optimized deep learning model is verified, and the collected positive and negative sample data are input into the model to predict whether it is a high-risk mobile phone model involved in fraud. The predicted value is calculated and compared with the true value to obtain the accuracy and recall rate, as shown in Table 7 below.

[0145] Table 7

[0146]

[0147] The predicted probability of whether each model is a high-risk mobile phone model involved in fraud is shown in Table 8 below.

[0148] Table 8

[0149]

[0150] The data analysis in Tables 7 and 8 proves that the prediction results of the deep learning model in this embodiment are consistent with the actual situation and have high prediction accuracy.

[0151] Furthermore, the terminal device model features are input into the trained deep learning model in feature vector format, and the target user's terminal device model is output as to whether it is fraudulent and the probability of fraud.

[0152] Then, based on the fraud probability value, determine the credibility score of the target user's terminal device model.

[0153] In order to facilitate the intuitive understanding of the probability value of the fraudulent user by the business personnel in the actual situation, whether the model output by the model is a high-risk mobile phone model involved in fraud and the probability value of the high-risk mobile phone model involved in fraud are converted into a model credibility score. In one embodiment, according to the fraud probability value, the step of determining the credibility score of the terminal device model of the target user includes:

[0154] Based on the fraud probability value and the normal user probability value, an evaluation model is constructed;

[0155] Wherein, the expression of the evaluation model includes:

[0156] score=A+B×ln(odds)

[0157] Where, odds is the ratio of the fraud probability value to the normal user probability value, p0 is the score corresponding to a certain probability value, and PDO is the score that increases when a certain probability value doubles;

[0158] The fraud probability value is substituted into the evaluation model to obtain the credibility score of the terminal device model of the target user.

[0159] Specifically, the underlying layer of the model credibility score is a binary classification problem. Through historical data, various characteristics of whether the model is a high-risk mobile phone model involved in fraud are used as dependent variables, and a mathematical model is established to predict the probability of whether the model is a high-risk mobile phone model involved in fraud.

[0160] Through the linear transformation relationship, the ratio between the probability of fraudulent users P and the probability of normal users 1-P is established with the evaluation model to determine the model credibility score.

[0161] in,

[0162] The evaluation model can be represented by an expression involving odds, which is defined as follows:

[0163] score=A+B×ln(odds)

[0164] Among them, A and B are constants that can be calculated by setting the score scale. One is the corresponding score p0 when a certain probability is x; the other is the score PDO that needs to be increased when the probability doubles to 2x. Substituting p0 and PDO into the evaluation model, we get:

[0165]

[0166] Solving the simultaneous equations, we can obtain:

[0167]

[0168] The prediction results of the deep learning model are then substituted into the evaluation model to obtain the credibility score of the target user's terminal device model.

[0169] Finally, determine whether the credibility score meets the fraud-related conditions, such as a score above 650 for low risk, 500-650 for medium risk, and below 500 for high risk. If the credibility score meets the fraud-related conditions, the target user is determined to be a fraud-related user. For models with lower scores, it is determined that there is a high probability of fraud; for models with higher scores, it is determined that there is a high probability of normal models.

[0170] For example, according to the actual situation of the data, take p0=500, PDO=100, odds=0.5, that is, the probability of predicting that it is a non-high-risk fraud model is 0.33.

[0171] According to the calculation, parameter A is approximately equal to 6000 and parameter B is approximately equal to 144.

[0172] Through the values ​​of A and B, the predicted probability value of the deep learning model can be linearly transformed into the model credibility score value. Among them, after the linear transformation of the prediction results, the highest score is 765 and the lowest score is 421.

[0173] In this embodiment, the credibility score of the model is divided into a range of 350 to 850 points, and the credibility score is divided into 8 equally spaced intervals. The fraud-related situation in each interval is calculated, and the results of the evaluation model are observed, as shown in Table 9 below.

[0174] Table 9

[0175]

[0176] It can be seen from Table 7 that as the score increases, the probability that the model is a normal model is also increasing, which is consistent with expectations and verifies that the evaluation model has a better effect. When the credibility score of a model reaches 650 points or more, it can be considered that it has a higher probability of being a normal model. At the same time, for models with lower scores below 500 points, most of them are fraudulent model samples, indicating that when a model has a low credibility score, it can basically be determined that there is a high probability of fraud. Models with middle scores also have a high probability of being normal models. When the score of a model is in this range, it is often difficult to judge its value, and the anti-fraud team staff need to make manual judgments. This is where the evaluation model needs to be strengthened. From the overall proportion of the number of samples, the credibility scores of most models are still distributed in a higher area, indicating that there are still a few fraudulent models, and most models are normal models, which is also consistent with the actual situation.

[0177] In one embodiment, different handling methods can be used for models with different risk levels, which helps to reduce the pressure on customer service to handle complaints and achieve more refined management and control of fraud governance.

[0178] In summary, a method for identifying fraudulent users based on a deep learning algorithm obtains the terminal device model of the target user, and then extracts the terminal device model features corresponding to the target user, so as to evaluate the fraud risk of the target user from a new perspective of the terminal brand model; the terminal device model features are input into a pre-trained deep learning model in a feature vector format, and whether the terminal device model of the target user is fraudulent and the probability value of fraud are output. This method can fully utilize the advantages of the deep learning algorithm to improve the accuracy of the assessment of fraudulent users, which is conducive to more accurate governance of telecommunications network fraud; according to the probability value of fraud, the credibility score of the terminal device model of the target user is determined to determine whether the fraud conditions are met, so as to realize the conversion of the probability value of the fraudulent terminal device model output by the model into an easy-to-understand model credibility score, which is intuitive and clear, and is conducive to rapid and comprehensive promotion in the communications industry.

[0179] A method for identifying fraudulent users based on deep learning algorithms mines fraudulent features from the perspective of terminal device models, and constructs accurate and comprehensive model characteristics of fraudulent users and normal users. In addition to traditional user behavior characteristics and user information characteristics, it provides a new dimension for risk assessment of fraudulent users, providing support for the assessment of fraudulent users.

[0180] A method for identifying users involved in fraud based on a deep learning algorithm is based on a deep learning algorithm. The model features, thresholds, weights, etc. used in calculating the credibility score of the terminal device model are all learned using a deep learning algorithm to obtain the optimal solution. Finally, a good performance evaluation model is trained to make the evaluation results objective, avoiding the limitations brought by human subjective cognition, and greatly improving the accuracy of identifying users involved in fraud.

[0181] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0182] The embodiment of the present application also provides a fraudulent user identification device based on a deep learning algorithm, which corresponds one-to-one to a fraudulent user identification method based on a deep learning algorithm in the above embodiment. The fraudulent user identification device based on a deep learning algorithm includes:

[0183] The device model module is used to obtain the terminal device model of the target user;

[0184] A feature module, used to extract a terminal device model feature corresponding to the target user based on the terminal device model;

[0185] A calculation module, used to input the terminal device model features into a pre-trained deep learning model in a feature vector format, and output whether the terminal device model of the target user is fraudulent and the probability value of fraudulent;

[0186] A mapping module, used to determine the credibility score of the terminal device model of the target user according to the fraud probability value;

[0187] A detection module, used to determine whether the credibility score meets the fraud-related condition;

[0188] The identification module is used to determine that the target user is a fraudulent user when the credibility score meets the fraudulent condition.

[0189] For the specific limitations of a fraudulent user identification device based on a deep learning algorithm, please refer to the limitations of a fraudulent user identification method based on a deep learning algorithm mentioned above, which will not be repeated here.

[0190] Each module in the above-mentioned fraudulent user identification device based on deep learning algorithm can be implemented in whole or in part by software, hardware and their combination. Each of the above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the corresponding operations of each of the above modules.

[0191] In one embodiment, a computer device is provided, which may be a server. The computer device includes a processor, a memory, a network interface, and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, any of the above-mentioned fraud-related user identification methods based on a deep learning algorithm is implemented.

[0192] In one embodiment, a computer-readable storage medium is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, any one of the above-mentioned methods for identifying fraudulent users based on a deep learning algorithm is implemented.

[0193] In one embodiment, a computer program product is provided, which includes a computer program, and when the computer program is executed by a processor, it implements any of the above-mentioned fraudulent user identification methods based on deep learning algorithms.

[0194] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present application. When the computer program is executed, it may include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in the present application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0195] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above.

Claims

1. A method for identifying fraudulent users based on a deep learning algorithm, characterized in that: The following steps are included: Obtain the terminal device model of the target user; Based on the terminal device model, extracting terminal device model features corresponding to the target user; Input the terminal device model features into a pre-trained deep learning model in a feature vector format, and output whether the terminal device model of the target user is involved in fraud and the probability value of fraud; Determining the credibility score of the terminal device model of the target user according to the fraud probability value; Determining whether the credibility score satisfies the fraud-related condition; If the credibility score meets the fraud-related condition, the target user is determined to be a fraud-related user.

2. The method for identifying fraudulent users based on deep learning algorithm according to claim 1 is characterized in that: The steps of constructing the deep learning model include: Based on the long short-term memory temporal neural network model, the model's input vector dimension is set to 18, the middle layer is 1, the number of nodes in the middle layer is 1, and the iteration step is 150; and The model is set to use a nonlinear activation function and a non-bidirectional neural network model.

3. The method for identifying fraudulent users based on deep learning algorithm according to claim 2 is characterized in that: The training steps of the deep learning model include: Extract the corresponding terminal device model historical features from the signaling history data of the fraudulent user and construct a sample data set; Dividing the sample data set into a training set and a test set in proportion; Based on the training set, combined with a preset excitation function and a loss function, the deep learning model is iteratively trained to output a fraud probability value, wherein the excitation function is a softplus function and the loss function is a MAPE function; When the iteration step of the deep learning model reaches a set value, the deviation between the current fraud probability value of the deep learning model and the true value is calculated to determine whether the deviation is reduced to a preset deviation threshold; When the deviation is reduced to the deviation threshold, the current deep learning model is output as the trained deep learning model.

4. The method for identifying fraudulent users based on deep learning algorithm according to claim 3 is characterized in that: The following steps are also included: When the deviation is not reduced to the deviation threshold, repeatedly iteratively training the deep learning model, and counting the number of calculations of the deviation; If the deviation is not equal to the deviation threshold, determining whether the number of calculations reaches a preset number threshold; When the number of calculations reaches the number threshold, the current deep learning model is output as a trained deep learning model.

5. The method for identifying fraudulent users based on deep learning algorithm according to claim 3 is characterized in that: The training step of the deep learning model also includes: Setting the optimization function of the deep learning model to a gradient descent function, performing optimization training on the deep learning model, and iteratively increasing the number of layers and the number of nodes in the middle layer; Until the loss function reaches a minimum value, the current deep learning model is output as the trained deep learning model. At this time, the middle layer is 3 and the number of nodes in the middle layer is 6.

6. The method for identifying fraudulent users based on a deep learning algorithm according to any one of claims 1 to 5, characterized in that: The step of determining the credibility score of the terminal device model of the target user according to the fraud probability value includes: Based on the fraud probability value and the normal user probability value, an evaluation model is constructed; Wherein, the expression of the evaluation model includes: score=A+B×ln(odds) Where, odds is the ratio of the fraud probability value to the normal user probability value, p0 is the score corresponding to a certain probability value, and PDO is the score that increases when a certain probability value doubles; The fraud probability value is substituted into the evaluation model to obtain the credibility score of the terminal device model of the target user.

7. A fraudulent user identification device based on a deep learning algorithm, characterized in that: include, The device model module is used to obtain the terminal device model of the target user; A feature module, used to extract a terminal device model feature corresponding to the target user based on the terminal device model; A calculation module, used to input the terminal device model features into a pre-trained deep learning model in a feature vector format, and output whether the terminal device model of the target user is fraudulent and the probability value of fraudulent; A mapping module, used to determine the credibility score of the terminal device model of the target user according to the fraud probability value; A detection module, used to determine whether the credibility score meets the fraud-related condition; The identification module is used to determine that the target user is a fraudulent user when the credibility score meets the fraudulent condition.

8. A computer device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product, characterized in that The invention comprises a computer program, which implements the steps of the method according to any one of claims 1 to 6 when being executed by a processor.