Number recognition method, device and equipment and computer readable storage medium

By acquiring user characteristics of numbers suspected of being fraudulent, and using probabilistic prediction models and environmental indicators to determine policy levels, the problem of low accuracy in identifying fraudulent numbers by operators has been solved, achieving higher processing accuracy.

CN115730211BActive Publication Date: 2025-11-28CHINA MOBILE GROUP ZHEJIANG +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111018471.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-31
Publication Date
2025-11-28
Estimated Expiration
2041-08-31

AI Technical Summary

Technical Problem

In existing technologies, operators have low accuracy in identifying fraudulent numbers, leading to mishandling of legitimate users and affecting their daily use.

Method used

By obtaining user characteristics of numbers involved in fraud, inputting them into a preset probability prediction model, obtaining the probability of reactivation, and combining environmental indicators to determine the policy level, the processing strategy is determined based on the probability of reactivation and the policy level, including shutdown, call restriction, or secondary real-name authentication.

Benefits of technology

It improves the accuracy of handling fraudulent numbers, ensuring that numbers with different probabilities of reactivation and environmental levels are handled in a targeted manner, reducing mishandling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115730211B_ABST
    Figure CN115730211B_ABST
Patent Text Reader

Abstract

The application discloses a number identification method, device and equipment and a computer readable storage medium. The number identification method comprises the following steps: obtaining a fraud-related number, and extracting user features of the fraud-related number; inputting the user features into a preset probability prediction model to obtain a computer simulation probability of the fraud-related number; obtaining an environment index and determining a strategy level according to the environment index; and determining a processing strategy of the fraud-related number according to the computer simulation probability and the strategy level. The application improves the accuracy of fraud-related number processing.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, and particularly relates to a number identification method, device and equipment and a computer readable storage medium. BACKGROUND

[0002] With the gradual improvement of the fraud control system of the operator, while increasing the efforts to combat communication fraud, the amount of misjudgment will inevitably increase, resulting in misdisposal behavior to normal users, which affects the daily use of normal users. At present, the operator mainly develops a fraud model through the analysis of the case-involved number and the experience of business personnel, and the identification accuracy of the fraud number is low, resulting in low disposal accuracy of the user number. SUMMARY

[0003] The main purpose of the present application is to provide a number identification method, device and equipment and a computer readable storage medium, which aims to solve the problem of how to improve the processing accuracy of the fraud number.

[0004] To achieve the above purpose, the present application provides a number identification method, which comprises the following steps:

[0005] Obtaining a fraud number, extracting user features of the fraud number;

[0006] Inputting the user features into a preset probability prediction model to obtain a re-machine probability of the fraud number;

[0007] Obtaining an environmental index and determining a strategy level according to the environmental index;

[0008] Determining a processing strategy of the fraud number according to the re-machine probability and the strategy level.

[0009] In an embodiment, the step of obtaining an environmental index and determining a strategy level according to the environmental index comprises:

[0010] Obtaining the environmental index, wherein the environmental index comprises a case-involved index, a basic index and a business index, the case-involved index at least comprises the number of fraud numbers within a preset time length, the basic index at least comprises a data coverage, and the business index at least comprises a roaming day number, a terminal switching number and / or a number channel;

[0011] Determining the strategy level according to the case-involved index, the basic index and the business index.

[0012] In an embodiment, the step of determining the strategy level according to the case-involved index, the basic index and the business index comprises:

[0013] determine a first strategy level of the identification strategy according to the involved index;

[0014] if the first strategy level is a preset level, taking the first strategy level as a strategy level of the identification strategy;

[0015] if the first strategy level is not a preset level, determining a second strategy level of the identification strategy according to the basic index and the business index, and taking the second strategy level as a strategy level of the identification strategy.

[0016] In an embodiment, the step of determining the processing strategy of the fraud number according to the machine probability and the strategy level comprises:

[0017] determining a probability threshold according to the strategy level;

[0018] judging whether the machine probability is less than the probability threshold;

[0019] when the machine probability of the fraud number is less than the probability threshold, shutting down the fraud number, calling limiting the fraud number, or secondary real-name authentication of the fraud number.

[0020] In an embodiment, before the step of inputting the user features into a preset probability prediction model, the method further comprises:

[0021] obtaining training samples, the training samples comprising shut-down machine numbers and non-shut-down machine numbers;

[0022] extracting target user features of the training samples;

[0023] training a preset neural network model according to the target user features;

[0024] taking the converged neural network model as a probability prediction model.

[0025] In an embodiment, the step of extracting the target user features of the training samples comprises:

[0026] extracting to-be-determined user features of the training samples, the number of the to-be-determined user features being greater than the number of the target user features;

[0027] determining feature importance and information value IV value of the to-be-determined user features according to a preset algorithm, the feature importance being determined by a random forest algorithm;

[0028] determining target user features from the to-be-determined user features according to the feature importance and the information value.

[0029] In an embodiment, the step of obtaining the fraud number comprises:

[0030] obtain a plurality of user numbers to be identified, and extract features of the user numbers;

[0031] input the features of the user numbers into a preset number identification model to obtain the fraud-related numbers.

[0032] To achieve the above object, the present application further provides a number identification device, which comprises:

[0033] an obtaining module, configured to obtain fraud-related numbers and extract user features of the fraud-related numbers;

[0034] a prediction module, configured to input the user features into a preset probability prediction model to obtain a probability of re-use of the fraud-related numbers;

[0035] a determination module, configured to obtain an environmental index and determine a strategy level according to the environmental index;

[0036] a processing module, configured to determine a processing strategy for the fraud-related numbers according to the probability of re-use and the strategy level.

[0037] To achieve the above object, the present application further provides a number identification device, which comprises a memory, a processor, and a number identification program stored in the memory and executable on the processor, and the number identification program, when executed by the processor, implements each step of the number identification method as described above.

[0038] To achieve the above object, the present application further provides a computer readable storage medium, which stores a number identification program, and the number identification program, when executed by a processor, implements each step of the number identification method as described above.

[0039] The number identification method, device, equipment and computer readable storage medium provided by the present application obtain fraud-related numbers and extract user features of the fraud-related numbers; input the user features into a preset probability prediction model to obtain a probability of re-use of the fraud-related numbers; obtain an environmental index and determine a strategy level according to the environmental index; and determine a processing strategy for the fraud-related numbers according to the probability of re-use and the strategy level. By determining the strategy level and the probability of re-use to set the processing strategy for the fraud-related numbers, fraud-related numbers with different probabilities of re-use are processed according to the strategy level, and fraud-related numbers with a high probability of re-use predicted by the probability prediction model are also processed when the strategy level is severe. The processing of the fraud-related numbers is determined according to the strategy level and the predicted probability of re-use, thereby improving the accuracy of the processing of the fraud-related numbers. BRIEF DESCRIPTION OF DRAWINGS

[0040] Figure 1A hardware structure schematic diagram of the number identification device involved in the embodiment of the present application is shown in the figure.

[0041] Figure 2 A flowchart of the first embodiment of the number identification method of the present application is shown in the figure.

[0042] Figure 3 A detailed flowchart of step S30 of the second embodiment of the number identification method of the present application is shown in the figure.

[0043] Figure 4 A detailed flowchart of step S30 of the third embodiment of the number identification method of the present application is shown in the figure.

[0044] Figure 5 A flowchart of the fourth embodiment of the number identification method of the present application is shown in the figure.

[0045] Figure 6 A logic structure schematic diagram of the number identification device of the present application is shown in the figure.

[0046] The implementation, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION

[0047] It should be understood that the specific embodiments described herein are merely intended to explain the present application and not to limit the present application.

[0048] The main solution of the embodiment of the present application is to obtain a fraud-involved number and extract user features of the fraud-involved number; input the user features into a preset probability prediction model to obtain a re-computer probability of the fraud-involved number; obtain an environment index and determine a strategy level according to the environment index; and determine a processing strategy of the fraud-involved number according to the re-computer probability and the strategy level.

[0049] The processing strategy of the fraud-involved number is set by determining the strategy level and the re-computer probability, and the fraud-involved number with different re-computer probabilities is processed according to the strategy level, so that the fraud-involved number is also processed when the re-computer probability predicted by the probability prediction model is high and the strategy level is severe, the processing of the fraud-involved number is determined according to the strategy level and the predicted re-computer probability, and the accuracy of the processing of the fraud-involved number is improved.

[0050] As an implementation scheme, the number identification device can be as shown in the figure. Figure 1

[0051] The embodiment of the present application relates to a number identification device, which comprises a processor 101 such as a CPU, a memory 102 and a communication bus 103. The communication bus 103 is used to realize the connection and communication between the components.

[0052] ​The memory 102 can be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. As shown in Figure 1 the memory 102 as a computer readable storage medium can include a number identification program; and the processor 101 can be configured to call the number identification program stored in the memory 102 and perform the following operations:

[0053] obtain a fraud number, and extract user features of the fraud number;

[0054] input the user features into a preset probability prediction model to obtain a re-computer probability of the fraud number;

[0055] obtain an environment index and determine a strategy level according to the environment index;

[0056] determine a processing strategy of the fraud number according to the re-computer probability and the strategy level.

[0057] In an embodiment, the processor 101 can be configured to call the number identification program stored in the memory 102 and perform the following operations:

[0058] obtain the environment index, wherein the environment index includes a case-related index, a basic index, and a business index, the case-related index at least includes a number of fraud numbers in a preset time period, the basic index at least includes a data coverage, and the business index at least includes a roaming day number, a terminal switching number, and / or a number channel;

[0059] determine the strategy level according to the case-related index, the basic index, and the business index.

[0060] In an embodiment, the processor 101 can be configured to call the number identification program stored in the memory 102 and perform the following operations:

[0061] determine a first strategy level of an identification strategy according to the case-related index;

[0062] if the first strategy level is a preset level, the first strategy level is taken as a strategy level of the identification strategy;

[0063] if the first strategy level is not the preset level, a second strategy level of the identification strategy is determined according to the basic index and the business index, and the second strategy level is taken as the strategy level of the identification strategy.

[0064] In an embodiment, the processor 101 can be configured to call the number identification program stored in the memory 102 and perform the following operations:

[0065] determine a probability threshold value according to the strategy level;

[0066] determining whether the probability of reconnection is less than the probability threshold value;

[0067] when the probability of reconnection of the fraud-related number is less than the probability threshold value, shutting down the fraud-related number, calling limiting the fraud-related number, or secondary real-name authentication of the fraud-related number.

[0068] In an embodiment, the processor 101 can be configured to call the number identification program stored in the memory 102 and perform the following operations:

[0069] obtaining training samples, the training samples including numbers that are reconnected after being shut down and numbers that are not reconnected after being shut down;

[0070] extracting target user features of the training samples;

[0071] training a preset neural network model according to the target user features;

[0072] using the converged neural network model as a probability prediction model.

[0073] In an embodiment, the processor 101 can be configured to call the number identification program stored in the memory 102 and perform the following operations:

[0074] extracting to-be-determined user features of the training samples, the number of to-be-determined user features being greater than the number of target user features;

[0075] determining feature importance and information value IV value of the to-be-determined user features according to a preset algorithm, the feature importance being determined by a random forest algorithm;

[0076] determining target user features from the to-be-determined user features according to the feature importance and the information value.

[0077] In an embodiment, the processor 101 can be configured to call the number identification program stored in the memory 102 and perform the following operations:

[0078] obtaining a plurality of to-be-identified user numbers and extracting features of the user numbers;

[0079] inputting the features of the user numbers into a preset number identification model to obtain the fraud-related number.

[0080] Based on the hardware architecture of the number identification device, embodiments of the number identification method are proposed.

[0081] Reference Figure 2 , Figure 2 For the first embodiment of the number identification method, the number identification method comprises the following steps:

[0082] Step S10, obtain a fraud-related number, and extract user features of the fraud-related number;

[0083] Specifically, the fraud-related number is a user number related to fraud identified by a model or manually screened. Since the model identification depends on the selection of the training data of the model, or the manual screening depends on the experience of the human, the model identification and the manual screening are both inaccurate, resulting in that the screened fraud-related number may be a non-fraud-related number. Therefore, the screened fraud-related number needs to be further processed.

[0084] The user features of the fraud-related number are extracted, wherein the user features can be user basic information, communication behavior, roaming behavior, and consumption behavior, etc. The fraud-related number can be obtained by model identification. A plurality of user numbers to be identified are obtained, and the features of the user numbers are extracted. The features of the user numbers are input into a preset number identification model to obtain the fraud-related number. The number identification model can be an identification model established by different rules, which can be a communication rule identification model, a channel rule identification model, a business rule identification model, a regional rule identification model, and / or a terminal rule identification model.

[0085] Step S20, input the user features into a preset probability prediction model to obtain a reconnection probability of the fraud-related number;

[0086] Specifically, the user features are input into the preset probability prediction model to obtain the reconnection probability of the fraud-related number predicted by the probability prediction model, wherein the reconnection probability is the probability of the user applying for reconnection after the fraud-related number is suspended.

[0087] Step S30, obtain an environmental index and determine a strategy level according to the environmental index;

[0088] Specifically, the environmental index is obtained, and the environmental index can include the number of fraud-related numbers in a period of time. The strategy level can be divided into a plurality of levels. For example, the strategy level can be divided into a severe mode, a general mode, and a relaxed mode. If the number of fraud-related numbers in March of each year is the highest in the whole year, the strategy level of the fraud-related number is set to the severe mode in March of each year. If the number of fraud-related numbers in November of each year is the lowest in the whole year, the strategy level of the fraud-related number is set to the relaxed mode in November of each year.

[0089] Step S40, determine a processing strategy of the fraud-related number according to the reconnection probability and the strategy level.

[0090] Specifically, the processing strategy of the fraud-related number is determined according to the re-machine probability and the policy level. The processing strategy can include that when the re-machine probability of the fraud-related number is less than a probability threshold, the fraud-related number is determined as an abnormal number, and the fraud-related number is processed. The processing manner can be shutting down the fraud-related number, call restriction, or secondary real-name authentication. When the re-machine probability of the fraud-related number is greater than or equal to the probability threshold, the fraud-related number is determined as a normal number, wherein the probability threshold is determined according to the policy level.

[0091] The processing strategy can also include determining the accuracy rate according to the policy level, and determining the fraud-related number according to the accuracy rate and the re-machine probability. For example, when the accuracy rate is 90%, the re-machine probability multiplied by the accuracy rate 90% equals the final re-machine probability of the fraud-related number, and whether the fraud-related number is processed is determined according to the final re-machine probability. The processing manner can be shutting down the fraud-related number, call restriction, or secondary real-name authentication.

[0092] In the technical scheme of the embodiment, the fraud-related number is obtained, and the user features of the fraud-related number are extracted. The user features are input into a preset probability prediction model to obtain the re-machine probability of the fraud-related number. The environment indicators are obtained, and the policy level is determined according to the environment indicators. The processing strategy of the fraud-related number is determined according to the re-machine probability and the policy level. The processing strategy of the fraud-related number is set by determining the policy level and the re-machine probability. According to the policy level, the fraud-related numbers with different re-machine probabilities are processed. When the policy level is severe, the fraud-related number will be processed even if the re-machine probability predicted by the probability prediction model is high. The processing of the fraud-related number is determined according to the policy level and the predicted re-machine probability, thereby improving the accuracy of the processing of the fraud-related number.

[0093] Reference Figure 3 , Figure 3 For the second embodiment of the number recognition method, based on the first embodiment, the step S30 includes:

[0094] In step S31, the environment indicators are obtained, the environment indicators include case-related indicators, basic indicators, and business indicators. The case-related indicators at least include the number of fraud-related numbers in a preset time period. The basic indicators at least include data coverage. The business indicators at least include roaming days, terminal switching times, and / or number channels.

[0095] In step S32, the policy level is determined according to the case-related indicators, the basic indicators, and the business indicators.

[0096] Specifically, the environment index represents data of the current network environment of the fraud number, and the environment index includes a case-related index, a basic index, and a business index. The case-related index at least includes the number of fraud numbers within a preset time period, such as the fraud numbers reported by the Ministry of Public Security and the fraud numbers reported by the Ministry of Industry and Information Technology or 10086. The basic index at least includes data coverage, and the fraud numbers reported by the Ministry of Public Security and the fraud numbers reported by the Ministry of Industry and Information Technology or 10086 are used as the basis to determine the coverage, timeliness, or case type statistics of the fraud numbers, and statistics such as same period, same period, average, etc. of each time period, such as the fraud number ring index reported by the case report in the last 7 days. The business index at least includes roaming days, terminal switching times, and / or number channels; it can also include business usage, such as silent days this month, regional conditions, such as provincial roaming days this month, communication conditions, such as non-family network calling times in the last 3 days, channel conditions, such as online channel number release identification, terminal conditions, such as terminal switching times this month, and statistics such as same period, same period, average, etc. of each time period, such as non-family network calling times ring index.

[0097] The strategy level can be determined according to the case-related index, the basic index, and the business index.

[0098] The first strategy level of the identification strategy can be determined according to the case-related index; if the first strategy level is a preset level, the first strategy level is used as the strategy level of the identification strategy; if the first strategy level is not a preset level, a second strategy level of the identification strategy is determined according to the basic index and the business index, and the second strategy level is used as the strategy level of the identification strategy.

[0099] In the technical solution of the embodiment, the environment index is obtained, the environment index includes the case-related index, the basic index, and the business index, and the strategy level is determined according to the case-related index, the basic index, and the business index. The processing strategy for the fraud number is determined according to the environment index, and the fraud numbers with different machine recovery probabilities are processed according to the strategy level, thereby improving the accuracy of fraud number processing.

[0100] Reference Figure 4 , Figure 4 For the third embodiment of the number identification method of the application, based on the first or second embodiment, the step S30 includes:

[0101] Step S33, determining a probability threshold according to the strategy level;

[0102] Step S34, determining whether the machine recovery probability is less than the probability threshold;

[0103] Step S35, when the reinstallation probability of the fraud number is less than the probability threshold, shutting down the fraud number, calling limiting the fraud number, or secondary real-name authentication of the fraud number.

[0104] Specifically, when the reinstallation probability of the fraud number is less than the probability threshold, the fraud number is determined as an abnormal number, and the fraud number is processed by shutting down, calling limiting, or secondary real-name authentication. When the reinstallation probability of the fraud number is greater than or equal to the probability threshold, the fraud number is determined as a normal number, and the fraud number is not processed. The probability threshold is determined according to the policy level. The policy level in different regions may be different. In a region with a high fraud probability, the policy level is a severe mode. In a region with a low fraud probability, the policy level is a lenient mode. In a region with a general fraud probability, the policy level is a general mode.

[0105] For example, the predicted reinstallation probability and the actual reinstallation probability of the fraud number are shown in the following table:

[0106]

[0107] From the above table, it can be seen that as the fraud probability value increases, the predicted reinstallation probability and the actual reinstallation probability of the fraud number gradually decrease.

[0108] For example, when the policy level is a severe mode, the probability threshold is 0.9. When the reinstallation probability of the fraud number is less than 0.9, the fraud number is determined as an abnormal number, and the fraud number is processed by shutting down, calling limiting, or secondary real-name authentication. When the reinstallation probability of the fraud number is greater than or equal to 0.9, the fraud number is determined as a normal number. When the policy level is a general mode, the probability threshold is 0.5. When the reinstallation probability of the fraud number is less than 0.5, the fraud number is determined as an abnormal number, and the fraud number is processed by shutting down, calling limiting, or secondary real-name authentication. When the reinstallation probability of the fraud number is greater than or equal to 0.5, the fraud number is determined as a normal number. When the policy level is a lenient mode, the probability threshold is 0.1. When the reinstallation probability of the fraud number is less than 0.1, the fraud number is determined as an abnormal number, and the fraud number is processed by shutting down, calling limiting, or secondary real-name authentication. When the reinstallation probability of the fraud number is greater than or equal to 0.1, the fraud number is determined as a normal number.

[0109] In the technical solution of the embodiment, when the reinstallation probability of the fraud number is less than the probability threshold, the fraud number is determined as an abnormal number; when the reinstallation probability of the fraud number is greater than or equal to the probability threshold, the fraud number is determined as a normal number. The probability threshold is determined according to the policy level. According to the policy level, different fraud numbers with different reinstallation probabilities are processed, which improves the accuracy of processing fraud numbers in different situations such as different times or regions.

[0110] Referring toFigure 5 , Figure 5 For the fourth embodiment of the number identification method of the application, based on any one of the first to third embodiments, before the step S20, further comprising:

[0111] Step S50, obtaining training samples, the training samples including the numbers of the machines restarted after shutdown and the numbers of the machines not restarted after shutdown;

[0112] Step S60, extracting target user features of the training samples;

[0113] Step S70, training a preset neural network model according to the target user features;

[0114] Step S80, taking the converged neural network model as a probability prediction model.

[0115] Specifically, the training samples are obtained, the training samples including the numbers of the machines restarted after shutdown and the numbers of the machines not restarted after shutdown. The target user features of the training samples are extracted, and the preset neural network model is trained according to the target user features. The converged neural network model is taken as a probability prediction model.

[0116] The target user features of the training samples can be extracted first, and the number of the to-be-determined user features of the training samples is greater than the number of the target user features. The feature importance and the information value IV value of the to-be-determined user features are determined according to a preset algorithm, wherein the feature importance is determined by a random forest algorithm. The target user features are determined from the to-be-determined user features according to the feature importance and the information value.

[0117] An XGBoost model with better performance in machine learning is used to predict the probability of the numbers of the machines restarted after the suspected fraud numbers are disposed of. The XGBoost algorithm is a gradient boosting algorithm based on decision trees, which has improved in algorithm applicability, running efficiency and robustness. The XGBoost algorithm iterates multiple times, each iteration produces a weak classifier, and each classifier is trained based on the residual error of the last round of classifiers. The requirements for weak classifiers are generally simple, low variance and high bias. The final total classifier is obtained by weighted summation of the weak classifiers obtained by each round of training. The main process of the XGBoost algorithm is as follows:

[0118] Let the training samples as input be I = {(x1, y1), (x2, y2),..., (xm, ym)}, the sample labels be y(i = 1, 2,..., m). The maximum number of iterations is T, the loss function is L, the regularization coefficient is λ, the weak learner is h(x), and the output is the strong learner f(x). For the tth iteration, the loss function is: m m i y(i = 1, 2,..., m). The maximum number of iterations is T, the loss function is L, the regularization coefficient is λ, the weak learner is h(x), and the output is the strong learner f(x). For the tth iteration, the loss function is:​​

[0119]

[0120] where f t-1 (x i ) is the strong learner output of the t-1th round of training, h i (x i ) is the weak learner of the tth round of training, J is the number of leaf nodes of the tth decision tree, and ω ij is the optimal value of the jth leaf node. For the tth iteration, we have:

[0121] Calculate the i-th sample, where i = 1, 2,..., m; the loss function L t based on f t-1 (x i ) is the first-order derivative g ii , the second-order derivative h ii . Calculate the first-order derivative sum and the second-order derivative sum of all samples:

[0122]

[0123] Based on the current node to split the decision tree, the default score score = 0, G t and H t are the sum of the first and second derivatives of the current node that needs to be split. Let the first and second derivative sums of the left and right child nodes of the current node be G t , H t , G R and H R , and the total number of sample features is N. For the kth sample feature r k (k = 1, 2,..., N):

[0124] G L = 0, H L = 0

[0125] Sort the samples by r k , and calculate the first and second derivative sums of the left and right child nodes of the i-th sample in turn:

[0126] G L = G L + g ti , G R = G - G L

[0127] H L = H L + h ti , H R = H - H L

[0128] Update the score value:

[0129]

[0130] Split the sub-tree based on the split feature and feature value corresponding to the maximum score value. If the maximum score value is 0, the current decision tree is established, and the ω of all leaf regions is calculated ij , to obtain the weak learner h i (x), update the strong learner f i (x), and enter the next round of weak learner iteration. If the maximum score value is not equal to 0, continue to try to split the decision tree.

[0131] A different number recognition model corresponds to a probability prediction model, and the probability prediction model is trained by the training sample of the data type corresponding to the number recognition model. For example, the identification model of the communication type rule corresponds to the probability prediction model of the communication type rule, the identification model of the channel type rule corresponds to the probability prediction model of the channel type rule, the identification model of the business type rule corresponds to the probability prediction model of the business type rule, the identification model of the regional type rule corresponds to the probability prediction model of the regional type rule, and / or the identification model of the terminal type rule corresponds to the probability prediction model of the terminal type rule.

[0132] In the technical scheme of the embodiment, the training sample is obtained, and the target user feature of the training sample is extracted; the preset neural network model is trained according to the target user feature; and the converged neural network model is used as the probability prediction model. The probability prediction model obtained by training is used to predict the re-machine probability of the fraud-related number, thereby improving the efficiency of the fraud-related number processing.

[0133] Referring to Figure 6 , the application further provides a number identification device, which comprises:

[0134] The acquisition module 100 is used to acquire a fraud-related number and extract a user feature of the fraud-related number;

[0135] The prediction module 200 is used to input the user feature into a preset probability prediction model to obtain a re-machine probability of the fraud-related number;

[0136] The determination module 300 is used to acquire an environmental index and determine a strategy level according to the environmental index;

[0137] The processing module 400 is used to determine a processing strategy of the fraud-related number according to the re-machine probability and the strategy level.

[0138] In an embodiment, in terms of acquiring an environmental index and determining a strategy level according to the environmental index, the determination module 300 is specifically used to:

[0139] obtaining the environment indicators, the environment indicators comprising a case-related indicator, a basic indicator, and a business indicator, the case-related indicator comprising at least a number of fraud-related numbers in a preset time period, the basic indicator comprising at least a data coverage, and the business indicator comprising at least a roaming day number, a terminal handover number, and / or a number channel;

[0140] determining the policy level according to the case-related indicator, the basic indicator, and the business indicator.

[0141] In an embodiment, in determining the policy level according to the case-related indicator, the basic indicator, and the business indicator, the determining module 300 is specifically configured to:

[0142] determining a first policy level of an identification policy according to the case-related indicator;

[0143] if the first policy level is a preset level, taking the first policy level as the policy level of the identification policy;

[0144] if the first policy level is not the preset level, determining a second policy level of an identification policy according to the basic indicator and the business indicator, and taking the second policy level as the policy level of the identification policy.

[0145] In an embodiment, in determining the processing policy of the fraud-related number according to the reconnection probability and the policy level, the processing module 400 is specifically configured to:

[0146] determining a probability threshold according to the policy level;

[0147] judging whether the reconnection probability is less than the probability threshold;

[0148] when the reconnection probability of the fraud-related number is less than the probability threshold, shutting down the fraud-related number, calling limiting the fraud-related number, or secondary real-name authentication of the fraud-related number.

[0149] In an embodiment, before inputting the user features into a preset probability prediction model, the prediction module 200 is specifically configured to:

[0150] obtaining training samples, the training samples comprising reconnected numbers after being shut down and non-reconnected numbers after being shut down;

[0151] extracting target user features of the training samples;

[0152] training a preset neural network model according to the target user features;

[0153] taking the converged neural network model as a probability prediction model.

[0154] In an embodiment, in extracting the target user features of the training samples, the prediction module 200 is specifically configured to:

[0155] extracting the to-be-determined user features of the training samples, the number of the to-be-determined user features being greater than the number of the target user features;

[0156] determining the feature importance and the information value IV value of the to-be-determined user features according to a preset algorithm, the feature importance being determined by a random forest algorithm;

[0157] determining the target user features from the to-be-determined user features according to the feature importance and the information value.

[0158] In an embodiment, in obtaining the fraud-related number, the obtaining module 100 is configured to:

[0159] obtaining a plurality of to-be-identified user numbers and extracting features of the user numbers;

[0160] inputting the features of the user numbers into a preset number identification model to obtain the fraud-related number.

[0161] The present application also provides a number identification device, which comprises a memory, a processor, and a number identification program stored in the memory and executable on the processor, and the number identification program, when executed by the processor, implements each step of the number identification method as described in the above embodiments.

[0162] The present application also provides a computer readable storage medium, which stores a number identification program, and the number identification program, when executed by a processor, implements each step of the number identification method as described in the above embodiments.

[0163] The above-mentioned serial numbers of the embodiments of the present application are only for description, and do not represent the advantages or disadvantages of the embodiments.

[0164] It should be noted that, in this document, the terms “comprising”, “including”, or any other variant thereof are intended to cover non-exclusive inclusion, so that processes, systems, articles, or apparatuses that include a series of elements not only include those elements, but also include other elements not explicitly listed, or further include elements inherent to such processes, systems, articles, or apparatuses. Without more limitations, the element defined by the phrase “comprising a” does not exclude the presence of additional identical elements in the process, system, article, or apparatus that includes the element.

[0165] Through the above description of the embodiments, those skilled in the art of the neighborhood can clearly understand that the above-mentioned embodiment system can be realized by means of software and the necessary general hardware platform, of course, it can also be realized by hardware, but in many cases the former is a better embodiment. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a computer readable storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes a plurality of instructions for making a terminal device (which can be a mobile phone, computer, parking management device, air conditioner, or network device, etc.) execute the system described in each embodiment of the present application.

[0166] The above is only the preferred embodiment of the present application, and does not limit the patent scope of the present application, and any equivalent structure or equivalent flow transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied to other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A number recognition method, characterized in that, The number identification method includes: Obtain the phone number suspected of being involved in fraud, and extract the user characteristics of the phone number suspected of being involved in fraud; The user characteristics are input into a preset probability prediction model to obtain the probability of the fraudulent number being called back. Obtain environmental indicators, which include case-related indicators, basic indicators, and business indicators. The case-related indicators include at least the number of fraudulent phone numbers within a preset time period. The basic indicators include at least the data coverage. The business indicators include at least the number of roaming days, the number of terminal switching times, and / or the number channel. The strategy level is determined based on the aforementioned case-related indicators, the aforementioned basic indicators, and the aforementioned business indicators; The processing strategy for the fraudulent number is determined based on the probability of reactivation and the strategy level.

2. The number identification method as described in claim 1, characterized in that, The step of determining the strategy level based on the case-related indicators, the basic indicators, and the business indicators includes: The first strategy level of the identification strategy determined based on the aforementioned indicators involved in the case; If the first strategy level is a preset level, then the first strategy level is used as the strategy level of the identification strategy; If the first strategy level is not a preset level, then a second strategy level for the identification strategy is determined based on the basic indicators and the business indicators, and the second strategy level is used as the strategy level of the identification strategy.

3. The number identification method as described in claim 1, characterized in that, The step of determining the processing strategy for the suspected fraudulent number based on the call reactivation probability and the strategy level includes: Determine the probability threshold based on the strategy level; Determine whether the probability of re-run is less than the probability threshold; When the probability of a fraudulent number being reconnected is less than the probability threshold, the fraudulent number will be shut down, its calls will be restricted, or it will require secondary real-name authentication.

4. The number identification method as described in claim 1, characterized in that, Before the step of inputting the user features into a preset probability prediction model, the method further includes: Obtain training samples, which include numbers that were shut down and then reactivated, as well as numbers that were shut down and then not reactivated. Extract the target user features from the training samples; A pre-defined neural network model is trained based on the target user characteristics; The converged neural network model is used as a probabilistic prediction model.

5. The number identification method as described in claim 4, characterized in that, The step of extracting the target user features of the training samples includes: Extract the user features to be determined from the training samples, wherein the number of user features to be determined is greater than the number of target user features; The feature importance and information value (IV) of the user features to be determined are determined according to a preset algorithm, wherein the feature importance is determined by the random forest algorithm; Target user features are determined from the user features to be determined based on the importance of the features and the value of the information.

6. The number identification method as described in claim 1, characterized in that, The steps for obtaining the suspected fraudulent phone number include: Obtain multiple user numbers to be identified and extract the features of the user numbers; The characteristics of the user's number are input into a preset number recognition model to obtain the number suspected of being involved in fraud.

7. A number recognition device, characterized in that, The number recognition device includes: The acquisition module is used to acquire fraudulent phone numbers and extract user characteristics of the fraudulent phone numbers; The prediction module is used to input the user characteristics into a preset probability prediction model to obtain the probability of the fraudulent number being called back. The determination module is used to acquire environmental indicators, which include case-related indicators, basic indicators, and business indicators. The case-related indicators include at least the number of fraudulent phone numbers within a preset time period. The basic indicators include at least the data coverage. The business indicators include at least the number of roaming days, the number of terminal switching times, and / or the number channel. The strategy level is determined based on the case-related indicators, the basic indicators, and the business indicators. The processing module is used to determine the processing strategy for the fraudulent number based on the reactivation probability and the strategy level.

8. A number recognition device, characterized in that, The number recognition device includes a memory, a processor, and a number recognition program stored in the memory and executable on the processor, wherein the number recognition program, when executed by the processor, implements the various steps of the number recognition method as described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a number recognition program, which, when executed by a processor, implements the steps of the number recognition method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Method for dynamically fusing rules to fraud behavior recognition model

    CN111709472A

  • Number identification method and device, equipment and medium

    CN112839335A

  • Identification method and device of information identification model based on transfer learning, and equipment

    CN113055208A