Number identification methods, devices, equipment, and computer-readable storage media

CN115878990BActive Publication Date: 2026-08-14CHINA MOBILE GROUP ZHEJIANG +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-26
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0005]本发明的主要目的在于提供一种号码识别方法、装置、设备以及计算机可读存储介质,旨在解决涉诈号码的识别准确度较低的问题

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115878990B_ABST
    Figure CN115878990B_ABST
Patent Text Reader

Abstract

This invention discloses a number identification method, apparatus, device, and computer-readable storage medium. The method includes: acquiring a verification set and inputting the verification set into a corresponding target identification model; acquiring fraudulent numbers output by each target identification model; determining multiple P-values ​​corresponding to each target identification model, and determining the number of fraudulent numbers corresponding to each P-value and the recurrence probability of the fraudulent numbers, wherein the P-value represents the degree of influence of the target identification model on a preset initial identification model; determining a target P-value range for the target identification model based on the number of numbers corresponding to each P-value and the recurrence probability; and determining an identification strategy for the fraudulent numbers based on the target P-value ranges corresponding to each target identification model. This invention makes the identification of fraudulent numbers more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Technology Neighborhood

[0002] This invention relates to the field of computer technology, and in particular to a number identification method, apparatus, device, and computer-readable storage medium. Background Technology

[0003] In recent years, with the development and popularization of mobile communication technology, society has gradually entered the information age, and telecommunications fraud has become increasingly rampant, causing significant financial losses to the public. As operators gradually improve their fraud control systems, while increasing efforts to combat telecommunications fraud, the number of misjudgments will inevitably increase, leading to incorrect actions against legitimate users and impacting their daily lives.

[0004] Currently, operators mainly develop fraud models by analyzing the numbers involved in cases and relying on the experience of their staff. This usually requires manual extraction of a large number of features to identify and handle user numbers, resulting in low accuracy in identifying fraudulent numbers. Summary of the Invention

[0005] The main objective of this invention is to provide a number identification method, apparatus, device, and computer-readable storage medium, aiming to solve the problem of low accuracy in identifying fraudulent numbers.

[0006] To achieve the above objectives, the present invention provides a number identification method, which includes the following steps:

[0007] Obtain a verification set and input the verification set into the corresponding target recognition model to obtain the fraudulent numbers output by each target recognition model;

[0008] Multiple P-values ​​are determined for each target identification model, and the number of fraudulent numbers and the repeat probability of each fraudulent number are determined for each P-value. The P-value represents the degree of influence of the target identification model on the preset initial identification model.

[0009] The target P-value range of the target recognition model is determined based on the number of numbers corresponding to each P-value of each target recognition model and the recurrence rate.

[0010] The identification strategy for the fraudulent number is determined based on the target P-value range corresponding to each target identification model.

[0011] In one embodiment, the step of determining the target P-value range of the target recognition model based on the number of numbers corresponding to each P-value of each target recognition model and the probability of repetition includes:

[0012] Determine the probability distribution and immediate reward value for each of the P values ​​corresponding to each target recognition model;

[0013] The total reward value is determined based on the probability distribution of each of the P values ​​and the immediate reward value;

[0014] The target P-value range is determined based on the total return value.

[0015] In one embodiment, before the step of inputting the verification set into the corresponding target recognition model, the method further includes:

[0016] Obtain a training set, which includes a first sub-training set, a second sub-training set, and / or a third sub-training set. The first sub-training set includes numbers involved in fraud within a preset time period and numbers removed from the whitelist of all users on the network. The second sub-training set includes numbers that have been shut down and not reactivated as identified by a preset initial identification model and numbers removed from the whitelist of all users on the network. The third sub-training set includes numbers that have been shut down and not reactivated as identified by a preset initial identification model and numbers that have been shut down and reactivated.

[0017] A first target recognition model is obtained by training a preset neural network model according to a preset first algorithm and the first training set;

[0018] And / or, a second target recognition model is obtained by training a preset neural network model according to a preset second algorithm and the second training set;

[0019] And / or, a third target recognition model is obtained by training a preset neural network model according to a preset third algorithm and the third training set.

[0020] In one embodiment, after the step of obtaining the training set, the method further includes:

[0021] Obtain the user features corresponding to the training set;

[0022] Determine the importance and information value of each of the user features in the random forest;

[0023] Target features are determined based on the importance of the random forest and the value of the information. The target features include at least user basic information, communication behavior, roaming behavior and / or consumption behavior.

[0024] A preset neural network model is trained based on the target features.

[0025] In one embodiment, the step of obtaining a validation set and inputting the validation set into each target recognition model includes:

[0026] Obtain the verification set, which includes a first sub-verification set, a second sub-verification set, and a third sub-verification set. The first sub-verification set includes numbers suspected of fraud within a preset time period and numbers removed from the whitelist of all users on the network. The second sub-verification set includes numbers that have been shut down and not reactivated as identified by a preset initial identification model and numbers removed from the whitelist of all users on the network. The third sub-verification set includes numbers that have been shut down and not reactivated as identified by a preset initial identification model and numbers that have been shut down and reactivated.

[0027] Input the first verification set into the first target recognition model;

[0028] Input the second validation set into the second target recognition model;

[0029] The third verification set is input into the third target recognition model.

[0030] In one embodiment, the step of determining the identification strategy for the fraudulent number based on the target P-value range corresponding to each target identification model includes:

[0031] The identification strategy for the fraudulent number is determined based on a preset decision tree algorithm and a target P-value range. The identification strategy includes each target initial model and the corresponding P-value of the target initial model.

[0032] In one embodiment, the step of determining the number of fraudulent numbers corresponding to each P value and the repeat probability of the fraudulent numbers includes:

[0033] The numbers involved in the fraud should be deactivated.

[0034] Obtain the number of the fraudulent phone numbers that have been reactivated within a preset time period;

[0035] The call recovery rate is determined based on the number of fraudulent phone numbers and the number of fraudulent phone numbers that have been reactivated.

[0036] To achieve the above objectives, the present invention also provides a number recognition device, the number recognition device comprising:

[0037] The acquisition module is used to acquire a verification set and input the verification set into the corresponding target recognition model to acquire the fraudulent numbers output by each target recognition model.

[0038] The determination module is used to determine multiple P values ​​corresponding to each target recognition model, and to determine the number of fraudulent numbers corresponding to each P value and the recurrence rate of the fraudulent numbers. The P value represents the degree of influence of the target recognition model on the preset initial recognition model.

[0039] The calculation module is used to determine the target P-value range of the target recognition model based on the number of numbers corresponding to each P-value of each target recognition model and the probability of repetition.

[0040] The identification module is used to determine the identification strategy for the fraudulent number based on the target P-value range corresponding to each target identification model.

[0041] To achieve the above objectives, the present invention also provides a number recognition device, the number recognition device including a memory, a processor, and a number recognition program stored in the memory and executable on the processor, wherein the number recognition program, when executed by the processor, implements the various steps of the number recognition method as described above.

[0042] To achieve the above objectives, the present invention also provides a computer-readable storage medium storing a number recognition program, which, when executed by a processor, implements the various steps of the number recognition method described above.

[0043] This invention provides a number identification method, apparatus, device, and computer-readable storage medium. The method involves acquiring a verification set and inputting it into a corresponding target identification model to obtain the fraudulent numbers output by each model. Multiple P-values ​​are determined for each target identification model, along with the number of fraudulent numbers corresponding to each P-value and the recurrence rate of these numbers. A target P-value range for each target identification model is determined based on the number of numbers corresponding to each P-value and the recurrence rate. Finally, an identification strategy for fraudulent numbers is determined based on the target P-value range for each target identification model. By determining the target P-value range for each target identification model based on the number of fraudulent numbers and the recurrence rate, the number of numbers output by the target identification model within the target P-value range is stable, and the recurrence rate is low. The identification strategy determined based on the target P-value range makes the identification of fraudulent numbers more accurate. Attached Figure Description

[0044] Figure 1 This is a schematic diagram of the hardware structure of the number recognition device according to an embodiment of the present invention;

[0045] Figure 2 This is a flowchart illustrating the first embodiment of the number recognition method of the present invention;

[0046] Figure 3 This is a detailed flowchart of step S30 of the second embodiment of the number recognition method of the present invention;

[0047] Figure 4 This is a flowchart illustrating a third embodiment of the number recognition method of the present invention;

[0048] Figure 5This is a schematic diagram of the logical structure of the number recognition device of the present invention.

[0049] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0050] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0051] The main solution of this invention is as follows: obtain a verification set and input the verification set into the corresponding target recognition model to obtain the fraudulent numbers output by each target recognition model; determine multiple P values ​​corresponding to each target recognition model, and determine the number of fraudulent numbers corresponding to each P value and the recurrence rate of the fraudulent numbers; determine the target P value range of the target recognition model based on the number of numbers corresponding to each P value and the recurrence rate; and determine the identification strategy for fraudulent numbers based on the target P value range of each target recognition model.

[0052] The target P-value range corresponding to the target identification model is determined by the number of fraudulent numbers and the recurrence rate. Within the target P-value range, the number of numbers output by the target identification model is stable and the recurrence rate is low. The identification strategy determined based on the target P-value range makes the identification of fraudulent numbers more accurate.

[0053] As one implementation solution, number recognition devices can be like Figure 1 As shown.

[0054] The embodiments of the present invention relate to a number recognition device, which includes: a processor 101, such as a CPU, a memory 102, and a communication bus 103. The communication bus 103 is used to enable communication between these components.

[0055] Memory 102 can be high-speed RAM or stable memory (non-volatile memory), such as disk storage. Figure 1 As shown, the memory 102, which is a computer-readable storage medium, may include a number recognition program; and the processor 101 may be used to call the number recognition program stored in the memory 102 and perform the following operations:

[0056] Obtain a verification set and input the verification set into the corresponding target recognition model to obtain the fraudulent numbers output by each target recognition model;

[0057] Multiple P-values ​​are determined for each target identification model, and the number of fraudulent numbers and the repeat probability of each fraudulent number are determined for each P-value. The P-value represents the degree of influence of the target identification model on the preset initial identification model.

[0058] The target P-value range of the target recognition model is determined based on the number of numbers corresponding to each P-value of each target recognition model and the recurrence rate.

[0059] The identification strategy for the fraudulent number is determined based on the target P-value range corresponding to each target identification model.

[0060] In one embodiment, the processor 101 can be used to invoke a number recognition program stored in the memory 102 and perform the following operations:

[0061] Determine the probability distribution and immediate reward value for each of the P values ​​corresponding to each target recognition model;

[0062] The total reward value is determined based on the probability distribution of each of the P values ​​and the immediate reward value;

[0063] The target P-value range is determined based on the total return value.

[0064] In one embodiment, the processor 101 can be used to invoke a number recognition program stored in the memory 102 and perform the following operations:

[0065] Obtain a training set, which includes a first sub-training set, a second sub-training set, and / or a third sub-training set. The first sub-training set includes numbers involved in fraud within a preset time period and numbers removed from the whitelist of all users on the network. The second sub-training set includes numbers that have been shut down and not reactivated as identified by a preset initial identification model and numbers removed from the whitelist of all users on the network. The third sub-training set includes numbers that have been shut down and not reactivated as identified by a preset initial identification model and numbers that have been shut down and reactivated.

[0066] A first target recognition model is obtained by training a preset neural network model according to a preset first algorithm and the first training set;

[0067] And / or, a second target recognition model is obtained by training a preset neural network model according to a preset second algorithm and the second training set;

[0068] And / or, a third target recognition model is obtained by training a preset neural network model according to a preset third algorithm and the third training set.

[0069] In one embodiment, the processor 101 can be used to invoke a number recognition program stored in the memory 102 and perform the following operations:

[0070] Obtain the user features corresponding to the training set;

[0071] Determine the importance and information value of each of the user features in the random forest;

[0072] Target features are determined based on the importance of the random forest and the value of the information. The target features include at least user basic information, communication behavior, roaming behavior and / or consumption behavior.

[0073] A preset neural network model is trained based on the target features.

[0074] In one embodiment, the processor 101 can be used to invoke a number recognition program stored in the memory 102 and perform the following operations:

[0075] Obtain the verification set, which includes a first sub-verification set, a second sub-verification set, and a third sub-verification set. The first sub-verification set includes numbers suspected of fraud within a preset time period and numbers removed from the whitelist of all users on the network. The second sub-verification set includes numbers that have been shut down and not reactivated as identified by a preset initial identification model and numbers removed from the whitelist of all users on the network. The third sub-verification set includes numbers that have been shut down and not reactivated as identified by a preset initial identification model and numbers that have been shut down and reactivated.

[0076] Input the first verification set into the first target recognition model;

[0077] Input the second validation set into the second target recognition model;

[0078] The third verification set is input into the third target recognition model.

[0079] In one embodiment, the processor 101 can be used to invoke a number recognition program stored in the memory 102 and perform the following operations:

[0080] The identification strategy for the fraudulent number is determined based on a preset decision tree algorithm and a target P-value range. The identification strategy includes each target initial model and the corresponding P-value of the target initial model.

[0081] In one embodiment, the processor 101 can be used to invoke a number recognition program stored in the memory 102 and perform the following operations:

[0082] The numbers involved in the fraud should be deactivated.

[0083] Obtain the number of the fraudulent phone numbers that have been reactivated within a preset time period;

[0084] The call recovery rate is determined based on the number of fraudulent phone numbers and the number of fraudulent phone numbers that have been reactivated.

[0085] Based on the hardware architecture of the aforementioned number recognition device, embodiments of the number recognition method of the present invention are proposed.

[0086] Reference Figure 2 , Figure 2 This is a first embodiment of the number recognition method of the present invention, the number recognition method includes the following steps:

[0087] Step S10: Obtain a verification set and input the verification set into the corresponding target recognition model to obtain the fraudulent numbers output by each target recognition model;

[0088] Specifically, the target identification model is used to identify fraudulent phone numbers. The validation set is used to verify the accuracy of the target identification model. The validation set includes a first sub-validation set, a second sub-validation set, and a third sub-validation set. The first sub-validation set includes fraudulent phone numbers within a preset time period, as well as phone numbers removed from the whitelist by all users across the network. The second sub-validation set includes phone numbers identified by the preset initial identification model that have been shut down and not reactivated, as well as phone numbers that have been shut down and reactivated. The initial identification model is used to identify fraudulent phone numbers, and its accuracy in identifying fraudulent phone numbers is lower than that of the target identification model.

[0089] Different validation sets correspond to different target recognition models. Different validation sets are input into their respective target recognition models. For example, the first validation set is input into the first target recognition model corresponding to the first validation set; the second validation set is input into the second target recognition model; and the third validation set is input into the third target recognition model. Each target recognition model outputs the suspected fraudulent numbers based on its corresponding validation set. For example, the first target recognition model predicts and identifies suspected fraudulent numbers among all non-whitelisted users across the entire network using the first validation set. The second target recognition model predicts and identifies suspected fraudulent numbers among all non-whitelisted users across the entire network using the second validation set. The third target recognition model predicts and identifies suspected fraudulent numbers for services that have been shut down and then reactivated using the third validation set.

[0090] Step S20: Determine multiple P values ​​corresponding to each target identification model, and determine the number of fraudulent numbers corresponding to each P value and the repeat probability of the fraudulent numbers. The P value represents the degree of influence of the target identification model on the preset initial identification model.

[0091] Specifically, each target recognition model corresponds to multiple P values, where the P value represents the degree of influence of the target recognition model on the initial recognition model, that is, the degree of optimization of the target recognition model relative to the initial recognition model.

[0092] The system suspends the service of fraudulent phone numbers output by the target identification model and obtains the number of fraudulent phone numbers that resume service within a preset time period. The resumption rate is determined based on the number of fraudulent phone numbers and the number of resumed fraudulent phone numbers. In other words, the resumption rate of fraudulent phone numbers is determined by the number of fraudulent phone numbers that resume service after suspension and the number of fraudulent phone numbers output by the target identification model, as shown in the following formula:

[0093]

[0094] The number of fraudulent phone numbers and the repeat probability corresponding to the target identification model under different P values ​​are determined, as exemplified in the table below:

[0095]

[0096]

[0097] As shown in the table above, the larger the P-value of the target recognition model, the smaller the number of fraudulent phone numbers output, and the lower the recurrence rate. When the P-value is 0.9, although the recurrence rate is low, the number of fraudulent phone numbers output is small, and there may be fraudulent phone numbers that are not identified. When the P-value is 0.1, although the number of fraudulent phone numbers output is large, the recurrence rate is high, and there are many cases where legitimate phone numbers are identified as fraudulent phone numbers.

[0098] Step S30: Determine the target P-value range of the target recognition model based on the number of numbers corresponding to each P-value of each target recognition model and the probability of repetition.

[0099] Specifically, the target P-value range of the target identification model is determined based on the number of numbers corresponding to each P-value of each target identification model and the recurrence rate. Within the target P-value range, the number of fraudulent numbers output by the target identification model is stable and the recurrence rate is low. For example, as shown in the table above, when the P-value is between 0.5 and 0.8, the number of fraudulent numbers output by the target identification model is stable and the recurrence rate is low. The range [0.5, 0.8] is taken as the target P-value range.

[0100] Step S40: Determine the identification strategy for the fraudulent number based on the target P-value range corresponding to each target identification model.

[0101] Specifically, the identification strategy for fraudulent phone numbers is determined based on the target P-value range corresponding to each target identification model. For example, the identification strategy can be determined based on a preset decision tree algorithm and the target P-value range. The identification strategy includes each initial target model and its corresponding P-value. For instance, when the target P-value range for the first target identification model is [0.4, 0.7], the target P-value range for the second target identification model is [0.6, 0.7], and the target P-value range for the third target identification model is [0.3, 0.5], the identification strategy for fraudulent phone numbers could be: setting the first target identification model and its corresponding P-value to 0.5; setting the second target identification model and its corresponding P-value to 0.6; and setting the third target identification model and its corresponding P-value to 0.4.

[0102] In this embodiment, a verification set is obtained and input into the corresponding target recognition model. The fraudulent phone numbers output by each target recognition model are then obtained. Multiple P-values ​​are determined for each target recognition model, along with the number of fraudulent phone numbers corresponding to each P-value and the recurrence rate of the fraudulent phone numbers. A target P-value range for the target recognition model is determined based on the number of phone numbers corresponding to each P-value and the recurrence rate. Finally, a fraudulent phone number identification strategy is determined based on the target P-value range for each target recognition model. By determining the target P-value range for the target recognition model based on the number of fraudulent phone numbers and the recurrence rate, the number of phone numbers output by the target recognition model within the target P-value range is stable, and the recurrence rate is low. The identification strategy determined based on the target P-value range makes the identification of fraudulent phone numbers more accurate.

[0103] Reference Figure 3 , Figure 3 This is a second embodiment of the number identification method of the present invention. Based on the first embodiment, step S30 includes:

[0104] Step S31: Determine the probability distribution of each P value and the immediate reward value corresponding to each target recognition model;

[0105] Step S32: Determine the total reward value based on the probability distribution of each of the P values ​​and the immediate reward value;

[0106] Step S33: Determine the target P value range based on the total return value.

[0107] Specifically, the target P-value range for each target recognition model is determined based on the number of numbers corresponding to each P-value and the probability of recurrence. The probability distribution and immediate reward value for each P-value corresponding to each target recognition model are determined; the total reward value is then determined based on the probability distribution and immediate reward value of each P-value, as shown in the following formula:

[0108] Q * (s,a)=max π Q * (s,a);

[0109] Q * (s,a)=∑ s′ P(s′|s,a)(R(s,a,s′)+γmax a′ Q * (s′,a′));

[0110] Among them, Q * Q represents the total return. * Let be the optimal value action function; s represents the current state, a represents the action from the current state to the next state, s′ represents the next state, a′ is the next action, γ represents the learning parameter, and R represents the immediate reward value.

[0111] The target P-value range is determined based on the total return value. The total return values ​​corresponding to each P-value are sorted. If the ranking of the total return value corresponding to a P-value is lower than the preset ranking, the target P-value range is determined based on the P-values ​​that are lower than the preset ranking.

[0112] In the technical solution of this embodiment, the probability distribution and immediate reward value of each P-value corresponding to each target recognition model are determined; the total reward value is determined based on the probability distribution and immediate reward value of each P-value; and the target P-value range is determined based on the total reward value. Determining the total reward value based on the probability distribution and immediate reward value, and determining the target P-value range based on the total reward value, ensures that the number of numbers output by the target recognition model is stable and the recurrence rate is low, making the identification of fraudulent numbers more accurate.

[0113] Reference Figure 4 , Figure 4 This is a third embodiment of the number identification method of the present invention. Based on the first or second embodiment, before step S10, it further includes:

[0114] Step S50: Obtain a training set, which includes a first sub-training set, a second sub-training set, and / or a third sub-training set. The first sub-training set includes numbers involved in fraud within a preset time period and numbers removed from the whitelist of all users on the network. The second sub-training set includes numbers that have been shut down and not reactivated as identified by a preset initial identification model and numbers removed from the whitelist of all users on the network. The third sub-training set includes numbers that have been shut down and not reactivated as identified by a preset initial identification model and numbers that have been shut down and reactivated.

[0115] Step S60: Train a preset neural network model according to a preset first algorithm and the first training set to obtain a first target recognition model; and / or, train a preset neural network model according to a preset second algorithm and the second training set to obtain a second target recognition model; and / or, train a preset neural network model according to a preset third algorithm and the third training set to obtain a third target recognition model.

[0116] Specifically, a training set is obtained, which includes a first sub-training set, a second sub-training set, and / or a third sub-training set. The first sub-training set includes numbers suspected of fraud within a preset time period and numbers removed from the whitelist of all users on the network. The second sub-training set includes numbers identified by the preset initial identification model that have been shut down and not reactivated, as well as numbers removed from the whitelist of all users on the network. The third sub-training set includes numbers identified by the preset initial identification model that have been shut down and not reactivated, as well as numbers that have been shut down and reactivated.

[0117] Obtain user features corresponding to the training set; determine the random forest importance and information value of each user feature; determine target features based on the random forest importance and information value, wherein the random forest importance is obtained by the random forest algorithm, and the target features include at least user basic information, communication behavior, roaming behavior and / or consumption behavior.

[0118] A pre-defined neural network model is trained based on the target features corresponding to the training set to obtain a corresponding target recognition model. A first target recognition model is obtained by training the pre-defined neural network model using a first algorithm and a first training set, where the first algorithm can be a logistic regression algorithm in classification algorithms. Alternatively, a second target recognition model is obtained by training the pre-defined neural network model using a second algorithm and a second training set, where the second algorithm can be a boosting iterative algorithm. And / or, a third target recognition model is obtained by training the pre-defined neural network model using a third algorithm and a third training set, where the third algorithm can be an XGBoost gradient boosting tree algorithm.

[0119] In this embodiment, a training set is obtained, which includes a first sub-training set, a second sub-training set, and / or a third sub-training set. A first target recognition model is obtained by training a preset neural network model according to a preset first algorithm and the first training set. Alternatively, a second target recognition model is obtained by training a preset neural network model according to a preset second algorithm and the second training set. And / or, a third target recognition model is obtained by training a preset neural network model according to a preset third algorithm and the third training set. Multiple target recognition models are obtained by training different sub-training sets and algorithms, providing multiple methods for number recognition. This facilitates the subsequent selection of the target recognition model and its corresponding P-value as the identification strategy for fraudulent numbers, ensuring a stable number of numbers output by each target recognition model and a low probability of duplication, thus making the identification of fraudulent numbers more accurate.

[0120] Reference Figure 5 The present invention also provides a number recognition device, the number recognition device comprising:

[0121] The acquisition module 100 is used to acquire a verification set and input the verification set into the corresponding target recognition model to acquire the fraudulent numbers output by each target recognition model.

[0122] The determination module 200 is used to determine multiple P values ​​corresponding to each target identification model, and to determine the number of fraudulent numbers corresponding to each P value and the repeat probability of the fraudulent numbers. The P value represents the degree of influence of the target identification model on the preset initial identification model.

[0123] The calculation module 300 is used to determine the target P-value range of the target recognition model based on the number of numbers corresponding to each P-value of each target recognition model and the recurrence probability.

[0124] The identification module 400 is used to determine the identification strategy of the fraudulent number based on the target P-value range corresponding to each target identification model.

[0125] In one embodiment, in determining the target P-value range of the target recognition model based on the number of numbers corresponding to each P-value of each target recognition model and the probability of repetition, the calculation module 300 is specifically used for:

[0126] Determine the probability distribution and immediate reward value for each of the P values ​​corresponding to each target recognition model;

[0127] The total reward value is determined based on the probability distribution of each of the P values ​​and the immediate reward value;

[0128] The target P-value range is determined based on the total return value.

[0129] In one embodiment, before inputting the verification set into the corresponding target recognition model, the acquisition module 100 is specifically used for:

[0130] Obtain a training set, which includes a first sub-training set, a second sub-training set, and / or a third sub-training set. The first sub-training set includes numbers involved in fraud within a preset time period and numbers removed from the whitelist of all users on the network. The second sub-training set includes numbers that have been shut down and not reactivated as identified by a preset initial identification model and numbers removed from the whitelist of all users on the network. The third sub-training set includes numbers that have been shut down and not reactivated as identified by a preset initial identification model and numbers that have been shut down and reactivated.

[0131] A first target recognition model is obtained by training a preset neural network model according to a preset first algorithm and the first training set;

[0132] And / or, a second target recognition model is obtained by training a preset neural network model according to a preset second algorithm and the second training set;

[0133] And / or, a third target recognition model is obtained by training a preset neural network model according to a preset third algorithm and the third training set.

[0134] In one embodiment, after acquiring the training set, the acquisition module 100 is specifically used for:

[0135] Obtain the user features corresponding to the training set;

[0136] Determine the importance and information value of each of the user features in the random forest;

[0137] Target features are determined based on the importance of the random forest and the value of the information. The target features include at least user basic information, communication behavior, roaming behavior and / or consumption behavior.

[0138] A preset neural network model is trained based on the target features.

[0139] In one embodiment, the acquisition module 100 is specifically used for: acquiring a validation set and inputting the validation set into each target recognition model.

[0140] Obtain the verification set, which includes a first sub-verification set, a second sub-verification set, and a third sub-verification set. The first sub-verification set includes numbers suspected of fraud within a preset time period and numbers removed from the whitelist of all users on the network. The second sub-verification set includes numbers that have been shut down and not reactivated as identified by a preset initial identification model and numbers removed from the whitelist of all users on the network. The third sub-verification set includes numbers that have been shut down and not reactivated as identified by a preset initial identification model and numbers that have been shut down and reactivated.

[0141] Input the first verification set into the first target recognition model;

[0142] Input the second validation set into the second target recognition model;

[0143] The third verification set is input into the third target recognition model.

[0144] In one embodiment, in determining the identification strategy for the fraudulent number based on the target P-value range corresponding to each target identification model, the identification module 400 is specifically used for:

[0145] The identification strategy for the fraudulent number is determined based on a preset decision tree algorithm and a target P-value range. The identification strategy includes each target initial model and the corresponding P-value of the target initial model.

[0146] In one embodiment, the determining module 200 is specifically used to determine the number of fraudulent numbers corresponding to each P value and the repeat probability of the fraudulent numbers, in order to:

[0147] The numbers involved in the fraud should be deactivated.

[0148] Obtain the number of the fraudulent phone numbers that have been reactivated within a preset time period;

[0149] The call recovery rate is determined based on the number of fraudulent phone numbers and the number of fraudulent phone numbers that have been reactivated.

[0150] The present invention also provides a number recognition device, the number recognition device including a memory, a processor, and a number recognition program stored in the memory and executable on the processor, wherein when the number recognition program is executed by the processor, it implements the various steps of the number recognition method as described in the above embodiments.

[0151] The present invention also provides a computer-readable storage medium storing a number recognition program, which, when executed by a processor, implements the various steps of the number recognition method described in the above embodiments.

[0152] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0153] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, system, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, system, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, system, article, or apparatus that includes that element.

[0154] Through the above description of the embodiments, those skilled in the art can clearly understand that the systems described in the above embodiments can be implemented using software plus necessary general-purpose hardware platforms. Of course, they can also be implemented using hardware, but in many cases, the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a computer-readable storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, parking management device, air conditioner, or network device, etc.) to execute the systems described in the various embodiments of the present invention.

[0155] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.

Claims

1. A number recognition method, characterized in that, The number identification method includes: Obtain a verification set and input the verification set into the corresponding target recognition model to obtain the fraudulent numbers output by each target recognition model; Multiple P-values ​​are determined for each target recognition model, and the number of fraudulent numbers and the recurrence rate of the fraudulent numbers are determined for each P-value. The larger the P-value of the target recognition model, the smaller the number of fraudulent numbers output and the lower the recurrence rate. The P-value represents the degree of influence of the target recognition model on the preset initial recognition model. The target P-value range of the target recognition model is determined based on the number of numbers corresponding to each P-value of each target recognition model and the recurrence rate. The identification strategy for the fraudulent number is determined based on the target P-value range corresponding to each target identification model.

2. The number identification method as described in claim 1, characterized in that, The step of determining the target P-value range of the target recognition model based on the number of numbers corresponding to each P-value of each target recognition model and the recurrence probability includes: Determine the probability distribution and immediate reward value for each of the P values ​​corresponding to each target recognition model; The total reward value is determined based on the probability distribution of each of the P values ​​and the immediate reward value; The target P-value range is determined based on the total return value.

3. The number identification method as described in claim 1, characterized in that, Before the step of inputting the validation set into the corresponding target recognition model, the method further includes: Obtain a training set, which includes a first sub-training set, a second sub-training set, and / or a third sub-training set. The first sub-training set includes numbers involved in fraud within a preset time period and numbers removed from the whitelist of all users on the network. The second sub-training set includes numbers that have been shut down and not reactivated as identified by a preset initial identification model and numbers removed from the whitelist of all users on the network. The third sub-training set includes numbers that have been shut down and not reactivated as identified by a preset initial identification model and numbers that have been shut down and reactivated. A first target recognition model is obtained by training a preset neural network model according to a preset first algorithm and the first sub-training set; And / or, a second target recognition model is obtained by training a preset neural network model according to a preset second algorithm and the second sub-training set; And / or, a third target recognition model is obtained by training a preset neural network model according to a preset third algorithm and the third sub-training set.

4. The number identification method as described in claim 3, characterized in that, Following the step of obtaining the training set, the method further includes: Obtain the user features corresponding to the training set; Determine the importance and information value of each of the user features in the random forest; Target features are determined based on the importance of the random forest and the value of the information. The target features include at least user basic information, communication behavior, roaming behavior and / or consumption behavior. A preset neural network model is trained based on the target features.

5. The number identification method as described in claim 3, characterized in that, The step of obtaining the validation set and inputting the validation set into each target recognition model includes: Obtain the verification set, which includes a first sub-verification set, a second sub-verification set, and a third sub-verification set. The first sub-verification set includes numbers suspected of fraud within a preset time period and numbers removed from the whitelist of all users on the network. The second sub-verification set includes numbers that have been shut down and not reactivated as identified by a preset initial identification model and numbers removed from the whitelist of all users on the network. The third sub-verification set includes numbers that have been shut down and not reactivated as identified by a preset initial identification model and numbers that have been shut down and reactivated. Input the first sub-validation set into the first target recognition model; Input the second sub-validation set into the second target recognition model; The third sub-validation set is input into the third target recognition model.

6. The number identification method as described in claim 1, characterized in that, The step of determining the identification strategy for the fraudulent number based on the target P-value range corresponding to each target identification model includes: The identification strategy for the fraudulent number is determined based on a preset decision tree algorithm and a target P-value range. The identification strategy includes each target identification model and the corresponding P-value of the target identification model.

7. The number identification method as described in claim 1, characterized in that, The repeat call rate of the fraudulent phone number is obtained through the following steps: The numbers involved in the fraud should be deactivated. Obtain the number of the fraudulent phone numbers that have been reactivated within a preset time period; The call recovery rate is determined based on the number of fraudulent phone numbers and the number of fraudulent phone numbers that have been reactivated.

8. A number recognition device, characterized in that, The number recognition device includes: The acquisition module is used to acquire a verification set and input the verification set into the corresponding target recognition model to acquire the fraudulent numbers output by each target recognition model. The determination module is used to determine multiple P values ​​corresponding to each target recognition model, and to determine the number of fraudulent numbers corresponding to each P value and the recurrence rate of the fraudulent numbers. The larger the P value of the target recognition model, the smaller the number of fraudulent numbers output and the lower the recurrence rate. The P value represents the degree of influence of the target recognition model on the preset initial recognition model. The calculation module is used to determine the target P-value range of the target recognition model based on the number of numbers corresponding to each P-value of each target recognition model and the probability of repetition. The identification module is used to determine the identification strategy for the fraudulent number based on the target P-value range corresponding to each target identification model.

9. A number recognition device, characterized in that, The number recognition device includes a memory, a processor, and a number recognition program stored in the memory and executable on the processor, wherein the number recognition program, when executed by the processor, implements the various steps of the number recognition method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a number recognition program, which, when executed by a processor, implements the steps of the number recognition method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Fraud number identification method and device, computer equipment and storage medium

    CN112291424A

  • Number identification method and device, equipment and medium

    CN112839335A