A method and device for identifying fraud numbers

By combining call behavior characteristics, user business characteristics and trajectory characteristics data, and using evidence theory to fusion the recognition results, the problem of inaccurate identification of fraud numbers in the prior art is solved, and higher recognition accuracy and stability are achieved.

CN115334510BActive Publication Date: 2025-06-10CHINA TELECOM CORP LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210899524.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-28
Publication Date
2025-06-10
Estimated Expiration
2042-07-28

AI Technical Summary

Technical Problem

It is difficult for the existing technology to achieve accurate analysis and identification of fraud numbers, especially in the GOIP fraud model, the cost and difficulty of operators' discovery, analysis, interception and traceability have increased significantly.

Method used

By entering the call behavior characteristic data of the target number, the user's business characteristic data and the trajectory characteristic data into different identification models, and combining the identification results based on the evidence theory, we determine whether the target number is a fraud number.

Benefits of technology

It improves the accuracy and stability of fraud number identification, enhances the fault tolerance of the model, can more effectively identify fraud number, and reduces operator supervision and crackdown costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115334510B_ABST
    Figure CN115334510B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a method and device for identifying fraud numbers. The method for identifying fraud numbers may include: inputting call behavior characteristic data and user service characteristic data of a target number into a first identification model to obtain a first probability value that the target number is a fraud number; inputting trajectory characteristic data of the target number within a preset duration into a second identification model to obtain a second probability value that the target number is a fraud number; and determining whether the target number is a fraud number based on the evidence theory and the first probability value and the second probability value. The technical solution provided by the embodiment of the present application can solve the problem that it is difficult to accurately analyze and identify fraud numbers by traditional fraud number identification methods in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence technology, and particularly relates to a method and device for identifying fraud numbers. Background Art

[0002] With the continuous development of information and communication technology and artificial intelligence technology, telecommunications services have provided great convenience for social development and people's lives. However, while bringing convenience, telecommunications services also provide some "opportunities" for some lawbreakers. Nowadays, the number of cases of fraud crimes committed through telecommunications technology has shown an obvious increasing trend, causing serious harm and impact to the life and property safety of telecommunications users.

[0003] Traditional telecommunications fraud methods have the characteristics that the fraudsters are not separated from the SIM (Subscriber Identity Module) card, the outbound call frequency is high, and a large number of terminal locations are concentrated. Therefore, it is very easy to be identified by the anti-fraud models of operators. In order to evade the supervision and crackdown of operators, using GOIP (GSM Over Internet Protocol) devices that seamlessly connect the GSM (Global System for Mobile Communications) network and the VoIP (Voice Over Internet Protocol) network for fraud has become a new trend and means of telecommunications fraud.

[0004] Compared with the traditional telecommunications fraud mode, the GOIP fraud mode has the characteristics of separation of people and cards, virtual dialing, arbitrary switching of numbers for dialing, and call-back can be connected, which greatly increases the cost and difficulty for operators to discover, judge, intercept, and trace such fraud modes. Moreover, the traditional methods for identifying fraud numbers mostly focus on analyzing single call behavior data, with a single type of information source, insufficient comprehensive feature analysis, poor adaptability, and it is difficult to achieve accurate analysis and identification of fraud numbers. Summary of the Invention

[0005] The purpose of the embodiments of this application is to provide a method and device for identifying fraud numbers to solve the problem that the traditional methods for identifying fraud numbers in the prior art are difficult to achieve accurate analysis and identification of fraud numbers.

[0006] In a first aspect, the embodiments of this application provide a method for identifying a fraud number, the method including:

[0007] Input the call behavior feature data and user service feature data of the target number into a first recognition model to obtain a first probability value that the target number is a fraud number;

[0008] Input the trajectory feature data of the target number within a preset time period into a second recognition model to obtain a second probability value that the target number is a fraud number;

[0009] Based on the evidence theory and the first probability value and the second probability value, determine whether the target number is a fraud number.

[0010] In a second aspect, an embodiment of the present application provides an identification device for fraud numbers, and the device includes:

[0011] A first recognition module, configured to input the call behavior feature data and user service feature data of a target number into a first recognition model to obtain a first probability value that the target number is a fraud number;

[0012] A second recognition module, configured to input the trajectory feature data of the target number within a preset time period into a second recognition model to obtain a second probability value that the target number is a fraud number;

[0013] A third recognition module, configured to determine whether the target number is a fraud number based on the evidence theory and the first probability value and the second probability value.

[0014] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory. The memory stores a program or instruction that can run on the processor. When the program or instruction is executed by the processor, the steps in the identification method for fraud numbers as described in the first aspect are implemented.

[0015] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps in the identification method for fraud numbers as described in the first aspect are implemented.

[0016] In the embodiments of the present application, multiple information sources such as call behavior feature data, user service feature data, and trajectory feature data are used for fraud number identification. The feature analysis is more abundant, which is beneficial to the precise analysis of fraud numbers and the improvement of the accuracy of fraud number identification. In addition, in the embodiments of the present application, the fraud numbers are first identified by the first recognition model and the second recognition model respectively, and then the recognition results of the two recognition models are fused based on the evidence theory to obtain the final recognition result. Compared with the method of identifying fraud numbers through a single recognition model, identifying fraud numbers based on two parallel recognition models can further improve the accuracy of fraud number identification while also improving the stability and fault tolerance of the model, which is helpful for practical scenario applications. Description of the Drawings

[0017] Figure 1 Schematic flowchart of the method for identifying fraud numbers provided by the embodiments of the present application;

[0018] Figure 2 Schematic flowchart of an example of the method for identifying fraud numbers provided by the embodiments of the present application;

[0019] Figure 3 Schematic flowchart of creating a second recognition model provided by the embodiments of the present application;

[0020] Figure 4 Schematic block diagram of the device for identifying fraud numbers provided by the embodiments of the present application. Detailed implementation manners

[0021] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0022] It should be understood that the "one embodiment" or "an embodiment" mentioned in the specification means that a specific feature, structure or characteristic related to the embodiment is included in at least one embodiment of the present application. Therefore, the appearances of "in one embodiment" or "in an embodiment" throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in any suitable manner in one or more embodiments.

[0023] In various embodiments of the present application, it should be understood that the sequence numbers of the steps do not mean an absolute order of execution. The execution order of each step should be determined according to its function and internal logic. Therefore, the sequence numbers of the steps should not absolutely limit the implementation process of the embodiments of the present application.

[0024] Next, in conjunction with the accompanying drawings, the method for identifying fraud numbers provided by the embodiments of the present application will be described in detail through specific embodiments and their application scenarios.

[0025] The embodiments of the present application provide a method for identifying fraud numbers, which is applied to an electronic device, and the electronic device can be a server or a terminal device.

[0026] As Figure 1 shown, the method for identifying fraud numbers may include:

[0027] Step 101: Input the call behavior characteristic data and user service characteristic data of the target number into the first recognition model to obtain a first probability value that the target number is a fraud number.

[0028] In the embodiment of the present application, a first recognition model is pre-constructed. The first recognition model can estimate the probability that a telephone number is a fraud number according to the call behavior feature data and user service feature data of the telephone number. Therefore, the call behavior feature data and user service feature data of the target number can be used as input data and input into the first recognition model, so as to obtain a first probability value that the target number is a fraud number.

[0029] Among them, the fraud number described here can be a GOIP number, that is, the first recognition model can estimate the probability that the target number is a GOIP number. Of course, the fraud number described here can also be other types of fraud numbers.

[0030] Among them, the call behavior feature data of the target number described here can include but are not limited to at least one of the following: average value of daily outgoing call duration, average value of daily outgoing call times, standard deviation of daily outgoing call duration, standard deviation of daily outgoing call times, average value of daily incoming call duration, average value of daily incoming call times, standard deviation of daily incoming call duration, standard deviation of daily incoming call times, ratio of outgoing call time to incoming call time, incoming call entropy value, aggregation average value of overseas numbers in the same base station (that is, the aggregation average value of overseas numbers under the same base station), proportion of incoming call numbers from different networks (assuming the local network is a telecommunications network, then different networks refer to mobile networks, Unicom networks, etc. other than the telecommunications network), and average value of call times within a preset period (such as the period from 8:00 am to 5:00 pm).

[0031] Among them, the user service feature data of the target number described here can include but are not limited to at least one of the following: user's length of stay on the network (the longer the length of stay on the network, the lower the possibility that the target number is a fraud number), package tariff (generally, fraudsters will not activate a package with a higher cost to increase the fraud cost. Therefore, the package tariff can also help to judge fraud numbers to a certain extent), and credit rating (the higher the credit rating, the lower the possibility that the target number is a fraud number).

[0032] Step 102: Input the trajectory feature data of the target number within a preset duration into a second recognition model to obtain a second probability value that the target number is a fraud number.

[0033] In the embodiment of the present application, a second recognition model is also pre-constructed. The second recognition model can estimate the probability that a telephone number is a fraud number according to the trajectory feature data of the telephone number. Therefore, the trajectory feature data of the target number can be used as input data and input into the second recognition model, so as to obtain a second probability value that the target number is a fraud number.

[0034] Among them, the trajectory feature data of the target number described herein may be a trajectory sequence formed by multiple trajectory points arranged in chronological order. Among them, each trajectory point may be composed of a location area code (LAC) and a cell identity (CI).

[0035] Step 103: Based on the evidence theory, as well as the first probability value and the second probability value, determine whether the target number is a fraud number.

[0036] In the embodiments of the present application, the diagnostic results of two recognition models can be integrated based on the evidence theory (i.e., the D-S evidence theory), so as to obtain a more comprehensive and accurate fusion recognition result.

[0037] In the embodiments of the present application, multiple information sources such as call behavior feature data, user service feature data, and trajectory feature data are used to identify fraud numbers. The analyzed data is richer, which is beneficial to the precise analysis of fraud numbers and the improvement of the accuracy of fraud number identification. In addition, in the embodiments of the present application, the fraud numbers are first identified by the first recognition model and the second recognition model respectively, and then the recognition results of the two recognition models are fused based on the evidence theory to obtain the final recognition result. Compared with the method of identifying fraud numbers through a single recognition model, identifying fraud numbers based on two parallel recognition models can not only further improve the accuracy of fraud number identification, but also improve the stability and fault tolerance of the model, which is helpful for practical scenario applications.

[0038] As an alternative embodiment, the embodiments of the present application also provide a method for establishing a first recognition model, as described below:

[0039] Step A1: Based on a preset algorithm, determine n preferred call behavior features among N call behavior features, and determine m preferred user service features among M user service features. Wherein, n is less than or equal to N, m is less than or equal to M, and M, N, m, and n are all positive integers.

[0040] Step A2: Use the data of the preferred call behavior features and the data of the preferred user service features of multiple numbers as the sample data of the initial first recognition model, and train the initial first recognition model to obtain the first recognition model.

[0041] The following further explains steps A1 and A2 respectively.

[0042] Step A1: Based on a preset algorithm, determine n preferred call behavior features among N call behavior features, and determine m preferred user service features among M user service features.

[0043] In the embodiments of the present application, some call behavior characteristics and user service characteristics that may be beneficial to identifying fraud numbers can be statistically analyzed first, and then based on a preset algorithm, screening is performed among the statistically obtained call behavior characteristics to obtain call behavior characteristics that are more helpful for identifying fraud numbers (i.e., preferred call behavior characteristics), and based on the preset algorithm, screening is performed among the statistically obtained user service characteristics to obtain user service characteristics that are more helpful for identifying fraud numbers, that is, preferred user service characteristics.

[0044] Optionally, step A1: determining n preferred call behavior characteristics among N call behavior characteristics and m preferred user service characteristics among M user service characteristics may include:

[0045] Step A11: Obtaining data of N call behavior characteristics and data of M user service characteristics of multiple numbers.

[0046] As Figure 2 shown in step 201 of , on the premise of ensuring user privacy and security, a call behavior data set and a user service data set of multiple numbers (such as all numbers) can be obtained first. Among them, the call behavior data set may include: data such as the calling and called numbers, call initiation time, call end time, etc. The user service data set may include: data such as user account opening time, package tariff, credit rating, etc.

[0047] After that, as Figure 2 shown in step 203 of , statistical analysis is performed on the data in the call behavior data set and the data in the user service data set.

[0048] Among them, statistical analysis of the data in the call behavior data set can generate data of N call behavior characteristics such as the average daily calling duration average value of the number, the average daily calling frequency average value, the standard deviation of the average daily calling duration, the standard deviation of the average daily calling frequency, the average daily called duration average value, the average daily called frequency average value, the standard deviation of the average daily called duration, the standard deviation of the average daily called frequency, the ratio of the calling time to the called time, the called entropy value, the average value of the aggregation of overseas numbers in the same base station, the off-net ratio of the called number, and the average value of the call frequency within a preset period.

[0049] Among them, statistical analysis of the data in the user service data set can generate data of M user service characteristics such as the user's on-net duration (determined according to the account opening time and the current time), package tariff, credit rating, etc. Optionally, the credit rating of each number can also be normalized, and continuous features such as package tariff can be discretized to unify the data format and facilitate the calculation of the preset algorithm.

[0050] Optionally, as Figure 2As shown, before performing step 203, step 202 can be executed first to clean the data in the call behavior dataset and the user service dataset, removing null values, abnormal data, etc. in the dataset. For example, due to possible delays in information synchronization, it is found that a phone number has been cancelled only after data statistics are completed, then all data of this phone number are null values and can be removed. Another example is that in a call record, when data such as the same calling and called numbers and the call end time being earlier than the call completion time appear, it indicates that these data are abnormal data and can be removed.

[0051] Step A12: Determine the confidence level of each call behavior feature based on the data of N call behavior features of multiple numbers using the random forest algorithm, and determine the confidence level of each user service feature based on the data of M user service features of multiple numbers using the random forest algorithm.

[0052] Step A13: Determine the preferred call behavior features among the N call behavior features according to the confidence level of each call behavior feature, and determine the preferred user service features among the M user service features according to the confidence level of each user service feature.

[0053] As Figure 2 shown in step 204, in the embodiment of the present application, call behavior feature screening can be performed based on the random forest algorithm (i.e., the preset algorithm). Specifically, first, through step A12, the data of the same call behavior feature of all numbers can be used as the input of the random forest algorithm to obtain the confidence level of this call behavior feature. Then, through step A13, according to the confidence level of each call behavior feature, determine the preferred call behavior features among the N call behavior features. For example, sort the call behavior features in descending order of confidence level, and determine the call behavior features ranked in the first preset position (such as the first 4, the first 5, etc.) as the preferred call behavior features; or, determine the call behavior features with a confidence level greater than or equal to the first preset value as the preferred call behavior features.

[0054] As Figure 2As shown in step 204, in the embodiments of the present application, user service feature screening can also be performed based on the random forest algorithm. Specifically, first, through step A12, the data of the same user service feature of all numbers can be used as the input of the random forest algorithm, so as to obtain the confidence of this user service feature. Then, through step A13, according to the confidence of each user service feature, the preferred user service features are determined among the N user service features. Specifically, the user service features can be sorted in descending order of confidence, and the user service features ranked in the top second preset positions (such as the top 3, top 4, etc.) are determined as the preferred user service features; or, the user service features with a confidence greater than or equal to the second preset value are determined as the preferred user service features.

[0055] In the embodiments of the present application, the confidence of call behavior features and user service features is obtained through the random forest algorithm, and each feature is evaluated and optimized according to the confidence value to screen out redundant features in the features. Through this screening method, the uncertainty and manual dependence in the traditional manual feature selection process can be improved.

[0056] Step A2: Use the data of the preferred call behavior features and the data of the preferred user service features of multiple numbers as the sample data of the initial first recognition model, and train the initial first recognition model to obtain the first recognition model.

[0057] In the embodiments of the present application, the data of the preferred call behavior features and the data of the preferred user service features of multiple numbers can be used as the sample data of the initial first recognition model. Specifically, based on the generality of the model, in the embodiments of the present application, the data of the preferred call behavior features and the data of the preferred user service features of each number can be serially fused to generate a feature vector (i.e., sample data, including positive samples and negative samples), so that the information in the feature vector is more comprehensive and the information loss rate is reduced. Then, the initial first recognition model can be trained using the feature vector as the sample data to obtain the first recognition model.

[0058] Optionally, the leave-one-out method can be used to process the feature vectors of the numbers participating in model establishment. According to a ratio of 3:1 or 4:1, the feature vectors are divided into a training set (i.e., the sample data described above) and a test set. The initial first recognition model receives the training set and performs classification training. According to the output result of the initial first recognition model, the model parameters can be adjusted to optimize the model, and finally the first recognition model is obtained. Then, the generalization ability of the first recognition model is evaluated using the test set.

[0059] Optionally, the first recognition model can be an identification model based on a Support Vector Machine (SVM) and supporting probability assignment output.

[0060] During the analysis of fraud number data, there is a situation of uneven data distribution. The sample of fraud numbers is much less than that of normal numbers, which has certain limitations for algorithms such as decision trees. SVM is a classification algorithm proposed under the condition of limited samples, which solves the machine learning problem in the case of small samples. Therefore, in the embodiments of this application, a model based on SVM is adopted to analyze the call behavior feature data and user service feature data.

[0061] However, the model built based on SVM generally has a threshold-free output, that is, the output result is a value used to represent "yes" (such as +1) or a value used to represent "no" (such as -1). In the embodiments of this application, the recognition model built based on SVM can be transformed so that the model can output specific probability values, realizing the probability assignment output for fraud number recognition. In this way, compared with the traditional threshold-free output, the rigor and applicability of the model can be further improved, enabling the model to be applicable to more application scenarios.

[0062] Optionally, in the embodiments of this application, the Plat method can be used to transform the recognition model built based on SVM to obtain an initial first recognition model, which is defined as follows:

[0063] f(x)=∑ j y j β j k(x j ,)+c(1)

[0064] O(x)=f(x)+b(2)

[0065]

[0066] Among them, f(x) is the decision function expression of the support vector machine; x j is the sample data, y j is the label value corresponding to the sample data x j (the label value used to represent "yes" or the label value used to represent "no"); k(x j , x) is the kernel function of the support vector machine; β j is the Lagrange multiplier; O(x) is the threshold-free output of the support vector machine; P is the probability assignment output after the transformation of the support vector machine; c, b, Q, and W all represent a preset constant.

[0067] After obtaining the initial first recognition model, the initial first recognition model is trained with the preferred call behavior feature data and the preferred user service feature data as the sample data, and then the first recognition model based on SVM and supporting probability assignment output can be obtained, such as Figure 2as shown in step 205 in []. After that, the target number can be identified as a fraud number through the first identification model, such as Figure 2 as shown in step 206 in [].

[0068] As an optional embodiment, the embodiment of the present application also provides a method for establishing a second identification model, which is described as follows:

[0069] Step B1: Obtain the trajectory feature data of multiple numbers within a preset time period.

[0070] Such as Figure 2 in step 201 in [] and Figure 3 as shown in step 301 in [], on the premise of ensuring user privacy security, a trajectory data set of multiple numbers (such as all numbers) can be obtained. Among them, the trajectory data set can include data such as LAC, CI, and base station longitude and latitude.

[0071] Such as Figure 2 as shown in step 202 in []. After obtaining the trajectory data set, the data in the trajectory data set can be cleaned to remove null values, abnormal data, etc. in the data set. For example, due to possible delays in information synchronization, it is found that the phone number has been cancelled only after the data statistics are completed, then all the data of the phone number are null values and can be removed. For another example, abnormal data can be cleared based on the base station longitude and latitude. When the base station latitude exceeds 90°, the corresponding LAC and CI data are considered abnormal data and can be removed.

[0072] Such as Figure 2 as shown in step 207 in []. After completing the data cleaning, the data in the trajectory data set can be processed to obtain the trajectory feature data. Specifically, the data in the trajectory data set can be first subjected to preliminary statistical analysis, and the LAC and CI of all numbers are associated with the corresponding collection time points. Then, since some fraud numbers (such as GOIP numbers) have characteristics such as overlapping location trajectories and number clustering, in order to facilitate the analysis of the trajectory feature data of the numbers, it can be as Figure 3 as shown in step 302 in [], convert the LAC and CI of all numbers into trajectory points represented in the form of two-dimensional coordinates: (lac, ci). After that, as Figure 3 as shown in step 303 in [], taking a preset time period (such as 24 hours) as a cycle, summarize and sort all (lac, ci) coordinates of each number in chronological order to generate the trajectory sequence of each number. The trajectory sequence can be expressed as: P i =(P i1 , P i2 , …, P in ), where P in =(lac in , ci in) where i is the identifier of a certain number and n is the number of trajectory points within a preset duration. The trajectory sequence described here is the trajectory feature data to be obtained in this step.

[0073] Since LAC and CI can determine the sector position under a wireless base station, according to expert experience, they can be directly used as typical trajectory features of a number.

[0074] Step B2: Determine the length value of the longest common subsequence between the trajectory sequences of any two numbers among multiple numbers.

[0075] In this step, using the Longest Common Subsequence (LCSS) algorithm, calculate the length value LCSS(P u , P t ) of the longest common subsequence of the trajectory sequences of any two numbers within a preset duration, as described in step 304 of Figure 3 .

[0076] Among them, LCSS(P u , P t ) is defined as follows:

[0077]

[0078] Among them, u and t are identifiers of any two numbers; A represents the trajectory sequence of number u, n is the maximum value of the number of trajectory points in trajectory sequence A, B represents the trajectory sequence of number t, n' is the maximum value of the number of trajectory points in trajectory sequence B, n and n' can be the same or different; φ represents the empty set; dist(P un , P tn′ ) represents the distance value between trajectory point P un and trajectory point P tn’ , and γ is the first preset similarity threshold.

[0079] The above LCSS(P u , P t ) formula indicates that when the number of trajectory points in the trajectory sequence of any one number is 0 (i.e., A = φ or B = φ, or A = φ and B = φ), the length value LCSS(P u , P t ) of the longest common subsequence between the trajectory sequences of the two numbers = 0; when A ≠ φ and B ≠ φ, if the distance between the last trajectory points in the two trajectory sequences A and B is less than the first preset similarity threshold γ, it means the two trajectory points are the same, then the length value LCSS(P u , P t ) of the longest common subsequence of the trajectory sequences of the two numbers = 1 + CSS(P un-1 , P tn′-1), otherwise LCSS(P u , P t ) = max(LCSS(P un , P tn′-1 ), LCSS(P un-1 , P tn′ ))

[0080] Step B3: Determine the trajectory similarity between any two numbers according to the length value of the longest common subsequence between any two numbers

[0081] In this step, according to the preset normalization formula, the length value of the longest common subsequence between any two numbers (i.e., LCSS(P u , P t )) can be normalized to obtain the trajectory similarity L(P u , P t ), as shown in step 305 of Figure 3

[0082] Step B4: Screen out all numbers whose trajectory similarity is greater than the second preset similarity threshold

[0083] Since some fraud numbers (such as GOIP numbers) have the characteristic of overlapping location trajectories, while the trajectory overlap degree between normal phone numbers is relatively small, therefore, in this embodiment of the present application, based on this, the numbers can be screened, and the numbers with trajectory similarity L(P u , P t ) > the second preset similarity threshold K 1 are screened out, as shown in step 306 of Figure 3

[0084] Step B5: Cluster the numbers screened in step B4 based on LAC and CI to form multiple clusters

[0085] Among them, each cluster corresponds to a trajectory point, that is, the numbers included in a cluster are clustered into a group because they have a same trajectory point. Since a number can include multiple different trajectory points, the same number can appear in multiple clusters

[0086] Since some fraud numbers (such as GOIP numbers) have characteristics such as number clustering, therefore, in this embodiment of the present application, based on this, the numbers screened in step B4 are clustered to form clusters N 1 to N n , as shown in Figure 3 ​​as shown in step 307. Specifically, clustering can be performed according to the trajectory points and the pairwise correlation principle, that is: for any two numbers, if there is a same trajectory point (lac, ci) between them, then the two numbers are classified into the same clustering group.

[0087] Among them, the more the total number of numbers included in a certain clustering group, the greater the possibility that the numbers in this clustering group are fraud numbers.

[0088] Step B6: Define the longest common subsequence length value calculation formula (i.e., the LCSS(P u , P t ) formula), the trajectory similarity calculation formula (i.e., the preset normalization formula), the number screening algorithm (screening based on trajectory similarity), the clustering algorithm, the clustering groups obtained by clustering, and the preset basic probability assignment output formula as the second recognition model.

[0089] In the embodiments of the present application, the basic probability assignment output formula of the second recognition model (i.e., the preset basic probability assignment output formula) can be defined, as Figure 3 shown in step 308. This basic probability assignment output formula can perform the basic probability assignment output of the number according to the maximum value in the total number of numbers in the clustering group where the number is located.

[0090] Among them, the basic probability assignment output formula is:

[0091]

[0092] Among them, g i represents the basic probability assignment output that the i-th number is a fraud number; h is the maximum value in the total number of numbers in the clustering group where the i-th number is located. For example, if the i-th number appears in three clustering groups, and the total number of numbers in these three clustering groups are 7, 8, and 9 respectively, then h takes the maximum value of 9. Z represents the first clustering threshold, C represents the second clustering threshold, Z is less than C, and the numerical values of Z and C can be determined in advance according to expert experience. K 2 is a preset probability value greater than 0 and less than 1, which can be determined in advance according to expert experience.

[0093] Through the second recognition model established above, the fraud number recognition of the target number can be performed, as Figure 2 shown in step 209. Therefore, the embodiments of the present application also provide an implementation manner for obtaining the probability that the target number is a fraud number based on the second recognition model, as follows:

[0094] Step 102: Input the trajectory feature data of the target number within a preset time period into the second recognition model to obtain a second probability value that the target number is a fraud number, which may include:

[0095] Step C1: Input the trajectory feature data of the target number within the preset duration into the second recognition model.

[0096] Step C2: Through the second recognition model, determine the length value of the longest common subsequence between the trajectory sequences of any two target numbers, and determine the trajectory similarity between any two target numbers according to the length value of the longest common subsequence. Wherein, the number of target numbers is at least two.

[0097] Step C3: Through the second recognition model, for two target numbers with a trajectory similarity greater than the preset similarity threshold, determine the cluster group to which they belong based on the trajectory points. The cluster group mentioned here is obtained through Step B5.

[0098] Step C4: Through the second recognition model, determine the maximum value among the total number of numbers included in the cluster group to which each target number belongs, and determine the second probability value that the target number is a fraud number according to the size relationship between the maximum value and the preset clustering threshold.

[0099] Wherein, the preset clustering threshold mentioned here may include: the first clustering threshold (i.e., the aforementioned Z) and the second clustering threshold (i.e., the aforementioned C).

[0100] Optionally, "determine the second probability value that the target number is a fraud number according to the size relationship between the maximum value and the preset clustering threshold" in Step C4 may include:

[0101] Step C41: When the maximum value is less than the first clustering threshold, determine that the second probability value that the target number is a fraud number is 0.

[0102] Step C42: When the maximum value is greater than or equal to the first clustering threshold and less than the second clustering threshold, determine that the second probability value that the target number is a fraud number is the preset probability value. Wherein, the preset probability value mentioned here is the aforementioned K 2 .

[0103] Step C43: When the maximum value is greater than or equal to the second clustering threshold, determine that the second probability value that the target number is a fraud number is 1.

[0104] Based on the in-depth analysis of the trajectory characteristics of fraud numbers, the embodiments of the present application fully exploit and utilize the location information (i.e., trajectory points) of the numbers. First, the LCSS algorithm is used to realize the similarity measurement between number trajectories, and then combined with the idea of number trajectory aggregation to further judge suspected fraud numbers, providing another judgment basis for the identification of fraud numbers.

[0105] Wherein, after obtaining the recognition results of the first recognition model and the second recognition model for the target number, it can be as Figure 2As shown in step 210 in [reference], the recognition results of the first recognition model and the second recognition model for the target number, which are decision-level information, are fused, and then the final recognition result is output according to the fusion result, as Figure 2 shown in step 211 in [reference]. The following further explains this part of the content.

[0106] As an alternative embodiment, step 103: determining whether the target number is a fraud number based on the evidence theory and the first probability value and the second probability value may include:

[0107] Step D1: Establish a recognition framework according to the possible judgment results of the target number.

[0108] Among them, the recognition framework (also called the complete set) may include: the first element that the target number is a fraud number and the second element that the target number is a normal number.

[0109] In the embodiment of the present application, the recognition framework can be represented by θ, then θ = {θ 1 , θ 2}, where θ 1 represents the first element that the target number is a fraud number, and θ 2 represents the second element that the target number is a normal number.

[0110] Step D2: Construct a first evidence body based on the recognition framework according to the output result of the first recognition model for the target number, and construct a second evidence body based on the recognition framework according to the output result of the second recognition model for the target number.

[0111] Among them, the first evidence body includes: the first element and the first probability value; the second evidence body includes: the first element and the second probability value.

[0112] For any subset A of the recognition framework, as long as m(A) > 0, then A is called the focal element of the evidence, where m(A) is the basic probability assignment of the subset A (which can also be called the proposition A). The binary body (A, m(A)) composed of the focal element of the evidence and its basic probability assignment is called the evidence body.

[0113] In the embodiment of the present application, the subsets of the recognition framework θ include: {θ 1}, {θ 2}, and {θ 1 , θ 2} (i.e., θ).

[0114] For the first recognition model, m 1 (θ 1 ) = x, m 1 (θ 2) = 1 - x, m 1 (θ) = 0, where x is the first probability value, i.e., the output result of the first recognition model for the target number.

[0115] For the second recognition model, m 2 (θ 1 ) = y, m 2 (θ 2 ) = 1 - y, m 2 (θ) = 0, where y is the second probability value, i.e., the output result of the second recognition model for the target number.

[0116] In the embodiments of the present application, the first body of evidence may be (θ 1 , m 1 (θ 1 )) and the second body of evidence may be (θ 1 , m 2 (θ 1 ).

[0117] Step D3: Determine the combined basic probability assignment under the joint action of the first body of evidence and the second body of evidence according to the evidence theory combination function.

[0118] In the embodiments of the present application, the evidence theory combination function is:

[0119]

[0120] where m h (θ 1 ) is the combined basic probability assignment under the joint action of the first body of evidence and the second body of evidence, m 1 (θ 1 ) represents the basic probability assignment of the first body of evidence, m 2 (θ 1 ) represents the basic probability assignment of the second body of evidence, and K is the normalization constant or conflict coefficient, representing the degree of conflict between the evidences. The normalization constant increases as the degree of conflict increases.

[0121] Step D4: Determine the belief function value and the plausibility function value under the joint action of the first body of evidence and the second body of evidence according to the combined basic probability assignment.

[0122] where the belief function value Bel({θ 1}) = m h (θ 1 ).

[0123] where the plausibility function value

[0124] Step D5: Determine the credibility interval under the combined action of the first evidence body and the second evidence body according to the belief function value and the plausibility function value.

[0125] Among them, the lower limit value of the credibility interval is the belief function value, and the upper limit value is the plausibility function value, that is, the credibility interval is [Bel, pl].

[0126] Step D6: Determine whether the target number is a fraud number according to the preset decision rule and the credibility interval under the combined action of the first evidence body and the second evidence body.

[0127] Among them, the preset decision rules may include: the maximum belief decision rule, the absolute support decision rule, and the uncertainty limitation decision rule. When the above three decision rules are satisfied at the same time, the correct decision result can be obtained.

[0128] The above is the description of the fraud number identification method provided by the embodiments of the present application.

[0129] In summary, the technical solution provided by the embodiments of the present application uses multiple information sources such as call behavior feature data, user service feature data, and trajectory feature data to identify fraud numbers, and the analyzed data is richer, which is beneficial to the precise analysis of fraud numbers and the improvement of the accuracy of fraud number identification. In addition, the embodiments of the present application first identify fraud numbers through the first identification model and the second identification model respectively, and then fuse the identification results of the two identification models based on the evidence theory to obtain the final identification result. Compared with identifying fraud numbers through a single identification model and identifying fraud numbers based on two parallel identification models, while further improving the accuracy of fraud number identification, it can also improve the stability and fault tolerance of the model, which is helpful for practical scenario applications. In short, the fraud number identification method provided by the embodiments of the present application can better help operators discover fraud numbers and improve the operator's service quality and customer service quality to a certain extent.

[0130] The above introduces the fraud number identification method provided by the embodiments of the present application. Next, the fraud number identification device provided by the embodiments of the present application will be introduced with reference to the drawings.

[0131] As Figure 4 shown, the embodiments of the present application also provide a fraud number identification device, which is applied to an electronic device.

[0132] Among them, the fraud number identification device may include:

[0133] The first identification module 401 is configured to input the call behavior feature data and user service feature data of the target number into the first identification model to obtain the first probability value that the target number is a fraud number.

[0134] The second identification module 402 is configured to input the trajectory feature data of the target number within a preset time period into a second identification model to obtain a second probability value that the target number is a fraud number.

[0135] The third identification module 403 is configured to determine whether the target number is a fraud number based on the evidence theory and the first probability value and the second probability value.

[0136] Optionally, the first identification model is an identification model constructed based on a support vector machine and supporting probability assignment output.

[0137] Optionally, the device may further include:

[0138] A screening module is configured to determine n preferred call behavior features from N call behavior features and m preferred user service features from M user service features based on a preset algorithm.

[0139] A model training module is configured to use the data of the preferred call behavior features and the data of the preferred user service features of multiple numbers as sample data of an initial first identification model, and train the initial first identification model to obtain the first identification model.

[0140] Optionally, the screening module may include:

[0141] An acquisition unit is configured to acquire the data of the N call behavior features and the data of the M user service features of multiple numbers.

[0142] A first determination unit is configured to determine the confidence level of each call behavior feature based on the random forest algorithm and the data of the N call behavior features of multiple numbers, and determine the confidence level of each user service feature based on the random forest algorithm and the data of the M user service features of multiple numbers.

[0143] A screening unit is configured to determine the preferred call behavior features from the N call behavior features according to the confidence level of each call behavior feature, and determine the preferred user service features from the M user service features according to the confidence level of each user service feature.

[0144] Optionally, the second identification module may include:

[0145] An input unit is configured to input the trajectory feature data of the target number within a preset time period into a second identification model.

[0146] Wherein, the trajectory feature data is a trajectory sequence formed by a plurality of trajectory points arranged in chronological order, and each trajectory point is composed of a location area code and a cell identification code.

[0147] A second determination unit, configured to determine, by using the second recognition model, the length value of the longest common subsequence between the trajectory sequences of any two of the target numbers, and determine the trajectory similarity between any two of the target numbers according to the length value of the longest common subsequence.

[0148] Wherein, the number of the target numbers is at least two.

[0149] A third determination unit, configured to, by using the second recognition model, for two of the target numbers whose trajectory similarity is greater than a preset similarity threshold, determine the cluster group to which they belong according to the trajectory points; wherein, the cluster groups are obtained by pre-clustering, each cluster group corresponds to a trajectory point, and the same number belongs to at least one cluster group.

[0150] The third determination unit is configured to, by using the second recognition model, determine the maximum value among the total number of numbers included in the cluster group to which each target number belongs, and determine a second probability value that the target number is a fraud number according to the magnitude relationship between the maximum value and a preset clustering threshold.

[0151] Optionally, the preset clustering threshold may include: a first clustering threshold and a second clustering threshold, and the first clustering threshold is less than the second clustering threshold.

[0152] The third determination unit may include:

[0153] A first determination subunit, configured to determine that the second probability value that the target number is a fraud number is 0 when the maximum value is less than the first clustering threshold.

[0154] A second determination subunit, configured to determine that the second probability value that the target number is a fraud number is a preset probability value when the maximum value is greater than or equal to the first clustering threshold and less than the second clustering threshold.

[0155] Wherein, the preset probability value is greater than 0 and less than 1.

[0156] A third determination subunit, configured to determine that the second probability value that the target number is a fraud number is 1 when the maximum value is greater than or equal to the second clustering threshold.

[0157] Optionally, the third recognition module may include:

[0158] A recognition framework establishment unit, configured to establish a recognition framework according to possible judgment results of the target numbers.

[0159] Wherein, the recognition framework includes: a first element that the target number is a fraud number and a second element that the target number is a normal number.

[0160] An evidence body construction unit is configured to construct a first evidence body based on the identification framework according to the output result of the target number by the first identification model, and construct a second evidence body based on the identification framework according to the output result of the target number by the second identification model.

[0161] Wherein, the first evidence body includes: the first element and the first probability value; the second evidence body includes: the second element and the second probability value.

[0162] A fourth determination unit is configured to determine a combined basic probability assignment under the joint action of the first evidence body and the second evidence body according to an evidence theory combination function.

[0163] A fifth determination unit is configured to determine a belief function value and a plausibility function value under the joint action of the first evidence body and the second evidence body according to the combined basic probability assignment.

[0164] A sixth determination unit is configured to determine a credibility interval under the joint action of the first evidence body and the second evidence body according to the belief function value and the plausibility function value.

[0165] A seventh determination unit is configured to determine whether the target number is a fraud number according to a preset decision rule and the credibility interval.

[0166] The fraud number identification device provided by the embodiments of the present application can implement Figure 1 each process implemented by the fraud number identification device in the method embodiment shown. To avoid repetition, it will not be elaborated here.

[0167] In the embodiments of the present application, various information sources such as call behavior feature data, user service feature data, and trajectory feature data are used for fraud number identification. The analyzed data is richer, which is beneficial to the accurate analysis of fraud numbers and the improvement of the accuracy of fraud number identification. In addition, in the embodiments of the present application, the fraud numbers are first identified by the first identification model and the second identification model respectively, and then the identification results of the two identification models are fused based on the evidence theory to obtain the final identification result. Compared with identifying fraud numbers through a single identification model, identifying fraud numbers based on two parallel identification models can further improve the accuracy of fraud number identification while also improving the stability and fault tolerance of the model, which is helpful for practical scenario applications. In short, the fraud number identification method provided by the embodiments of the present application can better help the operator discover fraud numbers and improve the operator's service quality of business processing and customer service to a certain extent.

[0168] An embodiment of the present application further provides an electronic device, including a processor and a memory. A program or instruction that can run on the processor is stored on the memory. When the program or instruction is executed by the processor, each step of the above-described embodiment of the fraud number recognition method is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be elaborated here.

[0169] An embodiment of the present application further provides a readable storage medium. A program or instruction is stored on the readable storage medium. When the program or instruction is executed by the processor, each process of the above-described embodiment of the fraud number recognition method is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be elaborated here.

[0170] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0171] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM, RAM, magnetic disk, optical disc, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0172] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present application.

Claims

1. A method for identifying fraud numbers, characterized in that, applied to a server or a terminal device, the method includes: Input the call behavior characteristic data and user service characteristic data of the target number into a first recognition model to obtain a first probability value that the target number is a fraud number; Input the trajectory characteristic data of the target number within a preset time period into a second recognition model to obtain a second probability value that the target number is a fraud number; Based on the evidence theory and the first probability value and the second probability value, determine whether the target number is a fraud number; The determining whether the target number is a fraud number based on the evidence theory and the first probability value and the second probability value includes: According to the possible judgment results of the target number, establish an identification framework; wherein, the identification framework includes: a first element that the target number is a fraud number and a second element that the target number is a normal number; According to the output result of the first recognition model for the target number, construct a first evidence body based on the identification framework, and according to the output result of the second recognition model for the target number, construct a second evidence body based on the identification framework; wherein, the first evidence body includes: the first element and the first probability value; the second evidence body includes: the first element and the second probability value; According to the evidence theory combination function, determine the combined basic probability assignment under the joint action of the first evidence body and the second evidence body; According to the combined basic probability assignment, determine the belief function value and plausibility function value under the joint action of the first evidence body and the second evidence body; According to the belief function value and the plausibility function value, determine the credibility interval under the joint action of the first evidence body and the second evidence body; According to a preset decision rule and the credibility interval, determine whether the target number is a fraud number.

2. The method for identifying fraud numbers according to claim 1, characterized in that, The first recognition model is a recognition model constructed based on a support vector machine and supporting probability assignment output.

3. The method for identifying fraud numbers according to claim 1 or 2, characterized in that, Before inputting the call behavior characteristic data and user service characteristic data of the target number into the first recognition model to obtain a first probability value that the target number is a fraud number, the method further includes: Based on a preset algorithm, determine n preferred call behavior characteristics among N call behavior characteristics, and determine m preferred user service characteristics among M user service characteristics; Use the data of the preferred call behavior characteristics and the data of the preferred user service characteristics of multiple numbers as sample data of an initial first recognition model, and train the initial first recognition model to obtain the first recognition model.

4. The method for identifying fraud numbers according to claim 3, characterized in that, The determining n preferred call behavior characteristics among N call behavior characteristics and determining m preferred user service characteristics among M user service characteristics includes: Obtain data of the N call behavior characteristics of multiple numbers and data of the M user service characteristics; Based on the random forest algorithm and the data of the N call behavior characteristics of the multiple numbers, determine the confidence level of each call behavior characteristic, and based on the random forest algorithm and the data of the M user service characteristics of the multiple numbers, determine the confidence level of each user service characteristic; According to the confidence level of each call behavior characteristic, determine the preferred call behavior characteristic among the N call behavior characteristics, and according to the confidence level of each user service characteristic, determine the preferred user service characteristic among the M user service characteristics.

5. The method for identifying a fraud number according to claim 1, wherein, the inputting the trajectory feature data of the target number within a preset time period into a second identification model to obtain a second probability value that the target number is a fraud number includes: Input the trajectory feature data of the target number within a preset time period into the second identification model; wherein, the trajectory feature data is a trajectory sequence formed by a plurality of trajectory points arranged in chronological order, and each trajectory point is composed of a location area code and a cell identification code; Through the second identification model, determine the length value of the longest common subsequence between the trajectory sequences of any two target numbers, and determine the trajectory similarity between any two target numbers according to the length value of the longest common subsequence; the number of target numbers is at least two; Through the second identification model, for two target numbers with a trajectory similarity greater than a preset similarity threshold, determine the cluster group to which they belong according to the trajectory points; wherein, the cluster group is pre-clustered, each cluster group corresponds to a trajectory point, and the same number belongs to at least one cluster group; Through the second identification model, determine the maximum value among the total number of numbers included in the cluster group to which each target number belongs, and determine the second probability value that the target number is a fraud number according to the magnitude relationship between the maximum value and a preset clustering threshold.

6. The method for identifying a fraud number according to claim 5, wherein, the preset clustering threshold includes: a first clustering threshold and a second clustering threshold, and the first clustering threshold is less than the second clustering threshold; the determining the second probability value that each target number is a fraud number according to the magnitude relationship between the maximum value and the preset clustering threshold includes: When the maximum value is less than the first clustering threshold, determine that the second probability value that the target number is a fraud number is 0; When the maximum value is greater than or equal to the first clustering threshold and less than the second clustering threshold, determine that the second probability value that the target number is a fraud number is a preset probability value; wherein, the preset probability value is greater than 0 and less than 1; When the maximum value is greater than or equal to the second clustering threshold, determine that the second probability value that the target number is a fraud number is 1.

7. An apparatus for identifying a fraud number, wherein, applied to a server or a terminal device, the apparatus includes: The first recognition module is configured to input the call behavior feature data and user service feature data of the target number into the first recognition model to obtain a first probability value that the target number is a fraud number; The second recognition module is configured to input the trajectory feature data of the target number within a preset duration into the second recognition model to obtain a second probability value that the target number is a fraud number; The third recognition module is configured to determine whether the target number is a fraud number based on the evidence theory and the first probability value and the second probability value; The third recognition module includes: An identification framework establishment unit, configured to establish an identification framework according to the possible judgment results of the target number; wherein, the identification framework includes: a first element that the target number is a fraud number and a second element that the target number is a normal number; An evidence body construction unit, configured to construct a first evidence body based on the identification framework according to the output result of the first recognition model for the target number, and construct a second evidence body based on the identification framework according to the output result of the second recognition model for the target number; wherein, the first evidence body includes: the first element and the first probability value; the second evidence body includes: the first element and the second probability value; A fourth determination unit, configured to determine the combined basic probability assignment under the joint action of the first evidence body and the second evidence body according to the evidence theory combination function; A fifth determination unit, configured to determine the belief function value and plausibility function value under the joint action of the first evidence body and the second evidence body according to the combined basic probability assignment; A sixth determination unit, configured to determine the credibility interval under the joint action of the first evidence body and the second evidence body according to the belief function value and the plausibility function value; A seventh determination unit, configured to determine whether the target number is a fraud number according to a preset decision rule and the credibility interval.

8. The fraud number identification device according to claim 7, wherein, The first recognition model is a recognition model constructed based on a support vector machine and supporting probability assignment output.

9. The fraud number identification device according to claim 7 or 8, wherein, The device further includes: A screening module, configured to determine n preferred call behavior features among N call behavior features and m preferred user service features among M user service features based on a preset algorithm; A model training module, configured to use the data of the preferred call behavior features and the data of the preferred user service features of multiple numbers as sample data of the initial first recognition model, and train the initial first recognition model to obtain the first recognition model.

10. The fraud number identification device according to claim 9, wherein, The screening module includes: An acquisition unit, configured to acquire the data of the N call behavior features and the data of the M user service features of multiple numbers; A first determination unit, configured to determine the confidence level of each call behavior feature based on the random forest algorithm and the data of the N call behavior features of the multiple numbers, and determine the confidence level of each user service feature based on the random forest algorithm and the data of the M user service features of the multiple numbers; A screening unit, configured to determine the preferred call behavior features from the N call behavior features according to the confidence level of each call behavior feature, and determine the preferred user service features from the M user service features according to the confidence level of each user service feature.

11. The fraud number identification device according to claim 7, wherein, the second identification module includes: An input unit, configured to input the trajectory feature data of the target number within a preset time period into the second identification model; wherein, the trajectory feature data is a trajectory sequence formed by a plurality of trajectory points arranged in chronological order, and each trajectory point is composed of a location area code and a cell identification code; A second determination unit, configured to determine the length value of the longest common subsequence between the trajectory sequences of any two target numbers through the second identification model, and determine the trajectory similarity between any two target numbers according to the length value of the longest common subsequence; the number of target numbers is at least two; A third determination unit, configured to determine, through the second identification model, the cluster group to which two target numbers with a trajectory similarity greater than a preset similarity threshold belong according to the trajectory points; wherein, the cluster group is pre-clustered, each cluster group corresponds to a trajectory point, and the same number belongs to at least one cluster group; A third determination unit, configured to determine, through the second identification model, the maximum value of the total number of numbers included in the cluster group to which each target number belongs, and determine a second probability value that the target number is a fraud number according to the magnitude relationship between the maximum value and a preset clustering threshold.

12. The fraud number identification device according to claim 11, wherein, the preset clustering threshold includes: a first clustering threshold and a second clustering threshold, and the first clustering threshold is less than the second clustering threshold; the third determination unit includes: A first determination subunit, configured to determine that the second probability value that the target number is a fraud number is 0 when the maximum value is less than the first clustering threshold; A second determination subunit, configured to determine that the second probability value that the target number is a fraud number is a preset probability value when the maximum value is greater than or equal to the first clustering threshold and less than the second clustering threshold; wherein, the preset probability value is greater than 0 and less than 1; A third determination subunit, configured to determine that the second probability value that the target number is a fraud number is 1 when the maximum value is greater than or equal to the second clustering threshold.

13. An electronic device, wherein, It includes a processor and a memory. The memory stores programs or instructions that can run on the processor. When the programs or instructions are executed by the processor, the steps of the method for identifying fraud numbers as described in any one of claims 1 to 6 are implemented.

14. A readable storage medium, characterized in that programs or instructions are stored on the readable storage medium. When the programs or instructions are executed by a processor, the steps of the method for identifying fraud numbers as described in any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Fraud number identification method, device and equipment, and computer readable storage medium

    CN109547942A

  • Method, device and system for constructing identification model and identifying, and electronic equipment

    CN111800546A

  • Fraud number identification method and device, computer equipment and storage medium

    CN112291424A