Fraud risk number marking method and device, electronic equipment and medium

By determining the activation information of fraud-risk numbers and using the k-sigma algorithm and the minimum circle covering algorithm to divide risk areas, the problem of insufficient timeliness in fraud detection in existing technologies is solved, and the ability to mark numbers before activation is achieved, thus improving early warning capabilities.

CN115758065BActive Publication Date: 2026-07-31BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING XIAOMI MOBILE SOFTWARE CO LTD
Filing Date
2022-10-26
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Current technologies for detecting fraudulent phone numbers are typically reactive, resulting in late detection and an inability to provide early warnings, which can lead to economic or material losses.

Method used

By determining the activation information of sample numbers, the k-sigma algorithm and the minimum circle covering algorithm are used to divide the target risk area, and the number is marked as a fraud risk number based on its geographical location when it is activated.

Benefits of technology

It enables the marking of fraudulent numbers before activation, improving the comprehensiveness and timeliness of fraudulent number marking and reducing economic losses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115758065B_ABST
    Figure CN115758065B_ABST
Patent Text Reader

Abstract

This disclosure relates to a method, apparatus, electronic device, and medium for marking fraud-risk numbers. The method includes: determining activation information of sample numbers with fraud risk; determining a target risk area based on the activation information of each sample number in a sample set, wherein the target risk area is determined by dividing the sample numbers in the sample set into regions and based on the geographical location represented by the activation information of the sample numbers in each region; when a number is activated, if the geographical location represented by the activation information is within the target risk area, the number is marked as a target number with fraud risk. The target risk area can be determined based on the activation information of the fraud-risk sample numbers, and numbers whose corresponding geographical location is within the target risk area at the time of activation can be marked as numbers with fraud risk. This means marking occurs before fraud occurs, rather than during or after the fraud, improving the timeliness of fraud number marking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of telecommunications technology, and in particular to a method, apparatus, electronic device, and medium for marking fraud risk numbers. Background Technology

[0002] In technologies related to identifying fraudulent risk numbers, detection is usually based on call records, data usage, and SMS messages. Therefore, by the time a fraudulent risk number is detected, the fraud is already underway, and the user may have already been defrauded, resulting in economic or material losses. It is a reactive defense measure with relatively late detection timeliness. Summary of the Invention

[0003] To overcome the problems existing in related technologies, this disclosure provides a method, apparatus, electronic device, and medium for marking fraud risk numbers.

[0004] According to a first aspect of the present disclosure, a method for marking fraud risk numbers is provided, comprising:

[0005] Identify activation information for sample phone numbers that pose a risk of fraud;

[0006] Based on the activation information of each sample number in the sample set, a target risk area is determined, wherein the target risk area is determined by dividing the sample numbers in the sample set into regions and then using the geographical location represented by the activation information of the sample numbers in each region.

[0007] When a number is activated, if the geographical location of the activation information is within the target risk area, the number will be marked as a target number with a risk of fraud.

[0008] Optionally, the step of determining the target risk area based on the activation information of each sample number in the sample set includes:

[0009] Based on the activation information corresponding to each sample number in the sample set, sample numbers with fraud risk are divided into regions;

[0010] For any of the regions, calculate the mean and variance of the geographical locations represented by the activation information corresponding to the sample numbers within the region;

[0011] Based on the k-sigma algorithm, the target risk area is determined from the region according to the average value and variance corresponding to each region, and the geographical location represented by the activation information of the sample number in the region.

[0012] Optionally, the step of determining the target risk area from the region based on the k-sigma algorithm, according to the mean and variance corresponding to each region, and the geographical location represented by the activation information of the sample numbers within the region, includes:

[0013] Repeat the following steps:

[0014] Based on the value of k in the k-sigma algorithm, the average value and variance corresponding to the region, abnormal sample numbers whose geographical locations are outside the Laida confidence interval corresponding to the value are removed, and normal sample numbers corresponding to the region that are within the Laida confidence interval corresponding to the value are obtained.

[0015] Based on the number of offline sample numbers and the number of non-offline sample numbers in the normal sample numbers corresponding to the region, determine whether the value of k needs to be adjusted;

[0016] If it is necessary to adjust the value of k, adjust the value of k within a preset range of preset k values ​​according to a preset step size until it is determined that it is not necessary to adjust the value of k, and obtain the target value of k.

[0017] Based on the geographical location represented by the activation information of the normal sample number corresponding to the target value of k, the target risk area is determined from the region.

[0018] Optionally, the step of determining the target risk area from the region based on the geographical location represented by the activation information of the normal sample number corresponding to the target value of k includes:

[0019] The minimum circle covering algorithm is used to determine the minimum circle that covers the geographical location of the activation information of the normal sample number corresponding to the target value of k.

[0020] The area covered by the smallest circle is taken as the target risk area corresponding to the area.

[0021] Optionally, the step of determining whether the value of k needs to be adjusted based on the number of off-network sample numbers and the number of non-off-network sample numbers in the normal sample numbers of the region includes:

[0022] The churn rate of the region is calculated based on the number of churned sample numbers and the number of non-churned sample numbers in the normal sample numbers of the region.

[0023] Based on the number of non-offline sample numbers and the churn rate, determine whether the value of k needs to be adjusted.

[0024] Optionally, the step of determining whether to adjust the value of k based on the number of non-offline sample numbers and the churn rate includes:

[0025] Based on the relationship between the number of non-offline sample numbers and the preset number threshold, and the relationship between the churn rate and the preset churn rate threshold, it is determined whether the value of k needs to be adjusted.

[0026] Optionally, the method includes:

[0027] The maximum value of the preset range of k is taken as the initial value of k;

[0028] The step of adjusting the value of k within a preset range by a preset step size when it is necessary to adjust the value of k until it is determined that no adjustment of the value of k is needed, and obtaining the target value of k, includes:

[0029] If it is necessary to adjust the value of k, the value of k is reduced within a preset range by a preset step size until it is determined that no adjustment of the value of k is needed, thus obtaining the target value of k.

[0030] Optionally, the method includes:

[0031] If the value of k is decreased within a preset range by a preset step size until the value of k is the minimum value, the minimum value is taken as the target value of k.

[0032] Optionally, the step of determining the activation information of sample numbers that pose a fraud risk includes:

[0033] When the activation information includes multiple items, one item is selected from the multiple activation information items as the activation information of the sample number with fraud risk, according to the preset information priority.

[0034] The activation information includes at least two of the following: the latitude and longitude at the time of activation, the connected base station, and the Internet Protocol address used.

[0035] According to a second aspect of the present disclosure, a fraud risk number marking device is provided, the device comprising:

[0036] The information determination module is configured to determine the activation information of sample numbers that pose a risk of fraud.

[0037] The region determination module is configured to determine a target risk region based on the activation information of each sample number in the sample set, wherein the target risk region is determined by dividing the sample numbers in the sample set into regions and then determining the geographical location represented by the activation information of the sample numbers in each region.

[0038] The marking module is configured to mark a number as a target number with fraud risk if the geographical location of the activation information is within the target risk area when the number is activated.

[0039] Optionally, the region determination module includes:

[0040] The region segmentation submodule is configured to segment sample numbers that pose a risk of fraud based on the activation information corresponding to each sample number in the sample set.

[0041] The calculation submodule is configured to calculate, for any one of the regions, the average and variance of the geographical location represented by the activation information corresponding to the sample number within the region;

[0042] The determination submodule is configured to determine the target risk area from the region based on the k-sigma algorithm, according to the average value and variance corresponding to each region, and the geographical location represented by the activation information of the sample number in the region.

[0043] Optionally, the determining submodule is configured as follows:

[0044] Repeat the following steps:

[0045] Based on the value of k in the k-sigma algorithm, the average value and variance corresponding to the region, abnormal sample numbers whose geographical locations are outside the Laida confidence interval corresponding to the value are removed, and normal sample numbers corresponding to the region that are within the Laida confidence interval corresponding to the value are obtained.

[0046] Based on the number of offline sample numbers and the number of non-offline sample numbers in the normal sample numbers of the region, determine whether the value of k needs to be adjusted;

[0047] If it is necessary to adjust the value of k, adjust the value of k within a preset range of preset k values ​​according to a preset step size until it is determined that it is not necessary to adjust the value of k, and obtain the target value of k.

[0048] Based on the geographical location represented by the activation information of the normal sample number corresponding to the target value of k, the target risk area is determined from the region.

[0049] Optionally, the determining submodule is configured as follows:

[0050] The minimum circle covering algorithm is used to determine the minimum circle that covers the geographical location of the activation information of the normal sample number corresponding to the target value of k.

[0051] The area covered by the smallest circle is taken as the target risk area corresponding to that area.

[0052] Optionally, the determining submodule is configured as follows:

[0053] The churn rate of the region is calculated based on the number of churned sample numbers and the number of non-churned sample numbers in the normal sample numbers of the region.

[0054] Based on the number of non-offline sample numbers and the churn rate, determine whether the value of k needs to be adjusted.

[0055] Optionally, the determining submodule is configured as follows:

[0056] Based on the relationship between the number of non-offline sample numbers and the preset number threshold, and the relationship between the churn rate and the preset churn rate threshold, it is determined whether the value of k needs to be adjusted.

[0057] Optionally, the determining submodule is configured as follows:

[0058] The maximum value of the preset range of k is taken as the initial value of k;

[0059] If it is necessary to adjust the value of k, the value of k is reduced within a preset range by a preset step size until it is determined that no adjustment of the value of k is needed, thus obtaining the target value of k.

[0060] Optionally, the determining submodule is configured as follows:

[0061] If the value of k is decreased within a preset range by a preset step size until the value of k is the minimum value, the minimum value is taken as the target value of k.

[0062] Optionally, the information determination module is configured to:

[0063] When the activation information includes multiple items, one item is selected from the multiple activation information items as the activation information of the sample number with fraud risk, according to the preset information priority.

[0064] The activation information includes at least two of the following: the latitude and longitude at the time of activation, the connected base station, and the Internet Protocol address used.

[0065] According to a third aspect of the present disclosure, an electronic device is provided, comprising:

[0066] processor;

[0067] Memory used to store processor-executable instructions;

[0068] The processor is configured as follows:

[0069] Identify activation information for sample phone numbers that pose a risk of fraud;

[0070] Based on the activation information of each sample number in the sample set, a target risk area is determined, wherein the target risk area is determined by dividing the sample numbers in the sample set into regions and then using the geographical location represented by the activation information of the sample numbers in each region.

[0071] When a number is activated, if the geographical location of the activation information is within the target risk area, the number will be marked as a target number with a risk of fraud.

[0072] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided that stores computer program instructions thereon, which, when executed by a processor, implement the steps of the method described in any one of the first aspects.

[0073] The technical solutions provided by the embodiments of this disclosure may include the following beneficial effects:

[0074] The system can identify target risk areas based on the activation information of sample numbers that pose a risk of fraud. Numbers whose latitude and longitude are within the target risk area at the time of activation can then be marked as potentially fraudulent. This marking of risky numbers can be done at the time of activation, i.e., before the fraud occurs, rather than during or after the fraud, thus improving the comprehensiveness and timeliness of fraud number marking.

[0075] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0076] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0077] Figure 1 This is a flowchart illustrating a method for marking fraud risk numbers according to an exemplary embodiment.

[0078] Figure 2 This is an implementation illustrated according to an exemplary embodiment. Figure 1 A flowchart of the method in step S12.

[0079] Figure 3 This is an implementation illustrated according to an exemplary embodiment. Figure 2 A flowchart of the method in step S123.

[0080] Figure 4 This is a schematic diagram illustrating another method for marking fraud risk numbers according to an exemplary embodiment.

[0081] Figure 5 This is a block diagram illustrating a fraud risk number marking device according to an exemplary embodiment.

[0082] Figure 6 This is a block diagram illustrating an apparatus for marking fraud risk numbers according to an exemplary embodiment. Detailed Implementation

[0083] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0084] It should be noted that all actions involving the acquisition of signals, information, or data in this application are carried out in compliance with the relevant data protection laws and policies of the country where the application is located, and with the authorization granted by the owner of the relevant device.

[0085] Figure 1 This is a flowchart illustrating a method for marking fraudulently risky phone numbers according to an exemplary embodiment. This method can be applied to a carrier's number activation device, such as a cloud server for number activation. See also... Figure 1 As shown, the method includes the following steps.

[0086] In step S11, the activation information of sample numbers that pose a risk of fraud is determined;

[0087] In this embodiment of the disclosure, the sample numbers that pose a risk of fraud can be determined from historical fraud records based on the victim's number's gender structure, age structure, network access duration, and information such as call records, text messages, instant voice messages, and transaction records.

[0088] In one implementation, a communication network can be generated from the full call records, and the spatiotemporal characteristics of sample numbers can be extracted. Sample numbers with fraud risk can be identified from the full number of numbers in the form of a spatiotemporal graph.

[0089] In this embodiment of the disclosure, the step of determining the activation information of sample numbers with fraud risk in step S11 includes:

[0090] When the activation information includes multiple items, one item is selected from the multiple activation information items as the activation information of the sample number with fraud risk, according to the preset information priority.

[0091] The activation information includes at least two of the following: the latitude and longitude at the time of activation, the connected base station, and the Internet Protocol address used.

[0092] It can be explained that the preset information priority is determined based on the accuracy of the location of the information representing the sample number when it is activated. For example, latitude and longitude have the highest accuracy in representing the location of the sample number when it is activated, so geographical location has the highest priority, while base station has the lowest accuracy in representing the location of the sample number when it is activated, so base station has the lowest priority.

[0093] For example, based on the priority that the latitude and longitude of the activated number is greater than the Internet Protocol address used during activation, which is greater than the base station connected during activation, the following steps are taken: First, obtain the latitude and longitude of the sample number at the time of activation if it is at risk of fraud. If the latitude and longitude are obtained, use them as activation information. Next, if the latitude and longitude are not obtained, obtain the Internet Protocol address used during activation. If the Internet Protocol address is obtained, use the location information represented by that address as activation information. Finally, if the Internet Protocol address is not obtained, obtain the base station connected to the sample number during activation and use its location information as activation information.

[0094] In step S12, the target risk area is determined based on the activation information of each sample number in the sample set.

[0095] The target risk area is determined by dividing the sample numbers in the sample set into regions and then using the geographical location represented by the activation information of the sample numbers in each region.

[0096] It is understood that the sample set in this disclosure embodiment is constructed using sample numbers that pose a risk of fraud.

[0097] In one implementation, see Figure 2 As shown, in step S12, the step of determining the target risk area based on the activation information corresponding to each sample number in the sample set includes:

[0098] In step S121, based on the activation information corresponding to each sample number in the sample set, sample numbers with fraud risk are divided into regions.

[0099] It is worth noting that the activation information corresponding to the sample numbers in the sample set varies. Specifically, there are first-type sample numbers whose activation information is the latitude and longitude at the time of activation; second-type sample numbers whose activation information is the Internet Protocol address used at the time of activation; and third-type sample numbers whose activation information is the base station they were connected to at the time of activation. Therefore, the different types of sample numbers within the same area are aggregated based on location to complete the area division.

[0100] In this embodiment of the disclosure, the geographical location represented by the activation information corresponding to the sample number can be distinguished layer by layer according to the city, district, and street where the geographical location is located, and finally the area can be distinguished according to the street where the geographical location represented by the activation information corresponding to the sample number is located.

[0101] In step S122, for any region, the mean and variance of the geographical location represented by the activation information corresponding to the sample number in that region are calculated;

[0102] In one implementation, for any given region, the number of sample numbers within that region is counted. If the number of sample numbers is greater than or equal to a preset threshold, the average and variance of the geographical locations represented by the activation information corresponding to those sample numbers within that region are calculated. If the number of sample numbers is less than the preset threshold, the region is determined to have a low risk, and the calculation for that region is abandoned.

[0103] In step S123, based on the k-sigma algorithm, the target risk area is determined from the region according to the mean and variance of each region and the geographical location represented by the activation information of the sample number in that region.

[0104] It can be explained that k here represents the value of the standard deviation multiple in the algorithm. For example, if the standard deviation multiple is 1.5, then the Raida confidence interval is constructed with the mean plus or minus 1.5 times the standard deviation.

[0105] Optionally, see Figure 3 As shown, in step S123, based on the k-sigma algorithm, the step of determining the target risk area from the region according to the mean and variance corresponding to each region, and the geographical location represented by the activation information of the sample numbers in that region, includes:

[0106] Repeat the following steps:

[0107] In step S1231, based on the value of k in the k-sigma algorithm, the mean and variance of the region, abnormal sample numbers whose geographical locations are outside the corresponding Laida confidence interval are removed, and normal sample numbers that are within the corresponding Laida confidence interval of the region are obtained.

[0108] It is understandable that the mean and labeled difference of the geographical location represented by the activation information of the sample numbers in this region are determined, and the corresponding Raida confidence interval is constructed with the mean as the center and plus or minus k times the standard deviation.

[0109] In step S1232, based on the number of off-network sample numbers and the number of non-off-network sample numbers in the normal sample numbers corresponding to the region, it is determined whether the value of k needs to be adjusted.

[0110] It is understandable that detached sample numbers can be numbers that were voluntarily cancelled by the operator, or numbers that were passively cancelled by the operator after never being used within a preset period. Non-detached sample numbers are numbers that have not been cancelled.

[0111] Optionally, step S1232, which determines whether the value of k needs to be adjusted based on the number of off-network sample numbers and the number of non-off-network sample numbers in the normal sample numbers of the region, includes:

[0112] The churn rate of a region is calculated based on the number of churned sample numbers and the number of non-churned sample numbers in the normal sample numbers of the region.

[0113] In this embodiment of the disclosure, the churn rate R can be calculated using the following formula. k :

[0114]

[0115] Among them, Withdraw k The number of off-network sample numbers, Active k This represents the number of non-offline sample numbers.

[0116] Based on the number of non-offline sample numbers and the churn rate, determine whether the value of k needs to be adjusted.

[0117] Optionally, the step of determining whether to adjust the value of k based on the number of non-offline sample numbers and the churn rate includes:

[0118] Based on the relationship between the number of non-offline sample numbers and the preset number threshold, and the relationship between the churn rate and the preset churn rate threshold, determine whether the value of k needs to be adjusted.

[0119] In this embodiment, if the number of non-offline sample numbers exceeds a preset threshold, or the churn rate is less than a preset churn rate threshold, then the value of k needs to be adjusted. If the number of non-offline sample numbers is less than or equal to the preset threshold, and the churn rate is greater than or equal to the preset churn rate threshold, then the value of k does not need to be adjusted. Specifically, a number of non-offline sample numbers less than or equal to the preset threshold indicates that the number of historically activated users in the area is very small, the workload for marking the area is minimal, the probability of mismarking is low, and shutting down number activation in the area will not significantly impact the operator's continued development. A churn rate greater than or equal to the preset churn rate threshold indicates that most of the historically activated users in the area have left the network, which meets the characteristics of clustered fraud.

[0120] It can be explained that after adjusting the value of k, the Raida confidence interval is re-determined based on the adjusted value of k. When there is no need to adjust the value of k, there are sample numbers within the Raida confidence interval corresponding to k.

[0121] In step S1233, if the value of k needs to be adjusted, the value of k is adjusted within a preset range of preset k values ​​according to a preset step size until it is determined that the value of k does not need to be adjusted, and the target value of k is obtained.

[0122] In this embodiment of the disclosure, the preset step size can be an empirical value, and the preset range of k values ​​is set based on the characteristics of clustering of historical fraud risk numbers.

[0123] Optionally, the method includes:

[0124] The maximum value within the preset range of k is used as the initial value of k.

[0125] In this embodiment of the disclosure, for example, the preset range of k is 1.1 to 1.9, that is, 1.9 is used as the initial value of k, and sample numbers whose geographical location is outside the range of plus or minus 1.9 times the standard deviation of the average value are identified as abnormal sample numbers and removed.

[0126] When it is necessary to adjust the value of k, the steps of adjusting the value of k within a preset range according to a preset step size until it is determined that no further adjustment of the value of k is needed, and obtaining the target value of k, include:

[0127] If the value of k needs to be adjusted, the value of k is decreased within a preset range by a preset step size until it is determined that the value of k does not need to be adjusted, thus obtaining the target value of k.

[0128] Following the previous embodiment, when k is 1.9, if the obtained churn rate and the number of non-churn sample numbers cannot meet the preset conditions, the preset conditions are that the number of non-churn sample numbers is less than or equal to a preset quantity threshold, or the churn rate is greater than or equal to a preset churn rate threshold. The value of k is reduced by a preset step size of 0.1, and sample numbers whose geographical location is outside the range of ±1.8 times the standard deviation of the average value are removed as abnormal sample numbers. The number of non-churn sample numbers and the churn rate of normal sample numbers within the corresponding Raida confidence interval are calculated until the obtained churn rate and the number of non-churn sample numbers meet the preset conditions.

[0129] Optionally, the method includes:

[0130] If the value of k is reduced within a preset range according to a preset step size until the value of k is a minimum, the minimum value is taken as the target value of k.

[0131] In this embodiment of the disclosure, if the value of k has been reduced to the minimum, there is no need to determine the corresponding Raida confidence interval, nor is it necessary to calculate the corresponding churn rate and the number of corresponding non-churn sample numbers. Instead, the value of k is directly set to the minimum value, and then the target risk area is determined from the region based on the geographical location represented by the activation information of the sample numbers within the Raida confidence interval.

[0132] In step S1234, the target risk area is determined from the region based on the geographical location represented by the activation information of the normal sample number corresponding to the target value of k.

[0133] In this embodiment of the disclosure, sample numbers can be divided into regions by streets. The distance between each pair of sample numbers is calculated based on the geographical location represented by the activation information of the normal sample number corresponding to the target value of k. The two sample numbers with the greatest distance are taken as target sample numbers. A square is constructed with the value of the greatest distance as the side length, and the streets covered by the square are taken as target risk areas.

[0134] Optionally, the step of determining the target risk area from the region based on the geographical location represented by the activation information of the normal sample number corresponding to the target value of k includes:

[0135] The minimum circle covering algorithm is used to determine the minimum circle representing the geographical location of the activation information of the normal sample number corresponding to the target value of covering k.

[0136] The area covered by the smallest circle is taken as the target risk area corresponding to that area.

[0137] In step S13, when a number is activated, if the geographical location of the activation information is within the target risk area, the number is marked as a target number with a risk of fraud.

[0138] See Figure 4 As shown, the number in this embodiment can be, for example, an unactivated number bundled with a mobile phone purchase, an unactivated number given as a gift for participating in an event, an unactivated number obtained through an online live stream, or an unactivated number obtained through an online store. Then, when a user activates the unactivated number, for example through a MVNO's application or by inserting a MVNO's SIM card, at least one of the following is extracted: the latitude and longitude of the number at the time of activation, the connected base station, and the Internet Protocol address used. This yields activation information representing the geographical location of the unactivated number at the time of activation.

[0139] In one implementation, when the target risk area is determined using the minimum circle coverage algorithm, the latitude and longitude of the center of the minimum circle can be determined. Then, based on the distance between the geographical location of the activation information representation of the number to be activated and the latitude and longitude of the circle's center, it can be determined whether the number is located within any target risk area. For example, if the distance between the geographical location of the activation information representation of the number to be activated and the latitude and longitude of the circle's center is less than the radius of the target risk area, it is determined that the number is within the target risk area.

[0140] In this embodiment, a target risk area can be obtained for each region. Finally, for a set of target risk areas for a region, a city, or a province, when any number is activated, its activation information is extracted. This includes at least one of the following: the latitude and longitude of the number at activation, the Internet Protocol address used at activation, and the base station it was connected to. If the geographical location represented by the activation information of the number falls within any target risk area in the set of target risk areas, the activated number can be marked or not activated. For example, numbers activated by overseas base stations are not activated, numbers located in domestic target risk areas are not activated, and numbers whose latitude and longitude at activation do not match the Internet Protocol address used at activation are not activated.

[0141] The above technical solution can determine the target risk area based on the activation information of sample numbers with potential fraud risk, and then mark the numbers whose latitude and longitude are in the target risk area at the time of activation as numbers with potential fraud risk. The risk numbers can be marked at the time of activation, that is, before the fraud occurs, rather than during or after the fraud occurs, which improves the comprehensiveness and timeliness of fraud number marking.

[0142] According to embodiments of this disclosure, a fraud risk number marking device 500 is also provided, see [link to relevant documentation]. Figure 5 As shown, the device 500 includes: an information determination module 510, a region determination module 520, and a marking module 530.

[0143] The information determination module 510 is configured to determine the activation information of sample numbers that pose a risk of fraud.

[0144] The region determination module 520 is configured to determine a target risk region based on the activation information of each of the sample numbers in the sample set, wherein the target risk region is determined by dividing the sample numbers in the sample set into regions and then determining the geographical location represented by the activation information of the sample numbers in each region.

[0145] The marking module 530 is configured to mark a number as a target number with fraud risk if the geographical location of the activation information is within the target risk area when the number is activated.

[0146] The aforementioned device can determine the target risk area based on the activation information of sample numbers that pose a risk of fraud. It then marks numbers whose latitude and longitude are within the target risk area at the time of activation as numbers that pose a risk of fraud. The risk numbers can be marked at the time of activation, that is, before the fraud occurs, rather than during or after the fraud occurs, thus improving the comprehensiveness and timeliness of fraud number marking.

[0147] Optionally, the region determination module 520 includes:

[0148] The region segmentation submodule is configured to segment sample numbers that pose a risk of fraud based on the activation information corresponding to each sample number in the sample set.

[0149] The calculation submodule is configured to calculate, for any one of the regions, the average and variance of the geographical location represented by the activation information corresponding to the sample number within the region;

[0150] The determination submodule is configured to determine the target risk area from the region based on the k-sigma algorithm, according to the average value and variance corresponding to each region, and the geographical location represented by the activation information of the sample number in the region.

[0151] Optionally, the determining submodule is configured as follows:

[0152] Repeat the following steps:

[0153] Based on the value of k in the k-sigma algorithm, the average value and variance corresponding to the region, abnormal sample numbers whose geographical locations are outside the Laida confidence interval corresponding to the value are removed, and normal sample numbers corresponding to the region that are within the Laida confidence interval corresponding to the value are obtained.

[0154] Based on the number of offline sample numbers and the number of non-offline sample numbers in the normal sample numbers corresponding to the region, determine whether the value of k needs to be adjusted;

[0155] If it is necessary to adjust the value of k, adjust the value of k within a preset range of preset k values ​​according to a preset step size until it is determined that it is not necessary to adjust the value of k, and obtain the target value of k.

[0156] Based on the geographical location represented by the activation information of the normal sample number corresponding to the target value of k, the target risk area is determined from this region.

[0157] Optionally, the determining submodule is configured as follows:

[0158] The minimum circle covering algorithm is used to determine the minimum circle that covers the geographical location of the activation information of the normal sample number corresponding to the target value of k.

[0159] The area covered by the smallest circle is taken as the target risk area corresponding to that area.

[0160] Optionally, the determining submodule is configured as follows:

[0161] The churn rate of the region is calculated based on the number of churned sample numbers and the number of non-churned sample numbers in the normal sample numbers of the region.

[0162] Based on the number of non-offline sample numbers and the churn rate, determine whether the value of k needs to be adjusted.

[0163] Optionally, the determining submodule is configured as follows:

[0164] Based on the relationship between the number of non-offline sample numbers and the preset number threshold, and the relationship between the churn rate and the preset churn rate threshold, it is determined whether the value of k needs to be adjusted.

[0165] Optionally, the determining submodule is configured as follows:

[0166] The maximum value of the preset range of k is taken as the initial value of k;

[0167] If it is necessary to adjust the value of k, the value of k is reduced within a preset range by a preset step size until it is determined that no adjustment of the value of k is needed, thus obtaining the target value of k.

[0168] Optionally, the determining submodule is configured as follows:

[0169] If the value of k is decreased within a preset range by a preset step size until the value of k is the minimum value, the minimum value is taken as the target value of k.

[0170] Optionally, the information determination module 510 is configured to:

[0171] When the activation information includes multiple items, one item is selected from the multiple activation information items as the activation information of the sample number with fraud risk, according to the preset information priority.

[0172] The activation information includes at least two of the following: the latitude and longitude at the time of activation, the connected base station, and the Internet Protocol address used.

[0173] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0174] Those skilled in the art should understand that the device embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and other division methods may exist in actual implementation. For instance, multiple modules may be combined or integrated into one module. Furthermore, the modules described as separate components may or may not be physically separate. For example, the region determination module 520 and the marking module 530 may be the same module or different modules physically, and each module may be implemented wholly or partially through software, hardware, firmware, or any combination thereof. When implemented in software, it may be implemented wholly or partially in the form of a computer program product. When implemented in hardware, it may be implemented wholly or partially in the form of an integrated circuit or chip.

[0175] This disclosure also provides an electronic device, including:

[0176] processor;

[0177] Memory used to store processor-executable instructions;

[0178] The processor is configured as follows:

[0179] Identify activation information for sample phone numbers that pose a risk of fraud;

[0180] Based on the activation information of each sample number in the sample set, a target risk area is determined, wherein the target risk area is determined by dividing the sample numbers in the sample set into regions and then using the geographical location represented by the activation information of the sample numbers in each region.

[0181] When a number is activated, if the geographical location of the activation information is within the target risk area, the number will be marked as a target number with a risk of fraud.

[0182] It can be noted that the electronic device in the embodiments of this disclosure can execute executable instructions stored in the memory to implement the steps of the fraud risk number marking method described in any of the foregoing embodiments of this disclosure.

[0183] This disclosure also provides a computer-readable storage medium storing computer program instructions that, when executed by a processor, implement the steps of any of the methods described in the foregoing embodiments.

[0184] Figure 6 This is a block diagram illustrating an apparatus 1900 for marking fraudulently risky phone numbers according to an exemplary embodiment. For example, apparatus 1900 may be provided as a cloud server for number activation. (Refer to...) Figure 6 The device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by memory 1932 for storing instructions, such as application programs, that can be executed by the processing component 1922. The application programs stored in memory 1932 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 1922 is configured to execute instructions to perform the aforementioned fraud risk number marking method.

[0185] Device 1900 may also include a power supply component 1926 configured to perform power management of device 1900, a wired or wireless network interface 1950 configured to connect device 1900 to a network, and an input / output interface 1958. Device 1900 can operate on an operating system stored in memory 1932, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, or similar.

[0186] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, which can be executed by a processor 820 of the device 800 to complete the aforementioned fraud risk number marking method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0187] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of this disclosure. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0188] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A fraud risk number marking method, characterized by, include: Identify activation information for sample phone numbers that pose a risk of fraud; Based on the activation information of each sample number in the sample set, a target risk area is determined, wherein the target risk area is determined by dividing the sample numbers in the sample set into regions and then using the geographical location represented by the activation information of the sample numbers in each region. When a number is activated, if the geographical location of the activation information is within the target risk area, the number will be marked as a target number with a risk of fraud. The step of determining the target risk area based on the activation information of each sample number in the sample set includes: Based on the activation information corresponding to each sample number in the sample set, sample numbers with fraud risk are divided into regions; For any of the regions, calculate the mean and variance of the geographical locations represented by the activation information corresponding to the sample numbers within the region; Based on the k-sigma algorithm, the target risk area is determined from the region according to the mean and variance corresponding to each region, and the geographical location represented by the activation information of the sample number in the region; The step of determining the target risk area from the region based on the k-sigma algorithm, according to the average value and variance corresponding to each region, and the geographical location represented by the activation information of the sample number within the region, includes: Repeat the following steps: Based on the value of k in the k-sigma algorithm, the average value and variance corresponding to the region, abnormal sample numbers whose geographical locations are outside the Laida confidence interval corresponding to the value are removed, and normal sample numbers corresponding to the region that are within the Laida confidence interval corresponding to the value are obtained. Based on the number of offline sample numbers and the number of non-offline sample numbers in the normal sample numbers corresponding to the region, determine whether the value of k needs to be adjusted; If it is necessary to adjust the value of k, adjust the value of k within a preset range of preset k values ​​according to a preset step size until it is determined that it is not necessary to adjust the value of k, and obtain the target value of k. Based on the geographical location represented by the activation information of the normal sample number corresponding to the target value of k, the target risk area is determined from the region.

2. The method of claim 1, wherein, The step of determining the target risk area from the region based on the geographical location represented by the activation information of the normal sample number corresponding to the target value of k includes: The minimum circle covering algorithm is used to determine the minimum circle that covers the geographical location of the activation information of the normal sample number corresponding to the target value of k. The area covered by the smallest circle is taken as the target risk area corresponding to the area.

3. The method of claim 1, wherein, The step of determining whether the value of k needs to be adjusted based on the number of offline sample numbers and the number of non-offline sample numbers in the normal sample numbers of the region includes: The churn rate of the region is calculated based on the number of churned sample numbers and the number of non-churned sample numbers in the normal sample numbers of the region. Based on the number of non-offline sample numbers and the churn rate, determine whether the value of k needs to be adjusted.

4. The method of claim 3, wherein, The step of determining whether the value of k needs to be adjusted based on the number of non-offline sample numbers and the churn rate includes: Based on the relationship between the number of non-offline sample numbers and the preset number threshold, and the relationship between the churn rate and the preset churn rate threshold, it is determined whether the value of k needs to be adjusted.

5. The method of claim 1, wherein, The method includes: The maximum value within the preset range of k is taken as the initial value of k; The step of adjusting the value of k within a preset range by a preset step size when it is necessary to adjust the value of k until it is determined that no adjustment of the value of k is needed, and obtaining the target value of k, includes: If it is necessary to adjust the value of k, the value of k is reduced within a preset range by a preset step size until it is determined that no adjustment of the value of k is needed, thus obtaining the target value of k.

6. The method of claim 5, wherein, The method includes: If the value of k is reduced within a preset range by a preset step size until the value of k is a minimum, then the minimum value is taken as the target value of k.

7. The method according to any one of claims 1 to 6, characterized in that, The step of determining the activation information of sample numbers that pose a risk of fraud includes: When there are multiple activation messages, one is selected from the multiple activation messages as the activation message of the sample number with fraud risk, according to a preset information priority. The activation information includes at least two of the following: the latitude and longitude at the time of activation, the connected base station, and the Internet Protocol address used.

8. A fraud risk number marking device, characterized by, The device includes: The information determination module is configured to determine the activation information of sample numbers that pose a risk of fraud. The region determination module is configured to determine a target risk region based on the activation information of each sample number in the sample set, wherein the target risk region is determined by dividing the sample numbers in the sample set into regions and then determining the geographical location represented by the activation information of the sample numbers in each region. The marking module is configured to mark the number as a target number with fraud risk if the geographical location of the activation information is within the target risk area when the number is activated. The region determination module includes: The region segmentation submodule is configured to segment sample numbers that pose a risk of fraud based on the activation information corresponding to each sample number in the sample set. The calculation submodule is configured to calculate, for any one of the regions, the average and variance of the geographical location represented by the activation information corresponding to the sample number within the region; The determination submodule is configured to determine the target risk area from the region based on the k-sigma algorithm, according to the average value and variance corresponding to each region, and the geographical location represented by the activation information of the sample number in the region; The determining submodule is configured to repeatedly execute the following steps: Based on the value of k in the k-sigma algorithm, the average value and variance corresponding to the region, abnormal sample numbers whose geographical locations are outside the Laida confidence interval corresponding to the value are removed, and normal sample numbers corresponding to the region that are within the Laida confidence interval corresponding to the value are obtained. Based on the number of offline sample numbers and the number of non-offline sample numbers in the normal sample numbers corresponding to the region, determine whether the value of k needs to be adjusted; If it is necessary to adjust the value of k, adjust the value of k within a preset range of preset k values ​​according to a preset step size until it is determined that it is not necessary to adjust the value of k, and obtain the target value of k. Based on the geographical location represented by the activation information of the normal sample number corresponding to the target value of k, the target risk area is determined from this region.

9. An electronic device, comprising: include: processor; Memory used to store processor-executable instructions; The processor is configured as follows: Identify activation information for sample phone numbers that pose a risk of fraud; Based on the activation information of each sample number in the sample set, a target risk area is determined, wherein the target risk area is determined by dividing the sample numbers in the sample set into regions and then using the geographical location represented by the activation information of the sample numbers in each region. When a number is activated, if the geographical location of the activation information is within the target risk area, the number will be marked as a target number with a risk of fraud. The step of determining the target risk area based on the activation information of each sample number in the sample set includes: Based on the activation information corresponding to each sample number in the sample set, sample numbers with fraud risk are divided into regions; For any of the regions, calculate the mean and variance of the geographical location represented by the activation information corresponding to the sample number within the region; Based on the k-sigma algorithm, the target risk area is determined from the region according to the mean and variance corresponding to each region, and the geographical location represented by the activation information of the sample number in the region; The step of determining the target risk area from the region based on the k-sigma algorithm, according to the average value and variance corresponding to each region, and the geographical location represented by the activation information of the sample number within the region, includes: Repeat the following steps: Based on the value of k in the k-sigma algorithm, the average value and variance corresponding to the region, abnormal sample numbers whose geographical locations are outside the Laida confidence interval corresponding to the value are removed, and normal sample numbers corresponding to the region that are within the Laida confidence interval corresponding to the value are obtained. Based on the number of offline sample numbers and the number of non-offline sample numbers in the normal sample numbers corresponding to the region, determine whether the value of k needs to be adjusted; If it is necessary to adjust the value of k, adjust the value of k within a preset range of preset k values ​​according to a preset step size until it is determined that it is not necessary to adjust the value of k, and obtain the target value of k. According to the geographical position represented by the activation information of the normal sample number corresponding to the target value of k, a target risk area is determined from the area.

10. A computer-readable storage medium having stored thereon computer program instructions, wherein, The program instructions, when executed by a processor, implement the steps of the method of any one of claims 1-7.