Risk identification method and device for call number, electronic equipment and computer program product

Through weighted scoring and classification models, the inefficiency problem in telecom fraud detection is solved, automated risk assessment and accurate number identification are realized, and the efficiency and adaptability of fraud detection are improved.

CN120372480APending Publication Date: 2025-07-25CHINA TELECOM CORP LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510437120.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing technology is inefficient in telecom fraud detection, making it difficult to deal with massive call data and changing fraudulent means, and cannot quickly screen and locate risk numbers.

Method used

The weighted scoring model and preset classification model are used to analyze the call data, and risk numbers are screened through weighted sum scores, and qualitative analysis is carried out in combination with call and network behavior data to achieve automated risk assessment and identification.

Benefits of technology

It improves the efficiency of number risk identification, can accurately identify fraud-related calls, reduce false alarms and missed reports, timely block fraudulent behavior, and ensure user safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372480A_ABST
    Figure CN120372480A_ABST
Patent Text Reader

Abstract

The invention discloses a call number risk identification method and device, electronic equipment and a computer program product. Obtaining a plurality of call tickets to be analyzed; analyzing the call data corresponding to the call number by using a weighted scoring model to obtain a risk score of the call number, the weighted scoring model performing weighted summation on parameter scores corresponding to multiple call parameters of the same call number by using a feature weight corresponding to each call parameter in advance; screening call numbers of which the risk scores accord with a preset scoring condition from call numbers corresponding to the plurality of call tickets to be analyzed to obtain risk numbers to be analyzed; and performing risk analysis on the risk behavior data corresponding to the to-be-analyzed risk number by using a preset classification model to obtain a number category to which the to-be-analyzed risk number belongs, the number category at least comprising a risk category to which the risk number belongs. According to the invention, the technical problem of poor efficiency of traditional number risk identification is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of machine learning, and in particular, to a method, apparatus, electronic device, and computer program product for risk identification of call numbers. Background Art

[0002] In the telecommunications field, with the rapid development of technology, the problem of telecommunications fraud has become increasingly serious, bringing great economic losses and security risks to society. Traditional fraud detection methods mainly rely on manual review and user reports. Specifically:

[0003] Manual review: Usually, a dedicated team or institution reviews suspicious call records one by one, analyzing factors such as call duration, call objects, and call frequency to determine whether there is fraud. This method is time-consuming and laborious, with low efficiency, and it is difficult to handle a large amount of call data.

[0004] User reports: Depend on users to actively report suspicious fraud calls to telecommunications operators or relevant institutions. However, this method is often lagging, and users may fail to report in a timely manner due to lack of professional knowledge or insufficient vigilance.

[0005] In addition, some rule-based methods have also been used for fraud detection, such as identifying suspicious behaviors by setting fixed call patterns or features (such as frequently changing numbers, making a large number of calls to different numbers in a short time, etc.). However, these methods are often too simple, easily bypassed by fraudsters, and difficult to adapt to constantly changing fraud means.

[0006] However, the existing technical solutions have the following disadvantages:

[0007] Backward technical means: Manual review and rule-based methods are too simple and backward to handle complex and changeable fraud behaviors.

[0008] Insufficient data utilization: A large amount of call data is not fully utilized, lacking effective data analysis and mining means to extract useful information.

[0009] Lack of intelligence: Existing fraud detection means lack intelligent technical means such as machine learning and cannot achieve automated risk assessment and call classification.

[0010] Therefore, the existing technical solutions have the following main problems in telecommunications fraud detection:

[0011] Low efficiency: The method of manual review is time-consuming and laborious, and it is difficult to handle the increasing amount of call data.

[0012] Limited coverage: User reports and rule-based methods are often lagging and limited, and it is difficult to comprehensively cover all fraud behaviors.

[0013] Poor adaptability: Existing fraud detection methods often struggle to adapt to the ever-changing fraud methods and user call behaviors.

[0014] Among them, the most prominent technical problems are low efficiency and poor adaptability. Traditional fraud detection methods cannot achieve rapid screening and positioning of a large number of numbers, nor can they promptly adapt to the emergence of new fraud methods, resulting in the difficulty of timely detection and blocking of fraud behaviors.

[0015] Regarding the problem of poor efficiency in the above-mentioned traditional risk number identification, no effective solution has been proposed yet. Summary of the Invention

[0016] Embodiments of the present invention provide a method, device, electronic device, and computer program product for risk identification of call numbers, so as to at least solve the technical problem of poor efficiency in traditional number risk identification.

[0017] According to one aspect of an embodiment of the present invention, a method for risk identification of call numbers is provided, including: obtaining a plurality of call records to be analyzed, where each of the call records to be analyzed at least includes: call data corresponding to a plurality of call numbers, and the call data includes multiple call parameters; using a weighted scoring model to analyze the call data corresponding to the call numbers to obtain a risk score for the call numbers, where the weighted scoring model uses the pre-set feature weights corresponding to each call parameter to perform weighted summation on the parameter scores corresponding to multiple call parameters of the same call number; screening the call numbers whose risk scores meet the preset scoring conditions from the call numbers corresponding to the plurality of call records to be analyzed to obtain call numbers to be analyzed for risk; using a preset classification model to perform risk analysis on the risk behavior data corresponding to the call numbers to be analyzed for risk to obtain the number categories to which the call numbers to be analyzed for risk belong, where the risk behavior data at least includes: the call data, and the number categories at least include the risk categories to which the risk numbers belong.

[0018] Optionally, obtaining multiple call bills to be analyzed includes: receiving multiple pieces of original call bill data provided by a call record database, where each piece of the original call bill data records at least: the calling number, the called number, the call duration, and the number attribution of the called number; performing data preprocessing on the multiple pieces of original call bill data to obtain the multiple call bills to be analyzed, where the call bills to be analyzed include the call data corresponding to multiple call numbers determined based on the multiple pieces of original call bill data, and multiple call parameters in the call data corresponding to each call number at least include: the number of calls with the call number as the calling number, the call duration, and the calling proportion, and the number dispersion of multiple called numbers called by the call number, and the number dispersion is determined based on the number attribution of the multiple called numbers called by the call number.

[0019] Optionally, using a weighted scoring model to analyze the call data corresponding to each call number to obtain the risk score of each call number includes: performing data analysis on the call data corresponding to each call number to obtain the parameter score corresponding to each call parameter in the call data; using the weighted scoring model to perform weighted summation on the parameter scores corresponding to multiple call parameters of the same call number by using the feature weights preset for each call parameter to obtain the risk score of the call number.

[0020] Optionally, before using the weighted scoring model to analyze the call data corresponding to each call number to obtain the risk score of each call number, the method further includes: using the weighted scoring model to analyze the call data corresponding to each historical number in multiple groups of historical call data to obtain the estimated score of each historical number, where each group of historical call data includes: the call data corresponding to the historical number, and the risk score already determined for the historical number; detecting the score difference between the estimated score and the already determined risk score of the same historical number; adjusting the feature weights in the weighted scoring model according to the score difference.

[0021] Optionally, screening the call numbers whose risk scores meet the preset scoring conditions as the call numbers to be analyzed for risk includes: matching the risk score of each call number with the scoring intervals corresponding to multiple preset levels to obtain the target level to which the call number belongs; in the case where the target level belongs to the pre-set high-risk level, determining the call number as the call number to be analyzed for risk.

[0022] Optionally, a preset classification model is used to perform risk analysis on the risk behavior data corresponding to the risk number to be analyzed, and obtaining the number category to which the risk number to be analyzed belongs includes: obtaining network behavior data corresponding to the risk number to be analyzed, wherein the network behavior data at least includes: Internet behavior data and application usage records; determining the network behavior data and the call data corresponding to the risk number to be analyzed as the risk behavior data, and using the preset classification model to perform risk analysis to obtain the number category to which the risk number to be analyzed belongs.

[0023] Optionally, before using a preset classification model to perform risk analysis on the risk behavior data corresponding to the risk number to be analyzed and obtaining the number category to which the risk number to be analyzed belongs, the method also includes: obtaining a confusion matrix, wherein the confusion matrix records multiple sample numbers and sample behavior data corresponding to each of the sample numbers; setting a category label for each of the sample numbers in the confusion matrix, wherein the category label is used to indicate the number category to which the sample number belongs, and the number categories include at least: fraudulent number, victim, diversion number, delivery person, and normal use; using the confusion matrix with the category label added as training data to train the preset classification model.

[0024] According to another aspect of an embodiment of the present invention, there is also provided a risk identification device for call numbers, comprising: an acquisition module, used to acquire a plurality of call records to be analyzed, wherein each of the call records to be analyzed includes at least: call data corresponding to a plurality of call numbers, the call data including a plurality of call parameters; a first analysis module, used to analyze the call data corresponding to the call number using a weighted scoring model to obtain a risk score for the call number, wherein the weighted scoring model uses a feature weight pre-set for each call parameter to perform a weighted summation of parameter scores corresponding to a plurality of call parameters of the same call number; a screening module, used to screen the call numbers whose risk scores meet preset scoring conditions from the call numbers corresponding to the plurality of call records to be analyzed, to obtain the risk numbers to be analyzed; a second analysis module, used to perform a risk analysis on the risk behavior data corresponding to the risk numbers to be analyzed using a preset classification model, to obtain the number category to which the risk numbers to be analyzed belong, wherein the risk behavior data includes at least: the call data, and the number category includes at least the risk category to which the risk numbers belong.

[0025] According to another aspect of an embodiment of the present invention, there is also provided an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the risk identification method for the call number through the computer program.

[0026] According to another aspect of the embodiments of the present invention, there is also provided a computer program product, including computer instructions, which implement the steps of the above-mentioned risk identification method for call numbers when executed by a processor.

[0027] In the embodiments of the present invention, a plurality of call bills to be analyzed are obtained, where each call bill to be analyzed at least includes: call data corresponding to a plurality of call numbers, and the call data includes multiple call parameters; a weighted scoring model is used to analyze the call data corresponding to the call numbers to obtain a risk score for the call numbers, where the weighted scoring model uses the feature weights pre-assigned to each call parameter to perform weighted summation on the parameter scores corresponding to multiple call parameters of the same call number; from the call numbers corresponding to the plurality of call bills to be analyzed, the call numbers whose risk scores meet the preset scoring conditions are screened to obtain the call numbers to be analyzed for risk; a preset classification model is used to perform risk analysis on the risk behavior data corresponding to the call numbers to be analyzed for risk to obtain the number categories to which the call numbers to be analyzed for risk belong, where the number categories at least include the risk categories to which the risk numbers belong. Thus, by performing risk scoring of calls through a weighted scoring model and further performing call qualitative analysis in combination with a preset classification model, the number risk can be evaluated more comprehensively, fraudulent calls can be accurately identified, and the purpose of automatically performing risk assessment and call qualitative analysis is achieved. Compared with the method of manually screening risk numbers, the screening efficiency can be improved, and the technical effect of improving the efficiency of identifying risk numbers is achieved, thereby solving the technical problem of poor efficiency in traditional number risk identification. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The drawings described herein are used to provide a further understanding of the present invention, and constitute a part of this application. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0029] Figure 1 is a flowchart of a method for identifying the risk of a call number according to an embodiment of the present invention;

[0030] Figure 2 is a schematic diagram of an architecture used in a method for analyzing call characteristics based on pattern recognition results according to an embodiment of the present invention;

[0031] Figure 3 is a schematic diagram of a fraud call detection result according to an embodiment of the present invention;

[0032] Figure 4 is a schematic diagram of using a confusion matrix to train a model according to an embodiment of the present invention;

[0033] Figure 5 is a schematic diagram of a classification boundary graph according to an embodiment of the present invention;

[0034] Figure 6 It is a schematic diagram of a risk identification device for call numbers according to an embodiment of the present invention;

[0035] Figure 7 It is a structural block diagram of a computer terminal according to an embodiment of the present invention. Detailed implementation manners

[0036] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0037] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order different from those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0038] First, some nouns or terms that appear in the process of describing the embodiments of the present application are applicable to the following explanations:

[0039] Call Feature Analysis: Deeply analyze the call behavior characteristics of users, including key information such as call times, call durations, call locations, call objects, etc., as well as derived features calculated therefrom, such as the ratio of outgoing calls, the ratio of incoming calls, and the number dispersion of numbers.

[0040] Pattern Recognition: Based on call feature analysis, use machine learning algorithms or preset rule sets to identify whether there are patterns or features related to fraud activities in the call. This usually involves the display of classification result graphs, including confusion matrices and classification boundary graphs, to evaluate the performance of the model.

[0041] Weighted Scoring Model: Based on historical experience and data analysis results, reasonable weights are set for each call feature, and a comprehensive score is calculated using the method of weighted summation to evaluate the fraud risk level of the number.

[0042] Risk Level Judgment: According to the comprehensive score calculated by the weighted scoring model, the numbers are divided into different risk levels (such as levels 1-5), and the higher the level, the higher the risk.

[0043] Call Qualitative Analysis: Conduct in-depth analysis on the detailed information of specific calls. Combining factors such as call duration, call location, and call object, and using the results of pattern recognition to determine whether it belongs to a fraud call.

[0044] Model Optimization and Iteration: Evaluate the performance of the model on real data through methods such as cross-validation and A / B testing, including indicators such as accuracy, recall rate, and F1 score. And adjust the model parameters (such as feature weights, thresholds, etc.) according to the evaluation results, and regularly collect new data for retraining to optimize the model performance and ensure that the model always maintains the best state.

[0045] Data Preprocessing: After obtaining the call detail record data of users from the database of telecom operators, perform data preprocessing steps such as data cleaning (removing duplicate, missing, or abnormal records) and formatting (ensuring consistent data types) for subsequent analysis.

[0046] Confusion Matrix: In pattern recognition, a table used to evaluate the performance of a classification model, where the rows represent the actual classes, the columns represent the predicted classes, and the values in the table represent the number of data points for the corresponding class combinations.

[0047] Classification Boundary Diagram: Draw a decision boundary in the feature space to show how the model classifies data points into different classes according to the features. For multi-classification problems, the classification boundary diagram may be more complex because the interactions between multiple classes need to be considered.

[0048] According to an embodiment of the present invention, an embodiment of a method for risk identification of call numbers is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0049] Figure 1 is a flowchart of a method for risk identification of call numbers according to an embodiment of the present invention. As Figure 1 shown, the method includes the following steps:

[0050] Step S102, obtain multiple call bills to be analyzed. Among them, each call bill to be analyzed at least includes: call data respectively corresponding to multiple call numbers, and the call data includes multiple call parameters;

[0051] Step S104, use a weighted scoring model to analyze the call data corresponding to the call number, and obtain a risk score of the call number. Among them, the weighted scoring model uses the feature weights pre-assigned to each call parameter to perform weighted summation on the parameter scores respectively corresponding to multiple call parameters of the same call number;

[0052] Step S106, screen the call numbers whose risk scores meet the preset scoring conditions from the call numbers respectively corresponding to multiple call bills to be analyzed, and obtain the call numbers to be analyzed for risk;

[0053] Step S108, use a preset classification model to perform risk analysis on the risk behavior data corresponding to the call numbers to be analyzed for risk, and obtain the number category to which the call numbers to be analyzed for risk belong. Among them, the risk behavior data at least includes: call data, and the number category at least includes the risk category to which the risk call number belongs.

[0054] In an embodiment of the present invention, multiple call records to be analyzed are obtained. Each call record to be analyzed at least includes: call data corresponding to multiple call numbers, where the call data includes multiple call parameters; a weighted scoring model is used to analyze the call data corresponding to a call number to obtain a risk score of the call number. The weighted scoring model weights and sums the parameter scores corresponding to multiple call parameters of the same call number by using the feature weights pre-assigned to each call parameter; call numbers whose risk scores meet the preset scoring conditions are screened from the call numbers corresponding to the multiple call records to be analyzed to obtain call numbers to be analyzed for risk; a preset classification model is used to perform risk analysis on the risk behavior data corresponding to the call numbers to be analyzed for risk to obtain the number category to which the call numbers to be analyzed for risk belong. The number category at least includes the risk category to which the risk number belongs. Therefore, by using the weighted scoring model to perform risk scoring of calls and combining the preset classification model to further perform call qualitative analysis, the number risk can be evaluated more comprehensively, fraud-related calls can be accurately identified, and the purpose of automatically performing risk assessment and call qualitative analysis is achieved. Compared with the method of manually screening risk numbers, the screening efficiency can be improved, and the technical effect of improving the efficiency of identifying risk numbers is realized, thus solving the technical problem of poor efficiency in traditional number risk identification.

[0055] In step S102 above, the call records to be analyzed can be pre-stored in the call record database. The call record data updated in the call record database within the preset time interval can be obtained at the preset time interval and used as the call record data to be analyzed for analysis; alternatively, after the call record data recorded in the call record database is updated, the updated call record data can be obtained in real time and used as the call record data to be analyzed for analysis.

[0056] In step S102 above, the call records to be analyzed can include multiple call records. Each call record can represent the call behavior of a set of calling number and called number. By integrating the multiple call records in the call records to be analyzed, the call behaviors of multiple call numbers can be obtained, and then the call behavior of each call number is represented by call data.

[0057] Optionally, the call numbers include: calling number and called number.

[0058] In step S102 above, the call parameters of the call number can reflect the behavioral characteristics of the call behavior of the call number. Among them, the multiple call parameters of the call number can include: call times, calling times, called times, call duration, call location, etc., as well as calling ratio, called ratio, and number dispersion, etc.

[0059] Optionally, each call parameter in the call data can be scored separately, and the scoring result corresponding to each call parameter is the parameter score of the call parameter.

[0060] In the above step S104, the weighted scoring model is used to integrate the scoring results of multiple call parameters. The weighted scoring model can preset the corresponding feature weights for each call parameter in advance, and then combine the parameter score corresponding to each call parameter with the feature weight corresponding to this call parameter to obtain the weighted score of this call parameter. Furthermore, by performing a summation operation on the weighted scores corresponding to multiple call parameters of the same call number, the weighted score of this call number can be obtained.

[0061] In the above step S106, the preset scoring condition can be set according to the risk scores of high-risk numbers, such as fraud numbers. Then, the risk numbers to be analyzed screened out using this preset scoring condition are high-risk numbers, such as fraud numbers.

[0062] In the above step S108, the preset classification model can be a multi-classification model, which is used to classify the risk number to be analyzed, and can realize the qualitative analysis of the risk number to be analyzed.

[0063] In the above step S108, the preset classification model is obtained through machine learning training using multiple sets of training data. The multiple sets of training data include: multiple sample numbers corresponding to multiple number categories respectively, and the sample behavior data corresponding to each sample number. The sample behavior data includes at least: call data.

[0064] It should be noted that in the risk behavior data analyzed by the preset classification model, in addition to call data, it can also include other behavior data of the risk number to be analyzed, such as the Internet access behavior data and application usage records of the risk number to be analyzed.

[0065] Optionally, the Internet access behavior data can be generated by the risk number to be analyzed during Internet access behavior.

[0066] Optionally, the application usage record can be generated when the terminal using the risk number to be analyzed is using the application APP; it can also be the registration behavior of the risk number to be analyzed in the application APP.

[0067] In the above step S108, the number categories can include: fraud numbers, victims, lead numbers, deliverymen, normal usage. The risk numbers can be fraud numbers, where the risk category to which the risk number belongs is the same as the number category to which the fraud number belongs.

[0068] As an alternative embodiment, obtaining a plurality of call bills to be analyzed includes: receiving a plurality of original bill data provided by a call record database, where each original bill data records at least: the calling number, the called number, the call duration, and the location of the called number; performing data preprocessing on the plurality of original bill data to obtain a plurality of call bills to be analyzed, where the call bills to be analyzed include call data corresponding to a plurality of call numbers determined based on the plurality of original bill data, and multiple call parameters in the call data corresponding to each call number at least include: the number of calls with the call number as the calling number, the call duration, and the calling proportion, and the number dispersion of the multiple called numbers called by the call number, and the number dispersion is determined based on the locations of the multiple called numbers called by the call number.

[0069] In the above embodiments of the present application, after obtaining the original bill data, it is possible to perform cleaning and preprocessing to extract key call features, such as the number of calls and the call duration, from the calling number, the called number, the call duration, and the location of the called number recorded in the original bill data, and calculate derivative features, such as the calling proportion and the number dispersion. Thus, by using optimized data preprocessing techniques, a large amount of call data can be quickly processed, improving the overall processing efficiency of the system. And by extracting call features and calculating derivative features, the expression ability of the features is enhanced, providing a more accurate data basis for subsequent risk assessment. An effective data cleaning process can remove redundant and invalid information, improving data quality and analysis accuracy.

[0070] It should be noted that for normal numbers, that is, there are both calling and called behaviors, the difference in the number of calls and called is not too large, and for normal number users, due to the geographical area limitation of their own activities, the locations of the called objects also have regularity and are not too scattered; however, for risk numbers, especially for fraud numbers, their typical characteristics are: more calls and fewer called, and there is no specific geographical pattern for the called numbers called by the fraud number. Therefore, by calculating the number of calls, the call duration, and the calling proportion of the call number, the risk behavior of the call number can be evaluated, and then the risk score of the call number can be realized.

[0071] Optionally, the calling proportion can be determined by counting the number of times the same call number is used as the calling number and the called number respectively.

[0072] As an alternative embodiment, a weighted scoring model is used to analyze the call data corresponding to each call number, and the risk score of each call number is obtained as follows: the call data corresponding to each call number is analyzed to obtain the parameter score corresponding to each call parameter in the call data; the weighted scoring model uses the feature weights preset for each call parameter to perform weighted summation on the parameter scores corresponding to multiple call parameters of the same call number to obtain the risk score of the call number.

[0073] In the above embodiment of the present application, the risk of each call number is initially evaluated by using the weighted scoring model to evaluate the risk of each call number. Furthermore, the risk score obtained by the weighted scoring model can be used to preliminarily screen the call number. Subsequently, a preset classification model can be used to perform qualitative analysis on the screened call numbers, thereby reducing the number of call numbers that need to be analyzed by the preset classification model and improving the efficiency of identifying risk numbers.

[0074] Optionally, the feature weights corresponding to each call parameter in the weighted scoring model can be determined in advance using the call data of each historical number in the historical call data.

[0075] Optionally, the feature weights corresponding to each call parameter in the weighted scoring model can also be updated after each identification of a risk number.

[0076] It should be noted that the principles of the initial setting and update of the feature weights are similar. The following takes the determination of the feature weights using the call data of each historical number in the historical call data as an example for explanation.

[0077] As an alternative embodiment, before using the weighted scoring model to analyze the call data corresponding to each call number and obtain the risk score of each call number, the method further includes: using the weighted scoring model to analyze the call data corresponding to each historical number in multiple groups of historical call data to obtain the estimated score of each historical number, where each group of historical call data includes: the call data corresponding to the historical number and the determined risk score of the historical number; detecting the score difference between the estimated score and the determined risk score of the same historical number; and adjusting the feature weights in the weighted scoring model according to the score difference.

[0078] In the above embodiments of the present application, when setting the feature weights corresponding to each call parameter for the weighted scoring model, the weighted scoring model can be used to analyze the call data corresponding to the historical numbers whose risk scores have been determined in advance, obtain the estimated scores of each historical number, and then adjust the feature weights in the weighted scoring model according to the scoring difference between the estimated score of the historical number and the determined risk score until the scoring difference between the estimated score obtained by the weighted scoring model after the feature weights are adjusted and the determined risk score of the same historical number meets the preset difference requirement. Thus, by adjusting the feature weights in the weighted scoring model in the above manner, the training and update of the weighted scoring model can be realized.

[0079] As an optional embodiment, screening the call numbers whose risk scores meet the preset scoring conditions as the risk numbers to be analyzed includes: matching the risk score of each call number with the scoring intervals corresponding to multiple preset levels to obtain the target level to which the call number belongs; when the target level belongs to the preset high-risk level, determining the call number as the risk number to be analyzed.

[0080] In the above embodiments of the present application, after determining the risk score of each call number, the call numbers can be classified, and multiple call numbers can be divided in each preset level, where the multiple preset levels include a high-risk level. Furthermore, the call numbers belonging to the high-risk level are the risk numbers to be analyzed, realizing the screening of the risk numbers to be analyzed.

[0081] Optionally, each preset level is respectively set with a corresponding scoring interval. When the risk score of the call number meets a certain scoring interval, it is determined that the call number belongs to the preset level corresponding to the scoring interval.

[0082] As an optional embodiment, using a preset classification model to perform risk analysis on the risk behavior data corresponding to the risk numbers to be analyzed to obtain the number category to which the risk numbers to be analyzed belong includes: obtaining the network behavior data corresponding to the risk numbers to be analyzed, where the network behavior data at least includes: Internet access behavior data and application usage records; determining the network behavior data and call data corresponding to the risk numbers to be analyzed as risk behavior data, and using the preset classification model to perform risk analysis to obtain the number category to which the risk numbers to be analyzed belong.

[0083] In the above embodiments of the present application, when using a preset classification model to perform risk analysis on the risk behavior data corresponding to the risk numbers to be analyzed, the risk behavior data analyzed by the preset classification model can include, in addition to the call data, the network behavior data corresponding to the risk numbers to be analyzed. Furthermore, by combining the network behavior data and call data corresponding to the risk numbers to be analyzed for analysis, the accuracy of the evaluation can be improved.

[0084] As an optional embodiment, before using a preset classification model to perform risk analysis on the risk behavior data corresponding to the risk number to be analyzed and obtaining the number category to which the risk number to be analyzed belongs, the method also includes: obtaining a confusion matrix, wherein the confusion matrix records multiple sample numbers and sample behavior data corresponding to each sample number; setting a category label for each sample number in the confusion matrix, wherein the category label is used to indicate the number category to which the sample number belongs, and the number categories include at least: fraudulent number, victim, diversion number, delivery person, and normal use; and using the confusion matrix with added category labels as training data to train the preset classification model.

[0085] In the above-mentioned embodiments of the present application, the preset classification model can be a multi-classification model trained based on a confusion matrix. The confusion matrix can record multiple sample numbers of determined number categories and add a category label for each sample number. Then, the preset classification model is trained based on the confusion matrix, so that the trained preset classification model can perform multi-classification of number belonging categories and realize qualitative analysis of the numbers.

[0086] It should be noted that the risk numbers to be analyzed may have behaviors that are not fraudulent numbers. For example, the lead numbers need to promote products, which requires frequent active calls, and the number locations of the called parties (that is, the called numbers) are also relatively scattered. For example, delivery personnel need to deliver express and take-out food, which also requires frequent active calls. Therefore, the risk numbers to be analyzed selected by the risk score obtained based on the weighted scoring model may cover the above scenarios. Through qualitative analysis by the preset classification model, the numbers in the above situations can be distinguished, thereby improving the accuracy of recognition.

[0087] It should be noted that communication behavior requires the integrity of both the calling number and the called number. Therefore, when the calling number is a fraudulent number, the called number is the victim of the fraudulent account. By identifying the victim, timely intervention can be made in the fraudulent behavior carried out by the fraudulent number, and fraudulent calls can be discovered and blocked in a timely manner to ensure user safety.

[0088] The present invention also provides a preferred embodiment, which provides a call feature analysis method based on pattern recognition results. It aims to achieve rapid screening and positioning of numbers through automated technical means, improve fraud detection efficiency, and adapt to the ever-changing fraud methods and user call behaviors, as traditional fraud detection methods are inefficient and difficult to cope with massive call data; and the existing technology is difficult to comprehensively cover all fraud behaviors, and has poor adaptability and is unable to promptly adapt to the emergence of new fraud methods.

[0089] The technical solution provided by this application is mainly applied to the fraud detection system of communication operators, aiming to improve the fraud detection efficiency and ensure user safety through automated number risk assessment and call characterization technology.

[0090] As an optional example, the network element devices used in this application include: a call record database, a data analysis server, a risk judgment and characterization module, and a user interface.

[0091] Optionally, the call record database is used to store the call detail data of users, where the call detail data includes key information such as call time, call duration, and call object.

[0092] Optionally, the data analysis server is responsible for executing algorithms such as data preprocessing, feature extraction, weighted scoring, and pattern recognition to deeply analyze the call data.

[0093] Optionally, the risk judgment and characterization module is used to judge the risk level of the number based on the output result of the data analysis server and conduct qualitative analysis on specific calls.

[0094] Optionally, the user interface is used to provide a visual interface for communication operators or relevant security agencies to display the risk assessment results and call characterization information.

[0095] Optionally, the call record database and the data analysis server are connected through a high-speed network to ensure efficient data transmission.

[0096] Optionally, the data analysis server and the risk judgment and characterization module communicate through an internal interface to transfer the analysis results.

[0097] Optionally, the risk judgment and characterization module interacts with communication operators or relevant security agencies through the user interface.

[0098] Figure 2 It is a schematic diagram of the architecture used in a call feature analysis method based on pattern recognition results according to an embodiment of the present invention, as Figure 2 shown. This architecture adopts a layered architecture design, including a data layer, an algorithm layer (i.e., the model foundation layer), an application layer (i.e., the output layer), etc., to ensure the scalability and maintainability of the system.

[0099] Optionally, the data layer is responsible for data storage and management; the algorithm layer (i.e., the model foundation layer) is responsible for executing various analysis algorithms; the application layer (i.e., the output layer) is responsible for providing a user interface and displaying the analysis results.

[0100] As an optional embodiment, the main process steps of the call feature analysis method based on pattern recognition results provided by this application are as follows:

[0101] Step S11, data preprocessing and analysis. Extract key information from the original call record data, such as the number of calls, the number of outgoing calls, the number of incoming calls, call duration, call location, etc., and calculate derived features, such as outgoing call ratio, incoming call ratio, number dispersion, etc.

[0102] Step S12, weighted scoring model. According to historical experience in the database, set reasonable weights and corresponding ratios for each parameter, and use the method of weighted summation to calculate the comprehensive score of the number, which is used to evaluate the risk level of the number.

[0103] Step S13, risk level judgment and call characterization: Divide the numbers into different risk levels according to the comprehensive score, and conduct qualitative analysis on specific calls to determine whether they are fraud-related calls.

[0104] Step S14, model optimization and iteration: Use machine learning algorithms to train and optimize the model, and continuously update the model parameters through iteration to improve the prediction accuracy and robustness.

[0105] Step S15, system implementation and deployment: Develop a complete number risk assessment and call characterization system and deploy it on the servers of telecom operators or relevant security agencies.

[0106] The call feature analysis method based on pattern recognition results provided by this application can be applied to call behavior analysis and number risk assessment. It involves extracting call record data from the database of communication operators, and through a series of technical means such as data analysis, feature extraction, weighted scoring, and pattern recognition, deeply analyzing the call behavior of users to evaluate the fraud risk level of numbers and conduct qualitative judgments on specific calls. Therefore, the technical solution provided by this application can also be specifically applied to telecom fraud detection.

[0107] As an optional embodiment, the application of the call feature analysis method based on pattern recognition results in telecom fraud detection specifically includes the following steps:

[0108] Step S21, data preprocessing.

[0109] The execution entity of the above step S21 is: a data analysis server; the trigger condition is: receiving new data in the call record database regularly or in real-time; the processing object is: the original call record data; the processing actions are: cleaning the data, extracting key information (such as the number of calls, call duration, etc.), calculating derived features (such as outgoing call ratio, number dispersion, etc.); the processing result is: obtaining a preprocessed call feature data set.

[0110] Through the above step S21 of this application, an accurate and complete data basis is provided for subsequent analysis.

[0111] In the above embodiments of the present application, through step S21, by cleaning and preprocessing the original call record data, key call features are extracted and derivative features are calculated; thus, through the optimized data preprocessing technology, a large amount of call data can be quickly processed, improving the overall processing efficiency of the system. By extracting and calculating derivative features, the expression ability of the features is enhanced, providing a more accurate data basis for subsequent risk assessment. An effective data cleaning process can remove redundant and invalid information, reduce data redundancy, and improve data quality and analysis accuracy.

[0112] Step S22, weighted scoring model.

[0113] The execution subject of the above step S22 is: the data analysis server; the triggering condition is: the preprocessed call feature data set is ready; the processing object is: the call feature data set; the processing action is: setting feature weights according to historical experience and calculating the comprehensive score of each number; the processing result is: obtaining the risk score list of the numbers.

[0114] Through step S22 of the present application, a quantitative basis is provided for risk level judgment.

[0115] Step S23, risk level judgment and call qualification.

[0116] The execution subject of the above step S23 is: the risk judgment and qualification module; the triggering condition is: the risk score list is generated; the processing object is: each number in the risk score list and its call records; the processing action is: classifying the numbers into different risk levels according to a preset threshold, and performing pattern recognition on specific calls of high-risk numbers to determine whether they belong to fraud-related calls; the processing result is: obtaining the risk level report of the numbers and the fraud-related call list.

[0117] Through step S23 of the present application, fraud-related calls can be discovered and blocked in a timely manner to ensure user safety.

[0118] In the above embodiments of the present application, through step S23, the risk level of the numbers is evaluated using a weighted scoring model; pattern recognition is performed on specific calls of high-risk numbers to achieve call qualification analysis; compared with the prior art that mostly relies on a single scoring model or rule judgment, the present application can more comprehensively evaluate the risk of numbers, improve accuracy, and accurately identify fraud-related calls by combining risk level evaluation and call qualification analysis; through pattern recognition technology, it can learn and adapt to the characteristics of new fraud means, enhancing adaptability and making the system have stronger adaptability and robustness; through a dual verification mechanism, the probability of false positives and false negatives is effectively reduced, reducing false positives and false negatives and improving the reliability of the system.

[0119] Step S24, model optimization and iteration.

[0120] The execution entity of the above step S24 is: the data analysis server; the triggering condition is: regularly or according to the model performance evaluation results; the processing objects are: the weighted scoring model and the pattern recognition algorithm; the processing actions are: collecting new data, training and validating the model, and adjusting the feature weights and algorithm parameters to optimize the model performance; the processing result is: obtaining the updated weighted scoring model and pattern recognition algorithm.

[0121] Through the above step S24 of the present application, it can be ensured that the model always remains in the best state and adapts to the constantly changing fraud means.

[0122] In the above embodiments of the present application, through the above step S24, by regularly collecting new data to train and validate the model, the feature weights and algorithm parameters can be adjusted according to the model performance evaluation results. Compared with the models in the prior art that often lack a dynamic optimization mechanism, the present application ensures that the model always remains in the best state through continuous iterative updates; in addition, as the fraud means continue to change, the model can quickly adapt to the new situation through the dynamic optimization technology, has strong adaptability to changes, can improve the prevention and control effect, and reduces the need for manual intervention through the automated model optimization process, thereby reducing the system maintenance cost.

[0123] In the above step S23, different pattern recognition algorithms (such as decision trees, random forests, etc.) can be used to improve the recognition accuracy.

[0124] Optionally, in the weighted scoring model, the feature weights and proportions can be adjusted according to the actual situation to adapt to the fraud characteristics and user call habits in different regions.

[0125] As an alternative embodiment, the above step S23 can be combined with the analysis of user behavior for call feature analysis.

[0126] Optionally, on the basis of the above embodiments, the user behavior analysis technology is introduced to further improve the accuracy of fraud detection. The specific steps include:

[0127] Step S31, collecting additional information such as the user's Internet behavior data and APP usage records.

[0128] Step S32, using machine learning algorithms to perform feature extraction and correlation analysis on this information.

[0129] Step S33, combining the user behavior features with the call features to construct a more comprehensive risk assessment model.

[0130] Step S34, judging the risk level of the number and performing call qualitative analysis according to the comprehensive evaluation result.

[0131] In the above embodiments of the present application, by introducing user behavior analysis technology, the call behavior and daily habits of users can be more comprehensively understood, improving the accuracy and pertinence of fraud detection; at the same time, this extended technical solution also enhances the adaptability and scalability of the system, providing more possibilities for future fraud prevention and control work.

[0132] In the above embodiments of the present application, through the above steps S31 to S34, by collecting and analyzing additional information such as users' Internet behavior data and APP usage records; and combining user behavior characteristics with call characteristics to construct a comprehensive risk assessment model; compared with the prior art that mainly focuses on call characteristics, the present application can more comprehensively understand users' habits and improve the accuracy of risk assessment by introducing user behavior analysis; in addition, the introduction of user behavior data makes the risk assessment more personalized, enabling the provision of differentiated prevention and control strategies for different user groups; by analyzing the changing trends of users' behaviors, it is possible to give early warnings before fraud behaviors occur, improving the timeliness of prevention and control.

[0133] Figure 3 It is a schematic diagram of the detection results of fraud-related calls according to an embodiment of the present invention, as Figure 3 shown, based on the call detail records of XXX number in June, suspicious calls are qualitatively determined, a group of calls with fraud-related risks are identified, and the risk scores of this group of calls are obtained. And the fraud number is determined as "19933351787" from this group of calls, with a risk level of 5, a suggestion of high-risk dual suspension is generated, and the victim number is determined as "18180059052", with a risk level of 5, and a suggestion of protective single suspension is generated.

[0134] Figure 4 It is a schematic diagram of training a model using a confusion matrix according to an embodiment of the present invention, as Figure 4 shown, the categories are set as "fraud number", "victim", "drainage number", "courier" and "normal use"; the category labels of the confusion matrix will be updated, and the situation of multi-classification will be considered when generating the classification boundary diagram. Since the classification boundary diagram may be more complex because it is necessary to show the decision boundaries between multiple categories.

[0135] Figure 5 It is a schematic diagram related to a classification boundary diagram according to an embodiment of the present invention, as Figure 5 shown, for the classification boundary diagram, the decision_function or predict_proba method in scikit-learn is used to obtain the distance of each point to the decision boundary or the probability of belonging to each category, and the decision boundary is drawn through these values. However, directly drawing the decision boundary of multi-classification may be relatively complex because it is necessary to consider the interactions between multiple categories.

[0136] In the above embodiments of the present application, through automated number risk assessment and call qualitative techniques, rapid screening and positioning of a large number of numbers can be achieved, reducing the workload of manual review and significantly improving the efficiency of fraud detection; by combining call feature analysis and pattern recognition techniques, fraudulent calls can be more accurately identified, enhancing the accuracy of fraud detection and reducing the occurrence of false positives and false negatives; fraudulent calls can be promptly detected and blocked, effectively reducing users' economic losses and security risks, enhancing users' trust in telecom operators, and ensuring user safety; effective fraud prevention and control means are provided for telecom operators, improving the service quality of operators, and enhancing user satisfaction and trust; the technical solution provided by the present application can also be applied to risk assessment and anomaly detection in other fields, such as financial anti-fraud and network security monitoring, with strong scalability and broad application prospects.

[0137] According to an embodiment of the present invention, there is also provided an embodiment of a risk identification device for call numbers. It should be noted that this risk identification device for call numbers can be used to execute the risk identification method for call numbers in the embodiments of the present invention, and the risk identification method for call numbers in the embodiments of the present invention can be executed in this risk identification device for call numbers.

[0138] Figure 6 It is a schematic diagram of a risk identification device for call numbers according to an embodiment of the present invention. As Figure 6 shown, the device may include: an acquisition module 62, configured to acquire a plurality of call bills to be analyzed, where each call bill to be analyzed at least includes: call data corresponding to a plurality of call numbers, and the call data includes a plurality of call parameters; a first analysis module 64, configured to analyze the call data corresponding to the call numbers using a weighted scoring model to obtain a risk score of the call numbers, where the weighted scoring model uses the pre-set feature weights corresponding to each call parameter to perform weighted summation on the parameter scores corresponding to the plurality of call parameters of the same call number; a screening module 66, configured to screen the call numbers corresponding to the plurality of call bills to be analyzed respectively to obtain call numbers whose risk scores meet the pre-set scoring conditions, so as to obtain the call numbers to be analyzed at risk; a second analysis module 68, configured to perform risk analysis on the risk behavior data corresponding to the call numbers to be analyzed at risk using a pre-set classification model to obtain the number categories to which the call numbers to be analyzed at risk belong, where the risk behavior data at least includes: call data, and the number categories at least include the risk categories to which the risk numbers belong.

[0139] It should be noted that the acquisition module 62 in this embodiment can be used to execute step S102 in the embodiment of the present application, the first analysis module 64 in this embodiment can be used to execute step S104 in the embodiment of the present application, the screening module 66 in this embodiment can be used to execute step S106 in the embodiment of the present application, and the second analysis module 68 in this embodiment can be used to execute step S108 in the embodiment of the present application. The examples and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the contents disclosed in the above embodiments.

[0140] In an embodiment of the present invention, a plurality of call records to be analyzed are obtained, wherein each call record to be analyzed includes at least: call data corresponding to a plurality of call numbers, wherein the call data includes a plurality of call parameters; a weighted scoring model is used to analyze the call data corresponding to the call number to obtain a parameter score of the call number, wherein the weighted scoring model uses a feature weight corresponding to each call parameter in advance to perform a weighted sum of the parameter scores corresponding to a plurality of call parameters of the same call number; from the call numbers corresponding to the plurality of call records to be analyzed, call numbers whose risk scores meet preset scoring conditions are screened to obtain risk numbers to be analyzed; a preset classification model is used to perform a risk analysis on the risk behavior data corresponding to the risk numbers to be analyzed to obtain a number category to which the risk numbers to be analyzed belong, wherein the number category includes at least a risk category to which the risk numbers belong, thereby performing a risk score for the call through the weighted scoring model, and further performing a qualitative analysis of the call in combination with the preset classification model, so that the number risk can be more comprehensively evaluated, fraud-related calls can be accurately identified, and the purpose of automatically performing risk evaluation and call qualitative analysis can be achieved. Compared with the method of manually screening risk numbers, the screening efficiency can be improved, and the technical effect of improving the efficiency of identifying risk numbers is achieved, thereby solving the technical problem of poor efficiency of traditional number risk identification.

[0141] As an optional embodiment, the acquisition module includes: a receiving unit, which is used to receive multiple original call record data provided by a call record database, wherein each original call record data at least records: the calling number, the called number, the call duration and the number location of the called number; a preprocessing unit, which is used to perform data preprocessing on the multiple original call record data to obtain multiple call records to be analyzed, wherein the call records to be analyzed include call data corresponding to multiple call numbers determined based on the multiple original call record data, and the multiple call parameters in the call data corresponding to each call number include at least: the number of calls when the call number is the calling number, the call duration and the caller ratio, and the number dispersion of the multiple called numbers called by the call number, and the number dispersion is determined according to the number locations of the multiple called numbers called by the call number.

[0142] As an alternative embodiment, the first analysis module includes: a first analysis unit configured to analyze the call data corresponding to each call number to obtain a parameter score corresponding to each call parameter in the call data; and a calculation unit configured to use a weighted scoring model and the feature weights preset for each call parameter to perform weighted summation on the parameter scores corresponding to multiple call parameters of the same call number to obtain a risk score of the call number.

[0143] As an alternative embodiment, the apparatus further includes: an analysis sub-module configured to, before using the weighted scoring model to analyze the call data corresponding to each call number to obtain the risk score of each call number, use the weighted scoring model to analyze the call data corresponding to each historical number in multiple groups of historical call data to obtain an estimated score of each historical number, where each group of historical call data includes: the call data corresponding to the historical number and the determined risk score of the historical number; a detection sub-module configured to detect the score difference between the estimated score and the determined risk score of the same historical number; and an adjustment sub-module configured to adjust the feature weights in the weighted scoring model according to the score difference.

[0144] As an alternative embodiment, the screening module includes: a matching unit configured to match the risk score of each call number with the score intervals corresponding to multiple preset levels to obtain the target level to which the call number belongs; and a determination unit configured to, when the target level belongs to a preset high-risk level, determine the call number as a risk number to be analyzed.

[0145] As an alternative embodiment, the second analysis module includes: an acquisition unit configured to acquire the network behavior data corresponding to the risk number to be analyzed, where the network behavior data at least includes: Internet access behavior data and application usage records; and a third analysis unit configured to determine the network behavior data and the call data corresponding to the risk number to be analyzed as risk behavior data and perform risk analysis using a preset classification model to obtain the number category to which the risk number to be analyzed belongs.

[0146] As an alternative embodiment, the apparatus further includes: an acquisition sub-module configured to acquire a confusion matrix before using the preset classification model to perform risk analysis on the risk behavior data corresponding to the risk number to be analyzed to obtain the number category to which the risk number to be analyzed belongs, where the confusion matrix records multiple sample numbers and the sample behavior data corresponding to each sample number; a setting sub-module configured to set a category label for each sample number in the confusion matrix, where the category label is used to indicate the number category to which the sample number belongs, and the number categories at least include: fraud number, victim, traffic diversion number, deliveryman, normal use; and a training sub-module configured to use the confusion matrix with added category labels as training data to train the preset classification model.

[0147] Embodiments of the present invention may provide an electronic device, which may be a computer terminal, and the computer terminal may be any one of a group of computer terminal devices. Optionally, in this embodiment, the above computer terminal may also be replaced with a terminal device such as a mobile terminal.

[0148] Optionally, in this embodiment, the above computer terminal may be located in at least one of a plurality of network devices in a computer network.

[0149] In this embodiment, the above computer terminal may execute the program code of the following steps in the risk identification method of call numbers: obtaining a plurality of call bills to be analyzed, where each call bill to be analyzed at least includes: call data respectively corresponding to a plurality of call numbers, and the call data includes a plurality of call parameters; using a weighted scoring model to analyze the call data corresponding to the call numbers to obtain a risk score of the call numbers, where the weighted scoring model uses the feature weights pre-assigned to each call parameter to perform weighted summation on the parameter scores respectively corresponding to a plurality of call parameters of the same call number; screening the call numbers whose risk scores meet the preset scoring conditions from the call numbers respectively corresponding to the plurality of call bills to be analyzed to obtain the call numbers to be analyzed for risk; using a preset classification model to perform risk analysis on the risk behavior data corresponding to the call numbers to be analyzed for risk to obtain the number category to which the call numbers to be analyzed for risk belong, where the risk behavior data at least includes: call data, and the number category at least includes the risk category to which the risk call numbers belong.

[0150] Figure 7 is a structural block diagram of a computer terminal according to an embodiment of the present invention, as Figure 7 shown, the computer terminal 70 may include: one or more (only one is shown in the figure) processors 72 and a memory 74.

[0151] Among them, the memory may be used to store software programs and modules, such as the program instructions / modules corresponding to the risk identification method and device of call numbers in the embodiments of the present invention. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the above-mentioned risk identification method of call numbers. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely disposed relative to the processor, and these remote memories may be connected to the terminal 70 through a network. Examples of the above network include but are not limited to the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof.

[0152] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: obtain multiple call records to be analyzed, wherein each call record to be analyzed includes at least: call data corresponding to multiple call numbers, and the call data includes multiple call parameters; use a weighted scoring model to analyze the call data corresponding to the call number to obtain a risk score for the call number, wherein the weighted scoring model uses the feature weights pre-assigned to each call parameter to perform weighted summation on the parameter scores corresponding to multiple call parameters of the same call number; from the call numbers corresponding to the multiple call records to be analyzed, screen the call numbers whose risk scores meet the preset scoring conditions to obtain the risk numbers to be analyzed; use a preset classification model to perform risk analysis on the risk behavior data corresponding to the risk numbers to be analyzed, and obtain the number category to which the risk number to be analyzed belongs, wherein the risk behavior data includes at least: call data, and the number category includes at least the risk category to which the risk number belongs.

[0153] Optionally, the processor may also execute the program code of the following steps: receiving multiple original call record data provided by a call record database, wherein each original call record data records at least: the calling number, the called number, the call duration and the number location of the called number; performing data preprocessing on the multiple original call record data to obtain multiple call records to be analyzed, wherein the call records to be analyzed include call data corresponding to multiple call numbers determined based on the multiple original call record data, and the multiple call parameters in the call data corresponding to each call number include at least: the number of calls when the call number is the calling number, the call duration and the caller ratio, and the number dispersion of the multiple called numbers called by the call number, and the number dispersion is determined based on the number locations of the multiple called numbers called by the call number.

[0154] Optionally, the processor may also execute the program code of the following steps: performing data analysis on the call data corresponding to each call number to obtain a parameter score corresponding to each call parameter in the call data; using a weighted scoring model and utilizing the feature weights pre-set for each call parameter to perform weighted summation on the parameter scores corresponding to multiple call parameters of the same call number to obtain a risk score for the call number.

[0155] Optionally, the processor may also execute the program code of the following steps: using a weighted scoring model to analyze the call data corresponding to each historical number in multiple groups of historical call data to obtain an estimated score for each historical number, wherein each group of historical call data includes: the call data corresponding to the historical number, and a determined risk score for the historical number; detecting the score difference between the estimated score and the determined risk score of the same historical number; and adjusting the feature weights in the weighted scoring model based on the score difference.

[0156] Optionally, the above-mentioned processor may also execute the program code of the following steps: matching the risk score of each call number with the score intervals corresponding to multiple preset levels to obtain the target level to which the call number belongs; in the case that the target level belongs to the preset high-risk level, determining the call number as a risk number to be analyzed.

[0157] Optionally, the above-mentioned processor may also execute the program code of the following steps: obtaining the network behavior data corresponding to the risk number to be analyzed, where the network behavior data at least includes: Internet access behavior data and application usage records; determining the network behavior data and call data corresponding to the risk number to be analyzed as risk behavior data, and performing risk analysis using a preset classification model to obtain the number category to which the risk number to be analyzed belongs.

[0158] Optionally, the above-mentioned processor may also execute the program code of the following steps: obtaining a confusion matrix, where the confusion matrix records multiple sample numbers and the sample behavior data corresponding to each sample number; setting a category label for each sample number in the confusion matrix, where the category label is used to indicate the number category to which the sample number belongs, and the number categories at least include: fraud number, victim, drainage number, deliveryman, normal use; using the confusion matrix with added category labels as training data to train the preset classification model.

[0159] Those of ordinary skill in the art can understand that Figure 7 the structure shown is only for illustration, and the computer terminal may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, and terminal devices such as Mobile Internet Devices (MID), PAD, etc. Figure 7 It does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 70 may include more or fewer components (such as a network interface, a display device, etc.) than those shown in Figure 7 and may have a different configuration from that shown in Figure 7 the figure.

[0160] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the relevant hardware of the terminal device through a computer program, and the computer program can be stored in a non-volatile medium. The non-volatile storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disc, etc.

[0161] An embodiment of the present invention further provides a non-volatile storage medium. Optionally, in this embodiment, the non-volatile storage medium can be used to store the program code executed by the risk identification method for the call number provided in the above embodiment. Optionally, in this embodiment, the non-volatile storage medium can be located in any one of the computer terminals in the computer terminal group in the computer network, or in any one of the mobile terminals in the mobile terminal group.

[0162] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for executing the following steps: obtaining multiple call records to be analyzed, wherein each call record to be analyzed includes at least: call data corresponding to multiple call numbers, the call data including multiple call parameters; using a weighted scoring model to analyze the call data corresponding to the call number to obtain a risk score for the call number, wherein the weighted scoring model uses a feature weight pre-assigned to each call parameter to perform a weighted sum of parameter scores corresponding to multiple call parameters of the same call number; from the call numbers corresponding to the multiple call records to be analyzed, respectively, screen the call numbers whose risk scores meet preset scoring conditions to obtain the risk numbers to be analyzed; using a preset classification model to perform a risk analysis on the risk behavior data corresponding to the risk numbers to be analyzed, to obtain the number category to which the risk numbers to be analyzed belong, wherein the risk behavior data includes at least: call data, and the number category includes at least the risk category to which the risk numbers belong.

[0163] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for executing the following steps: receiving multiple original call record data provided by a call record database, wherein each original call record data records at least: the calling number, the called number, the call duration and the number location of the called number; performing data preprocessing on the multiple original call record data to obtain multiple call records to be analyzed, wherein the call records to be analyzed include call data corresponding to multiple call numbers determined based on the multiple original call record data, and the multiple call parameters in the call data corresponding to each call number include at least: the number of calls when the call number is the calling number, the call duration and the caller ratio, and the number dispersion of the multiple called numbers called by the call number, and the number dispersion is determined based on the number locations of the multiple called numbers called by the call number.

[0164] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for executing the following steps: performing data analysis on the call data corresponding to each call number to obtain a parameter score corresponding to each call parameter in the call data; using a weighted scoring model using feature weights pre-set for each call parameter, performing weighted summation on the parameter scores corresponding to multiple call parameters of the same call number to obtain a risk score for the call number.

[0165] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: analyzing the call data corresponding to each historical number in multiple groups of historical call data using a weighted scoring model to obtain an estimated score for each historical number, where each group of historical call data includes: the call data corresponding to the historical number and the determined risk score of the historical number; detecting the score difference between the estimated score and the determined risk score of the same historical number; and adjusting the feature weights in the weighted scoring model according to the score difference.

[0166] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: matching the risk score of each call number with the score intervals corresponding to multiple preset levels to obtain the target level to which the call number belongs; and determining the call number as a risk number to be analyzed when the target level belongs to a preset high-risk level.

[0167] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: obtaining the network behavior data corresponding to the risk number to be analyzed, where the network behavior data at least includes: Internet access behavior data and application usage records; determining the network behavior data and call data corresponding to the risk number to be analyzed as risk behavior data, and performing risk analysis using a preset classification model to obtain the number category to which the risk number to be analyzed belongs.

[0168] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: obtaining a confusion matrix, where the confusion matrix records multiple sample numbers and the sample behavior data corresponding to each sample number; setting a category label for each sample number in the confusion matrix, where the category label is used to indicate the number category to which the sample number belongs, and the number categories at least include: fraud number, victim, traffic diversion number, deliveryman, normal use; and using the confusion matrix with added category labels as training data to train a preset classification model.

[0169] An embodiment of the present invention also provides a computer program product, including a computer program. Optionally, in this embodiment, when the computer program is executed by a processor, it implements the steps of the risk identification method for call numbers provided in the above embodiment.

[0170] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments.

[0171] In the above embodiments of the present invention, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0172] In several embodiments provided in the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of units or modules can be in electrical or other forms.

[0173] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0174] In addition, in each embodiment of the present invention, the functional units can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0175] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a non-volatile storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a non-volatile storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned non-volatile storage media include: USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs, etc., which can store program codes.

[0176] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A risk identification method for call numbers, characterized in that, include: Acquire multiple call records to be analyzed, wherein each of the call records to be analyzed includes at least: call data corresponding to multiple call numbers, the call data including multiple call parameters; Using a weighted scoring model to analyze the call data corresponding to the call number to obtain a risk score for the call number, wherein the weighted scoring model uses a feature weight that is pre-set for each call parameter to perform a weighted sum of parameter scores corresponding to multiple call parameters of the same call number; From the call numbers respectively corresponding to the plurality of call records to be analyzed, screening the call numbers whose risk scores meet the preset scoring conditions to obtain the risk numbers to be analyzed; A preset classification model is used to perform risk analysis on the risk behavior data corresponding to the risk number to be analyzed, and the number category to which the risk number to be analyzed belongs is obtained, wherein the risk behavior data at least includes: the call data, and the number category at least includes the risk category to which the risk number belongs.

2. The method according to claim 1, wherein Obtaining multiple call records to be analyzed includes: Receive multiple original call record data provided by the call record database, wherein each of the original call record data records at least: the calling number, the called number, the call duration and the number location of the called number; Data preprocessing is performed on the multiple original call bill data to obtain the multiple call bills to be analyzed, wherein the call bills to be analyzed include the call data corresponding to the multiple call numbers determined based on the multiple original call bill data, and the multiple call parameters in the call data corresponding to each call number include at least: the number of calls, call duration and caller ratio of the call number as the calling number, and the number dispersion of the multiple called numbers called by the calling number, and the number dispersion is determined according to the number locations of the multiple called numbers called by the calling number.

3. The method according to claim 1, wherein The call data corresponding to each call number is analyzed using a weighted scoring model to obtain a risk score for each call number, including: Performing data analysis on the call data corresponding to each call number to obtain the parameter score corresponding to each call parameter in the call data; The weighted scoring model is used to utilize the feature weights pre-set for each call parameter to perform weighted summation on the parameter scores corresponding to multiple call parameters of the same call number to obtain the risk score of the call number.

4. The method according to claim 1, wherein Before analyzing the call data corresponding to each call number using a weighted scoring model to obtain a risk score for each call number, the method further includes: Using the weighted scoring model to analyze the call data corresponding to each historical number in the multiple groups of historical call data, to obtain an estimated score for each historical number, wherein each group of historical call data includes: the call data corresponding to the historical number, and the determined risk score of the historical number; Detect the scoring difference between the predicted score and the determined risk score of the same historical number; Adjust the feature weights in the weighted scoring model according to the scoring difference.

5. The method according to claim 1, characterized in that, Screen the call numbers whose risk scores meet the preset scoring conditions as the risk numbers to be analyzed, including: Match the risk score of each call number with the scoring intervals corresponding to multiple preset levels to obtain the target level to which the call number belongs; In the case where the target level belongs to the preset high-risk level, determine the call number as the risk number to be analyzed.

6. The method according to claim 1, characterized in that, Use a preset classification model to perform risk analysis on the risk behavior data corresponding to the risk number to be analyzed to obtain the number category to which the risk number to be analyzed belongs, including: Obtain the network behavior data corresponding to the risk number to be analyzed, where the network behavior data at least includes: Internet access behavior data and application usage records; Determine the network behavior data and the call data corresponding to the risk number to be analyzed as the risk behavior data, and use the preset classification model to perform risk analysis to obtain the number category to which the risk number to be analyzed belongs.

7. The method according to claim 1, characterized in that Before using the preset classification model to perform risk analysis on the risk behavior data corresponding to the risk number to be analyzed to obtain the number category to which the risk number to be analyzed belongs, the method further includes: Obtain a confusion matrix, where the confusion matrix records multiple sample numbers and the sample behavior data corresponding to each sample number; Set a category label for each sample number in the confusion matrix, where the category label is used to indicate the number category to which the sample number belongs, and the number category at least includes: fraud number, victim, traffic diversion number, deliveryman, normal use; Use the confusion matrix with the added category labels as training data to train the preset classification model.

8. A risk identification device for call numbers, characterized in that, Including: An acquisition module, configured to acquire multiple call bills to be analyzed, where each call bill to be analyzed at least includes: call data corresponding to multiple call numbers, and the call data includes multiple call parameters; A first analysis module, configured to use a weighted scoring model to analyze the call data corresponding to the call number to obtain the risk score of the call number, where the weighted scoring model uses the feature weights preset for each call parameter to perform weighted summation on the parameter scores corresponding to multiple call parameters of the same call number; A screening module, configured to screen the call numbers whose risk scores meet the preset scoring conditions from the call numbers corresponding to multiple call bills to be analyzed to obtain the risk numbers to be analyzed; A second analysis module, configured to use a preset classification model to perform risk analysis on the risk behavior data corresponding to the risk number to be analyzed to obtain the number category to which the risk number to be analyzed belongs, where the risk behavior data at least includes: the call data, and the number category at least includes the risk categories to which the risk numbers belong.

9. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to execute the risk identification method for the call number according to any one of claims 1 to 7 through the computer program.

10. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, the steps of the risk identification method for the call number according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Intermediate number call risk judgment method and device, electronic equipment and storage medium

    CN120583177A

  • Number identification method, device and equipment, medium and computer program product

    CN121771325A