Abnormal number filtering method and device, computer equipment and storage medium
By obtaining communication behavior data in real time and using multi-dimensional number features and machine learning model training, combined with a distributed computing framework, dynamically update the black and white list library, solving the real-time, accuracy and scalability of abnormal number recognition, and achieving efficient number filtering and recognition.
Patent Information
- Application Number
- CN202510433231.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-08
AI Technical Summary
The prior art has problems such as insufficient real-time, limited accuracy and poor scalability when identifying and filtering abnormal numbers, making it difficult to effectively identify and intercept empty numbers, invalid numbers and abnormal risk numbers.
Real-time acquisition of communication behavior data, multi-dimensional number feature extraction and machine learning model training, combined with distributed computing framework for parallel detection, dynamically update the black and white list library, and realize efficient identification and filtering of numbers.
It realizes accurate identification of empty numbers, invalid numbers and abnormal numbers, significantly improves identification efficiency and accuracy, shortens response time, reduces operational costs and resource waste, and adapts to massive concurrent environments.
Smart Images

Figure CN120282142A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of programming, and in particular to an abnormal number filtering method, device, computer device and storage medium. Background Art
[0002] With the rapid development of communication services, telephone numbers are increasingly widely used in scenarios such as SMS marketing, telephone customer service, and mobile application registration verification. However, the existence of empty numbers, invalid numbers, and abnormal numbers (such as numbers suspected of fraud or malicious calls) has brought significant technical challenges and business risks to communication service providers and enterprises, specifically manifested as:
[0003] Resource waste: Sending SMS or initiating call requests to empty numbers or invalid numbers results in the ineffective occupation of communication resources (such as bandwidth and computing resources), increasing system load and operation and maintenance costs;
[0004] Decline in marketing effectiveness: In the scenario of mass marketing, a too high proportion of invalid numbers will significantly reduce the user reach rate, directly affecting the marketing conversion rate and return on investment of enterprises;
[0005] Communication security risks: Abnormal numbers may trigger malicious behaviors such as fraud and harassment, threatening the stability of the communication network and the security of user information, and even causing legal compliance issues.
[0006] Currently, the industry mainly adopts the following technical means to address the above problems:
[0007] Third-party number verification tools: Rely on external service interfaces to simply query the number status, but the covered scenarios are limited and it is difficult to update in real time;
[0008] Manual or semi-manual cleaning: Periodically clean the number library through manual verification or script rules, but the efficiency is low and it is prone to missed detection and misjudgment;
[0009] Rule-based black and white list filtering: Use fixed rules (such as number segment matching, call frequency threshold) or static list libraries for preliminary screening, but the rules have poor flexibility and it is difficult to adapt to the dynamic changes of number abnormal behaviors (such as new fraud number patterns).
[0010] However, the existing technologies have the following core defects:
[0011] Lack of real-time performance: Manual cleaning and rule updates are lagging, unable to meet the real-time detection requirements of massive number libraries;
[0012] Limited accuracy: Traditional rules rely on manual experience and have weak recognition capabilities for complex abnormal features (such as device fingerprint forgery, number behavior time series patterns);
[0013] Poor scalability: The centralized processing architecture is difficult to support high-concurrency scenarios, and the maintenance cost of the rule library increases exponentially with the growth of data volume.
[0014] Therefore, it is necessary to provide a technical solution that integrates big data processing and intelligent analysis to solve the bottlenecks of the existing methods in terms of real-time performance, accuracy, and scalability, and at the same time achieve efficient identification and dynamic interception of invalid numbers, invalid numbers, and abnormal risk numbers. Summary of the Invention
[0015] The technical problem to be solved by the present invention is: to provide an abnormal number filtering method, device, computer device, and storage medium, aiming to solve the accuracy problem of abnormal number identification.
[0016] To solve the above technical problems, the technical solution adopted by the present invention is: an abnormal number filtering method, including the following steps:
[0017] S10. Real-time obtain the communication behavior data, historical verification records, and device feature data of the number to be detected;
[0018] S20. Preprocess the communication behavior data based on preset data cleaning rules, including denoising, missing value filling, and outlier removal, and extract multi-dimensional number features, where the multi-dimensional number features include number activity, call frequency, device fingerprint consistency, and historical verification failure rate;
[0019] S30. Train a number classification prediction model through cross-validation and hyperparameter tuning methods. The number classification prediction model uses valid numbers, invalid numbers, and abnormal numbers in historical labeled data as supervision labels, and introduces feature importance analysis to optimize feature selection;
[0020] S40. Input the preprocessed number features into the number classification prediction model, output the validity prediction result and confidence score, and trigger an artificial review process based on the confidence threshold;
[0021] S50. Use a distributed computing framework to perform parallel detection on a large number of numbers, allocate computing resources through a dynamic task scheduling algorithm, and monitor the system load in real time and generate alarm information;
[0022] S60. Update the blacklist and whitelist libraries based on the prediction results. The blacklist library supports multi-level classification and dynamically adjusts the level according to the behavior pattern. The whitelist library provides a fast channel to skip repeated detection;
[0023] S70. Output the detection results to the downstream system and receive business feedback data to close the loop and optimize the number classification prediction model and the list library.
[0024] Further, step S30 specifically includes:
[0025] S31. Use the stratified sampling technique to balance the class distribution in the historical labeled data, and perform time series segmentation on the training set, validation set, and test set;
[0026] S32. Use grid search or Bayesian optimization algorithm to tune the hyperparameters of the number classification prediction model, and generate the feature importance ranking through SHAP value analysis;
[0027] S33. Integrate the prediction results of multiple base models, and improve the classification accuracy through weighted voting or stacking generalization methods.
[0028] Further, step S40 specifically includes:
[0029] S41. Standardize the input number features during real-time detection, and cache the historical number feature data through the sliding window mechanism to capture the temporal correlation;
[0030] S42. Quickly integrate the newly labeled data into the number classification prediction model through incremental learning or online learning algorithms;
[0031] S43. Trigger the manual review process for the confidence prediction results below the preset threshold, and synchronize the review results to the training data set.
[0032] Further, step S50 specifically includes:
[0033] S51. Build a distributed computing cluster based on the Spark or Flink framework, and adopt an elastic resource allocation strategy to dynamically expand the computing nodes;
[0034] S52. Perform joint analysis on real-time stream data and historical batch data through a stream-batch integrated processing engine;
[0035] S53. Use a load balancing algorithm to allocate detection tasks, and perform real-time monitoring and fault tolerance recovery on the task execution status.
[0036] Further, step S60 specifically includes:
[0037] S61. Automatically update the blacklist based on business feedback data, and set the conversion rules between the temporary blacklist and the permanent blacklist;
[0038] S62. Periodically verify the numbers in the whitelist, and eliminate the invalid or abnormal numbers;
[0039] S63. Audit the list update operations through a rule engine to ensure compliance and traceability.
[0040] Further, step S70 specifically includes:
[0041] S71. Generate a multi-dimensional detection report, which includes a number validity score, an abnormal risk level, and recommended operations;
[0042] S72. Push the real-time detection results to the downstream system through a message queue, and support a backpressure mechanism to handle traffic peaks;
[0043] S73. Statistically analyze the business feedback data and generate suggestions for optimizing the number classification prediction model and adjusting rules.
[0044] Further, step S73 specifically includes:
[0045] S731. Transmit the business result data of the downstream system back to the number classification prediction model to trigger a retraining process;
[0046] S732. Dynamically adjust the update strategy of the black and white lists and the confidence threshold based on the feedback data;
[0047] S733. Verify the effect of the iterative version of the number classification prediction model through a control test and select the optimal version for production deployment.
[0048] The present invention also provides an abnormal number filtering device, including:
[0049] A data access module, which is used to obtain the communication behavior data, historical verification records, and device feature data of the numbers to be detected in real time;
[0050] A data preprocessing module, which is used to preprocess the communication behavior data based on preset data cleaning rules, including denoising, missing value filling, and outlier removal, and extract multi-dimensional number features, where the multi-dimensional number features include number activity, call frequency, device fingerprint consistency, and historical verification failure rate;
[0051] A model training module, which is used to train a number classification prediction model through cross-validation and hyperparameter tuning methods. The number classification prediction model uses valid numbers, empty numbers, and abnormal numbers in historical labeled data as supervision labels, and introduces feature importance analysis to optimize feature selection;
[0052] A model prediction module, which is used to input the preprocessed number features into the number classification prediction model, output a validity prediction result and a confidence score, and trigger an artificial review process based on the confidence threshold;
[0053] A distributed detection execution module, which is used to perform parallel detection on a large number of numbers using a distributed computing framework, allocate computing resources through a dynamic task scheduling algorithm, monitor the system load in real time, and generate alarm information;
[0054] The black and white list management module is used to update the blacklist and whitelist libraries based on the prediction results. The blacklist library supports multi-level classification and dynamically adjusts the levels according to the behavior patterns. The whitelist library provides a fast track to skip duplicate detections;
[0055] The result feedback module is used to output the detection results to the downstream system and receive business feedback data to close the loop and optimize the number classification prediction model and the list libraries.
[0056] The present invention also provides a computer device, which includes a memory and a processor. A computer program is stored on the memory, and when the processor executes the computer program, the abnormal number filtering method as described above is implemented.
[0057] The present invention also provides a storage medium, which stores a computer program. When the computer program is executed by a processor, the abnormal number filtering method as described above can be implemented.
[0058] The beneficial effects of the present invention are as follows: By using historical data and multi-dimensional features to train a machine learning model, more accurate identification of empty numbers, invalid numbers, and abnormal numbers is achieved, and both the accuracy and the identification efficiency are greatly improved, thereby effectively reducing the phenomena of missed detection and misjudgment; With the help of a distributed parallel processing framework, rapid detection and filtering of number data can be carried out in a massive concurrent environment, significantly shortening the response time of number cleaning and meeting the real-time requirements of large-scale marketing and communication services. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] The specific structure of the present invention is described in detail below with reference to the accompanying drawings.
[0060] Figure 1 It is a flowchart of the abnormal number filtering method according to an embodiment of the present invention;
[0061] Figure 2 It is a flowchart of the model training according to an embodiment of the present invention;
[0062] Figure 3 It is a flowchart of the model prediction according to an embodiment of the present invention;
[0063] Figure 4 It is a flowchart of the distributed detection execution according to an embodiment of the present invention;
[0064] Figure 5 It is a flowchart of the black and white list management according to an embodiment of the present invention;
[0065] Figure 6 It is a flowchart of the result feedback according to an embodiment of the present invention;
[0066] Figure 7 It is a flowchart of the model optimization according to an embodiment of the present invention;
[0067] Figure 8 Block diagram of the abnormal number filtering device according to an embodiment of the present invention;
[0068] Figure 9 Schematic block diagram of a computer device according to an embodiment of the present invention. Detailed implementation manners
[0069] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0070] It should be understood that when used in this specification and the appended claims, the terms "comprises" and "comprising" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0071] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in this specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0072] It should be further understood that the term " / and / " used in this specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0073] As Figure 1 shown, an embodiment of the present invention is: an abnormal number filtering method, including the following steps:
[0074] S10. Real-time obtain the communication behavior data, historical verification records, and device feature data of the number to be detected.
[0075] In this embodiment, numbers and their related feature information are collected from multiple data sources (such as operator-side data, historical call records, third-party number data, etc.). Multiple data protocols (HTTP, FTP, Kafka) and data formats (JSON, CSV) are supported. By initially cleaning the raw data, obvious invalid data (such as format errors, null values, etc.) is removed.
[0076] S20. Preprocess the communication behavior data based on preset data cleaning rules, including denoising, missing value filling, and outlier removal, and extract multi-dimensional number features, where the multi-dimensional number features include number activity, call frequency, device fingerprint consistency, and historical verification failure rate.
[0077] In this embodiment, duplicate removal, missing value processing, and feature field extraction (such as call times, SMS response rate, whether there is a high-frequency outgoing call record, etc.) are performed on the number data; the preprocessed data is saved in a structured or distributed storage form (such as HDFS, NoSQL database, etc.) for model training and detection calls.
[0078] S30. Train a number classification prediction model through cross-validation and hyperparameter tuning methods. The number classification prediction model uses valid numbers, empty numbers, and abnormal numbers in the historical annotation data as supervision labels, and introduces feature importance analysis to optimize feature selection.
[0079] In a specific embodiment, as Figure 2 shown, step S30 specifically includes:
[0080] S31. Use the stratified sampling technique to balance the class distribution in the historical annotation data, and perform time series segmentation on the training set, validation set, and test set;
[0081] S32. Use the grid search or Bayesian optimization algorithm to tune the hyperparameters of the number classification prediction model, and generate a feature importance ranking through SHAP value analysis;
[0082] S33. Integrate the prediction results of multiple base models to improve the classification accuracy through weighted voting or stacking generalization methods.
[0083] In this embodiment, a number classification prediction model is trained using historical annotation data (such as "valid number", "empty number", "abnormal number"). A classification model based on random forest, XGBoost, or deep learning algorithm is trained to improve the model performance through cross-validation and hyperparameter tuning. At the same time, feature importance analysis is introduced to optimize feature selection and improve the model generalization ability.
[0084] S40. Input the preprocessed number features into the number classification prediction model, output the validity prediction result and confidence score, and trigger the manual review process based on the confidence threshold.
[0085] In a specific embodiment, as Figure 3 shown, step S40 specifically includes:
[0086] S41. Standardize the input number features during real-time detection, and cache the historical number feature data through a sliding window mechanism to capture the temporal correlation;
[0087] S42. Rapidly integrate newly labeled data into the number classification prediction model through incremental learning or online learning algorithms;
[0088] S43. Trigger a manual review process for prediction results with confidence levels below a preset threshold, and synchronize the review results to the training dataset.
[0089] In this embodiment, during real-time detection, the preprocessed number features are input into the number classification prediction model to output prediction results (such as "valid", "invalid number", "abnormal"). Support hot updates of the number classification prediction model to ensure that new data can be quickly applied to predictions. Provide confidence scores to assist downstream system decision-making (such as manual review for numbers with low confidence).
[0090] S50. Use a distributed computing framework to perform parallel detection on a large number of numbers, allocate computing resources through a dynamic task scheduling algorithm, and monitor the system load in real time and generate alarm information.
[0091] In a specific embodiment, as Figure 4 shown, step S50 specifically includes:
[0092] S51. Build a distributed computing cluster based on the Spark or Fl ink framework, and dynamically expand computing nodes using an elastic resource allocation strategy;
[0093] S52. Jointly analyze real-time stream data and historical batch data through a stream-batch integrated processing engine;
[0094] S53. Use a load balancing algorithm to allocate detection tasks, and perform real-time monitoring and fault tolerance recovery on the task execution status.
[0095] In this embodiment, a distributed computing framework is used to perform parallel detection on a large number of numbers to ensure real-time performance and efficiency in high-concurrency scenarios. Implement distributed computing based on Spark or Fl ink, support horizontal expansion, use a dynamic task scheduling algorithm, optimize task allocation according to system load and resource utilization, and provide real-time monitoring and alarm functions to ensure stable operation of the system.
[0096] S60. Update the blacklist and whitelist libraries based on the prediction results. The blacklist library supports multi-level classification and dynamically adjusts levels according to behavior patterns. The whitelist library provides a fast track to skip duplicate detections.
[0097] In a specific embodiment, as Figure 5 shown, step S60 specifically includes:
[0098] S61. Automatically update the blacklist based on business feedback data, and set the conversion rules between the temporary blacklist and the permanent blacklist;
[0099] S62. Periodically verify the numbers in the whitelist and remove invalid or abnormal numbers;
[0100] S63. Audit the list update operations through a rules engine to ensure compliance and traceability. - In this embodiment, maintain and dynamically update the blacklist and whitelist to ensure real-time interception of abnormal numbers and efficient passage of signalable numbers. Mark and confirm newly emerged abnormal numbers and under certain strategies and threshold conditions.
[0101] Blacklist management:
[0102] Based on the prediction results of the machine learning model and business feedback (such as SMS return, call failure), dynamically add high-risk numbers. Support multi-level blacklists (such as temporary blacklists, permanent blacklists), and adjust the levels according to the behavior patterns.
[0103] Whitelist management: Based on historical data and business feedback, maintain a signalable number library. Provide a fast track for the whitelist to reduce repeated detection of signalable numbers.
[0104] Dynamic update mechanism:
[0105] Regularly clean up expired or no longer abnormal numbers to ensure the accuracy and timeliness of the list library. Support real-time effectiveness to ensure that newly discovered abnormal numbers can be intercepted immediately.
[0106] S70. Output the detection results to the downstream system and receive business feedback data to close the loop and optimize the number classification prediction model and the list library.
[0107] In a specific embodiment, as Figure 6 shown, step S70 specifically includes:
[0108] S71. Generate a multi-dimensional detection report, which includes number validity score, abnormal risk level, and recommended operations;
[0109] S72. Push the real-time detection results to the downstream system through a message queue and support a backpressure mechanism to handle traffic peaks;
[0110] S73. Statistically analyze the business feedback data and generate suggestions for optimizing the number classification prediction model and adjusting the rules.
[0111] In a specific embodiment, as Figure 7 shown, step S73 specifically includes:
[0112] S731. Transmit the business result data of the downstream system back to the number classification prediction model to trigger the retraining process;
[0113] S732. Dynamically adjust the update strategies and confidence thresholds of the black and white lists based on the feedback data;
[0114] S733. Verify the effect of the iterative version of the number classification prediction model through control tests, and select the optimal version for production deployment.
[0115] In this embodiment, the detection results and the filtering list are output to downstream applications, such as a short message gateway, a call center system, or a customer relationship management (CRM) system; multiple output forms (such as API interfaces, report files) are provided to support real-time and batch output. The output content includes the number validity score, the abnormal risk level, and recommended operations (such as interception, release, manual review). Collect the actual business results of downstream systems (such as the short message return rate, the call connection rate) to form a closed-loop feedback. The feedback data is used for retraining the number classification prediction model and optimizing the list library to improve the overall performance of the system.
[0116] In summary, the beneficial effects of the embodiments of the present application are as follows:
[0117] The real-time performance is significantly improved: With the help of the distributed parallel processing framework, the present invention can quickly detect and filter number data in a massive concurrent environment, significantly shortening the response time of number cleaning and meeting the real-time requirements of large-scale marketing and communication services.
[0118] The detection accuracy and efficiency are greatly improved: By using historical data and multi-dimensional features to train the machine learning model, more accurate identification of invalid numbers, invalid numbers, and abnormal numbers is achieved. Compared with traditional single rules or manual judgments, both the accuracy rate and the identification efficiency are greatly improved, thus effectively reducing the phenomena of missed detection and misjudgment.
[0119] Good scalability: By adopting a distributed computing architecture, the computing resources can be horizontally expanded according to the growth level of the business volume to achieve real-time processing of larger-scale number data, avoiding performance bottlenecks or downtime risks caused by explosive data growth in the system.
[0120] Dynamic adaptive ability: With the help of the continuously updated blacklist and whitelist mechanisms, as well as the feedback loop of the detection results, the present invention can automatically adapt to new abnormal number patterns or feature changes, timely optimize the detection model and filtering strategy, and reduce the dependence on manual intervention and regular batch updates.
[0121] The operation cost and resource waste are significantly reduced: The present invention significantly reduces the communication attempts of invalid numbers while also reducing the waste of marketing resources and network resources; the timely identification and filtering of abnormal numbers contribute to improving the user experience and the security of the communication network.
[0122] Wide range of applicable scenarios: The present invention can be compatible with multiple data sources and business scenarios, and can be applied to both short message marketing and call centers, and can also be docked with various customer management and business verification systems, with high flexibility and versatility.
[0123] As Figure 8 shown, an embodiment of the present invention further provides an abnormal number filtering device, including:
[0124] A data access module 10, configured to obtain communication behavior data, historical verification records, and device feature data of a number to be detected in real time.
[0125] A data preprocessing module 20, configured to preprocess the communication behavior data based on preset data cleaning rules, including denoising, missing value filling, and outlier removal, and extract multi-dimensional number features, where the multi-dimensional number features include number activity, call frequency, device fingerprint consistency, and historical verification failure rate.
[0126] A model training module 30, configured to train a number classification prediction model through cross-validation and hyperparameter tuning methods, where the number classification prediction model uses valid numbers, empty numbers, and abnormal numbers in historical labeled data as supervision labels, and introduces feature importance analysis to optimize feature selection.
[0127] In a specific embodiment, the model training module 30 is specifically configured to:
[0128] Use the stratified sampling technique to balance the class distribution in the historical labeled data, and perform time series segmentation on the training set, validation set, and test set;
[0129] Use the grid search or Bayesian optimization algorithm to tune the hyperparameters of the number classification prediction model, and generate a feature importance ranking through SHAP value analysis;
[0130] Integrate the prediction results of multiple base models, and improve the classification accuracy through weighted voting or stacking generalization methods.
[0131] A model prediction module 40, configured to input the preprocessed number features into the number classification prediction model, output a validity prediction result and a confidence score, and trigger an artificial review process based on a confidence threshold.
[0132] In a specific embodiment, the model prediction module 40 is specifically configured to:
[0133] Perform standardization processing on the input number features during real-time detection, and cache historical number feature data through a sliding window mechanism to capture temporal correlation;
[0134] Quickly integrate new labeled data into the number classification prediction model through incremental learning or online learning algorithms;
[0135] Trigger an artificial review process for confidence prediction results below a preset threshold, and synchronize the review results to the training data set.
[0136] The distributed detection execution module 50 is used to perform parallel detection on a large number of numbers using a distributed computing framework, allocate computing resources through a dynamic task scheduling algorithm, and monitor the system load in real time and generate alarm information.
[0137] In a specific embodiment, the distributed detection execution module 50 is specifically used for:
[0138] Build a distributed computing cluster based on the Spark or Fl ink framework, and dynamically expand computing nodes using an elastic resource allocation strategy;
[0139] Perform joint analysis on real-time stream data and historical batch data through a stream-batch integrated processing engine;
[0140] Use a load balancing algorithm to allocate detection tasks, and perform real-time monitoring and fault tolerance recovery on the task execution status.
[0141] The black and white list management module 60 is used to update the blacklist and whitelist libraries based on the prediction results. The blacklist library supports multi-level classification and dynamically adjusts the level according to the behavior pattern. The whitelist library provides a fast channel to skip duplicate detections.
[0142] In a specific embodiment, the black and white list management module 60 is specifically used for:
[0143] Automatically update the blacklist based on business feedback data, and set the conversion rules between the temporary blacklist and the permanent blacklist;
[0144] Periodically verify the numbers in the whitelist, and eliminate invalid or abnormal numbers;
[0145] Audit the list update operations through a rule engine to ensure compliance and traceability.
[0146] The result feedback module 70 is used to output the detection results to the downstream system, and receive business feedback data to close the loop and optimize the number classification prediction model and the list library.
[0147] In a specific embodiment, the result feedback module 70 is specifically used for:
[0148] Generate a multi-dimensional detection report, which includes the number validity score, abnormal risk level, and recommended operations;
[0149] Push the real-time detection results to the downstream system through a message queue, and support a backpressure mechanism to handle traffic peaks;
[0150] Perform statistical analysis on the business feedback data, and generate suggestions for optimizing the number classification prediction model and adjusting the rules.
[0151] In a specific embodiment, the statistical analysis of business feedback data to generate suggestions for optimizing the number classification prediction model and adjusting rules specifically includes:
[0152] Transmit the business result data of the downstream system back to the number classification prediction model to trigger the retraining process;
[0153] Dynamically adjust the update strategy of the black and white lists and the confidence threshold based on the feedback data;
[0154] Verify the effect of the iterative version of the number classification prediction model through a control test, and select the optimal version for production deployment. Conduct statistical analysis on the business feedback data to generate suggestions for optimizing the number classification prediction model and adjusting rules.
[0155] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the above abnormal number filtering device can refer to the corresponding description in the foregoing method embodiments. For the sake of convenience and brevity of description, it will not be elaborated here.
[0156] The above abnormal number filtering device can be implemented in the form of a computer program, and this computer program can run on a computer device as shown in Figure 9 shown.
[0157] Please refer to Figure 9 , Figure 9 which is a schematic block diagram of a computer device provided by an embodiment of the present application. The computer device 500 can be a terminal or a server. Among them, the terminal can be an electronic device with communication functions such as a smart phone, a tablet computer, a notebook computer, a desktop computer, a personal digital assistant, and a wearable device. The server can be an independent server or a server cluster composed of multiple servers.
[0158] Refer to Figure 9 , the computer device 500 includes a processor 502, a memory, and a network interface 505 connected through a system bus 501. Among them, the memory can include a non-volatile storage medium 503 and an internal memory 504.
[0159] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions, and when these program instructions are executed, the processor 502 can be made to execute an abnormal number filtering method.
[0160] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500.
[0161] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can be made to execute an abnormal number filtering method.
[0162] The network interface 505 is used for network communication with other devices. Those skilled in the art can understand that Figure 9 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device 500 to which the solution of this application is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0163] Among them, the processor 502 is used to run the computer program 5032 stored in the memory to implement the abnormal number filtering method as described above.
[0164] It should be understood that in the embodiments of this application, the processor 502 may be a central processing unit (CPU), and the processor 502 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0165] Those of ordinary skill in the art can understand that all or part of the process of implementing the method in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program includes program instructions, and the computer program can be stored in a storage medium, and the storage medium is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above method.
[0166] Therefore, the present invention also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program, where the computer program includes program instructions. When the program instructions are executed by the processor, the processor is made to execute the abnormal number filtering method as described above.
[0167] The storage medium may be a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk, an optical disk, or other computer-readable storage media that can store program codes.
[0168] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0169] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of each unit is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0170] The steps in the method embodiments of the present invention can be adjusted, combined, and deleted according to actual needs. The units in the device embodiments of the present invention can be combined, divided, and deleted according to actual needs. In addition, the functional units in each embodiment of the present invention can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0171] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention.
[0172] As described above, the above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. An abnormal number filtering method, characterized in that, The following steps are involved: S10, real-time acquisition of communication behavior data, historical verification records and device feature data of the number to be detected; S20, preprocessing the communication behavior data based on preset data cleaning rules, including denoising, missing value filling and outlier removal, and extracting multi-dimensional number features, wherein the multi-dimensional number features include number activity, call frequency, device fingerprint consistency and historical verification failure rate; S30. Training a number classification prediction model through cross-validation and hyperparameter tuning methods. The number classification prediction model uses valid numbers, empty numbers, and abnormal numbers in historical annotation data as supervision labels, and introduces feature importance analysis to optimize feature selection; S40, inputting the preprocessed number features into the number classification prediction model, outputting the validity prediction result and the confidence score, and triggering the manual review process based on the confidence threshold; S50, using a distributed computing framework to perform parallel detection on a large number of numbers, allocating computing resources through a dynamic task scheduling algorithm, monitoring system load in real time and generating alarm information; S60, updating the blacklist and whitelist libraries based on the prediction results, wherein the blacklist library supports multi-level classification and dynamically adjusts the level according to the behavior pattern, and the whitelist library provides a fast channel to skip repeated detection; S70. Output the detection results to the downstream system, and receive business feedback data to close the loop and optimize the number classification prediction model and list library.
2. The abnormal number filtering method according to claim 1, characterized in that Step S30 specifically includes: S31. Use stratified sampling techniques to balance the category distribution in historical annotation data and perform time series segmentation on the training set, validation set, and test set. S32. Use grid search or Bayesian optimization algorithm to tune the hyperparameters of the number classification prediction model, and generate feature importance rankings through SHAP value analysis; S33. Integrate the prediction results of multiple base models and improve the classification accuracy through weighted voting or stacked generalization methods.
3. The abnormal number filtering method according to claim 1, characterized in that, Step S40 specifically includes: S41, standardizing the input number features during real-time detection, and caching historical number feature data through a sliding window mechanism to capture temporal correlation; S42. Rapidly integrate new annotation data into the number classification prediction model through incremental learning or online learning algorithms; S43. Trigger a manual review process for the confidence prediction results that are lower than the preset threshold, and synchronize the review results to the training data set.
4. The abnormal number filtering method according to claim 1, characterized in that Step S50 specifically includes: S51. Build a distributed computing cluster based on the Spark or Flink framework and dynamically expand computing nodes using elastic resource allocation strategies; S52, jointly analyzing the real-time streaming data and historical batch data through the streaming and batch integrated processing engine; S53. Use a load balancing algorithm to distribute detection tasks, and perform real-time monitoring and fault-tolerant recovery on the task execution status.
5. The abnormal number filtering method according to claim 1, wherein Step S60 specifically includes: S61. Automatically update the blacklist based on the business feedback data, and set the conversion rules between the temporary blacklist and the permanent blacklist; S62, periodically verify the numbers in the whitelist and remove invalid or abnormal numbers; S63. Audit list update operations through the rule engine to ensure compliance and traceability.
6. The abnormal number filtering method according to claim 1, wherein Step S70 specifically includes: S71. Generate a multi-dimensional detection report, which includes a number validity score, an abnormal risk level, and recommended operations; S72. Push the real-time detection results to the downstream system through a message queue, and support a backpressure mechanism to handle traffic peaks; S73. Statistically analyze the business feedback data and generate suggestions for optimizing the number classification prediction model and adjusting rules.
7. The abnormal number filtering method according to claim 1, characterized in that Step S73 specifically includes: S731. Transmit the business result data of the downstream system back to the number classification prediction model to trigger a retraining process; S732. Dynamically adjust the update strategy and confidence threshold of the black and white lists based on the feedback data; S733. Verify the effect of the iterative version of the number classification prediction model through a control test and select the optimal version for production deployment.
8. An abnormal number filtering device, characterized in that, Including: A data access module for obtaining the communication behavior data, historical verification records, and device feature data of the numbers to be detected in real time; A data preprocessing module for preprocessing the communication behavior data based on preset data cleaning rules, including denoising, missing value filling, and outlier removal, and extracting multi-dimensional number features, where the multi-dimensional number features include number activity, call frequency, device fingerprint consistency, and historical verification failure rate; A model training module for training a number classification prediction model through cross-validation and hyperparameter tuning methods, where the number classification prediction model uses valid numbers, empty numbers, and abnormal numbers in historical labeled data as supervision labels and introduces feature importance analysis to optimize feature selection; A model prediction module for inputting the preprocessed number features into the number classification prediction model, outputting a validity prediction result and a confidence score, and triggering a manual review process based on the confidence threshold; A distributed detection execution module for performing parallel detection on a large number of numbers using a distributed computing framework, allocating computing resources through a dynamic task scheduling algorithm, and monitoring the system load in real time and generating alarm information; A black and white list management module for updating the blacklist and whitelist libraries based on the prediction results, where the blacklist library supports multi-level classification and dynamically adjusts the level according to the behavior pattern, and the whitelist library provides a fast channel to skip repeated detections; A result feedback module for outputting the detection results to the downstream system and receiving business feedback data to close the loop and optimize the number classification prediction model and the list library.
9. A computer device, characterized in that: The computer device includes a memory and a processor, and a computer program is stored on the memory. When the processor executes the computer program, it implements the abnormal number filtering method according to any one of claims 1 to 7.
10. A storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, it can implement the abnormal number filtering method according to any one of claims 1 to 7.