A method and device for number monitoring

By building a number stability prediction model and differentiated verification strategy, the problem of waste of resources and low accuracy in number library maintenance is solved, and the automation and intelligent management of the number library is realized, and the maintenance efficiency and accuracy of the number library are improved.

CN119922258BActive Publication Date: 2025-07-11BEIJING YULORE INNOVATION TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510415153.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-11
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

The existing number library maintenance technology has waste of resources, long update cycles, and the inability to detect number changes in time, resulting in low database accuracy.

Method used

By collecting outbound call verification data, user error correction data and recycling number data, generating training sample sets, using the XGBoost algorithm to build a number stability prediction model, combining Bayesian confidence accelerator and large language model, number stability prediction is carried out, and differentiated verification and automated updates are realized.

Benefits of technology

提高了号码库的维护效率和准确性,降低了资源浪费,实现了号码库的自动化和智能化管理,能够提前识别号码变更风险,为业务调整提供预警。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119922258B_ABST
    Figure CN119922258B_ABST
Patent Text Reader

Abstract

The present application provides a number monitoring method and apparatus. The method includes: generating a training sample set; extracting the historical call detail data of each number from a communication record device storing the historical call records of numbers based on each number included in the training sample set, and generating a number behavior data set; performing feature extraction and processing on the number behavior data set from three dimensions of number features, call behaviors, and call behavior changes, and generating a number feature variable set; based on the number feature variable set, constructing an initial number stability prediction model through the XGBoost algorithm, and generating a number stability prediction model; deploying the number stability prediction model to a business scenario, obtaining the yellow page number data in a number library, and using the number stability prediction model to perform stability prediction on the yellow page number data, and generating a number stability probability value; and updating the number library based on the number stability probability value. The embodiments of the present application achieve efficient maintenance of the number library.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of telecommunication data analysis, and particularly to a method and device for number monitoring. Background Art

[0002] With the development of telecommunication technologies and the growth of user demands, the construction and maintenance of various commercial number databases have become increasingly important. These number databases usually contain the contact numbers of a large number of merchants and are used in various scenarios such as customer service, marketing promotion, and industry information query.

[0003] Currently, the common number database maintenance technologies mainly rely on periodic manual verification or basic web crawling methods for updating. For example, manually call the numbers to confirm their validity, or use web crawlers to collect and update information item by item according to the data sources. Although these methods can ensure a certain degree of data accuracy, they suffer from long update cycles and low efficiency.

[0004] The existing relatively advanced number database maintenance technology adopts a simple regular batch verification strategy, that is, periodically verify all the numbers in the database without discrimination. This technology verifies the validity of the numbers through manual or automatic outbound calling devices and updates the database according to the verification results. This method can ensure a certain timeliness of the data, but it cannot accurately identify and predict the change characteristics of different numbers.

[0005] However, this non-discriminatory verification method has obvious technical defects: verifying all numbers at the same frequency consumes a large amount of manpower and material resources, resulting in waste of resources; due to the long update cycle, it is impossible to timely discover and handle number changes, leading to a large amount of outdated or incorrect information in the database, affecting the overall accuracy; it is impossible to conduct intelligent management according to the actual usage and change rules of the numbers, lacking targeted maintenance strategies. Summary of the Invention

[0006] In view of this, this application provides a method and device for number monitoring, which solves the problems in the prior art of consuming a large amount of manpower and material resources, having a long update cycle, and being unable to timely discover and handle number changes when verifying all numbers at the same frequency.

[0007] An embodiment of the present application provides a number monitoring method, including: collecting positive and negative sample data in outbound call verification data results, user error correction data, and recycled number data to generate a training sample set; based on each number included in the training sample set, extracting the historical call detail data of each number from a communication record device storing the historical call records of the numbers, and preprocessing the historical call detail data to generate a number behavior data set; extracting and processing features from the number behavior data set from three dimensions of number features, call behaviors, and call behavior changes to generate a number feature variable set; based on the number feature variable set, constructing an initial model for predicting number stability through the XGBoost algorithm, and evaluating and tuning parameters of the prediction result of the initial model for predicting number stability to generate a number stability prediction model; deploying the number stability prediction model to a business scenario, obtaining yellow page number data in a number library, and using the number stability prediction model to perform stability prediction on the yellow page number data to generate a number stability probability value; based on the number stability probability value, conducting key verification on numbers higher than a preset probability threshold, and updating the number library according to the verification result.

[0008] Optionally, the extracting the historical call detail data of each number from a communication record device storing the historical call records of the numbers based on each number included in the training sample set, and preprocessing the historical call detail data to generate a number behavior data set includes: determining a preset time window range based on the verification time period of the numbers in the training sample set; for each number in the training sample set, extracting the call detail records within the preset time window from the communication record device; sequentially performing format standardization, missing value processing, and outlier filtering on the call detail records to form preprocessed call detail records; and performing structured processing on the preprocessed call detail records to generate the number behavior data set.

[0009] Optionally, from the three dimensions of the number features, call behaviors, and changes in call behaviors, feature extraction and processing are performed on the number behavior data set to generate a number feature variable set, including: using the number behavior data set, from the number feature dimension, extracting the city of number attribution, operator type, and number segment type information, and outputting a number attribute feature set; according to the number attribute feature set, in combination with the number behavior data set, from the call behavior dimension, counting the number of calls and the number of call objects to form a call behavior feature set; using the call behavior feature set, from the dimension of changes in call behaviors, calculating the statistical features of each behavior index in the call behavior feature set to construct a call behavior change feature set, where the statistical features include at least one of the mean or variance; fusing the number attribute feature set, the call behavior feature set, and the call behavior change feature set to generate the number feature variable set.

[0010] Optionally, evaluating and tuning the parameters of the prediction result of the initial number stability prediction model to generate a number stability prediction model, including: dividing the number feature variable set into a training set and a test set according to a preset ratio; using the test set to perform performance evaluation on the initial number stability prediction model and calculating the performance evaluation index; according to the performance evaluation index, adjusting the parameters of the initial number stability prediction model to generate the number stability prediction model.

[0011] Optionally, after generating the number stability prediction model, it further includes: inputting the prediction result of the number stability prediction model into a reconfigurable streaming accelerator with Bayesian confidence, performing Monte Carlo sampling through the accelerator to generate the uncertainty distribution of the prediction result; according to the uncertainty distribution, performing probability calibration calculation to form a 90% confidence interval of the prediction result; based on the 90% confidence interval value of the prediction result, using a dynamic weight adjustment mechanism to generate a calibrated and optimized prediction result, and updating the calibrated and optimized prediction result to the number stability prediction model for the number stability prediction process.

[0012] Optionally, deploying the number stability prediction model to a business scenario includes: using a serialization method to convert the number stability prediction model into a model file in a persistable format and completing the environment configuration on a business server; using the configured business server environment to construct a timed task scheduler and establish an automated prediction process; according to the automated prediction process, implementing a prediction process monitoring mechanism to continuously record and output the prediction status and result distribution data.

[0013] Optionally, after using the phone number stability prediction model to make a preliminary judgment on the yellow page phone number data, the following steps are further included: using a large language model based on the active inference framework to deeply analyze the phone number stability probability value generated by the phone number stability prediction model, generating an analysis result at the semantic level; combining the analysis result at the semantic level, using a multi-objective optimization strategy to output optimized prediction parameters and an optimized verification strategy; based on the optimized prediction parameters and the optimized verification strategy, iteratively optimize the phone number stability prediction model to improve the accuracy of the phone number stability preliminary judgment.

[0014] Optionally, the step of using a large language model based on the active inference framework to deeply analyze the phone number stability probability value generated by the phone number stability prediction model and generating an analysis result at the semantic level includes: obtaining the phone numbers within a preset probability interval of the phone number stability probability value, and performing in-depth feature analysis on the phone numbers to form a preliminary analysis result; using the preliminary analysis result to retrieve the historical case database and analyzing the temporal changes of the phone number usage pattern to generate the phone number change risk characteristics; based on the active inference framework, fusing the phone number change risk characteristics, and at the same time extracting and processing the outbound call record text and merchant status information related to the phone number, and outputting a comprehensive evaluation result of the phone number change risk as the analysis result at the semantic level.

[0015] Optionally, the step of using a multi-objective optimization strategy to output optimized prediction parameters and an optimized verification strategy includes: collecting data on prediction accuracy, recognition coverage rate, and resource consumption efficiency, and establishing a multi-dimensional optimization index system; using the multi-dimensional optimization index system to calculate the gradient direction of each optimization objective, constructing a Pareto optimal solution set, and determining the optimal balance point; according to the optimal balance point, generating updated prediction weights and threshold parameters, and applying the updated parameters to the phone number stability preliminary judgment process.

[0016] The embodiment of the present application further provides a number monitoring device, and the device includes: a sample collection module, configured to collect positive sample and negative sample data from outbound call verification data results, user error correction data, and recycled number data, and generate a training sample set; a data processing module, configured to extract the historical call detail data of each number from a communication record device storing the historical call records of numbers based on each number included in the training sample set, and preprocess the historical call detail data to generate a number behavior data set; a feature extraction module, configured to perform feature extraction and processing on the number behavior data set from three dimensions of number features, call behaviors, and call behavior changes to generate a number feature variable set; a model training module, configured to construct an initial number stability prediction model through the XGBoost algorithm based on the number feature variable set, and evaluate and optimize parameters of the prediction result of the initial number stability prediction model to generate a number stability prediction model; a prediction module, configured to deploy the number stability prediction model to a service scenario, obtain yellow page number data in a number library, and use the number stability prediction model to perform stability prediction on the yellow page number data to generate a number stability probability value; a verification and update module, configured to perform key verification on numbers higher than a preset probability threshold based on the number stability probability value, and update the number library according to the verification result.

[0017] The embodiment of the present application further provides a computer device, and the computer device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method of the above number monitoring method.

[0018] The embodiment of the present application further provides a computer-readable storage medium, which stores computer instructions for causing a computer to execute the method of the above number monitoring method. The embodiment of the present application further provides a computer program product, including computer instructions, and the computer instructions implement the steps of the above number monitoring method when executed by a processor.

[0019] The present application has the following technical effects: By comprehensively capturing the characteristics of the number status from three dimensions: number features, call behaviors, and changes in call behaviors, the accuracy of the prediction model is improved; The XGBoost algorithm is used to train the number stability prediction model, and through the statistical analysis of the call behaviors of the number, the accurate prediction of the number change risk is realized; Based on the differential verification strategy of the probability threshold, the key monitoring of high-risk numbers is realized, and the problem of resource waste caused by indiscriminate verification is greatly reduced; A closed-loop device for maintaining the number database is constructed, and through the cyclic mechanism of model prediction, key verification, database update, and feedback optimization, the automated and intelligent management of the number database maintenance is realized. Description of the Drawings

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings required for use in the embodiments will be briefly introduced below. The drawings herein are incorporated into the specification and constitute a part of this specification. These drawings show the embodiments that conform to the present disclosure and are used together with the specification to illustrate the technical solutions of the present disclosure. It should be understood that the following drawings only show some embodiments of the present disclosure and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can also be obtained based on these drawings without creative efforts.

[0021] Figure 1 It is a schematic flowchart of a number monitoring method provided by an embodiment of the present application;

[0022] Figure 2 It is a schematic flowchart of step S1 provided by an embodiment of the present application;

[0023] Figure 3 It is a schematic flowchart of step S2 provided by an embodiment of the present application;

[0024] Figure 4 It is a schematic flowchart of step S3 provided by an embodiment of the present application;

[0025] Figure 5 It is a schematic flowchart of step S4 provided by an embodiment of the present application. Detailed Embodiments

[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are only a part rather than all of the embodiments of the present disclosure. Components of the embodiments of the present disclosure generally described and illustrated in the accompanying drawings herein may be arranged and designed in a variety of different configurations. Therefore, the detailed description of the embodiments of the present disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed present disclosure, but merely represents selected embodiments of the present disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of the present disclosure without creative efforts fall within the scope of protection of the present disclosure.

[0027] It should be noted that like reference numerals and letters denote like items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0028] As used herein, the term "and / or" merely describes an association relationship and indicates that three relationships may exist. For example, A and / or B may represent three cases: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the term "at least one" as used herein means any one or any combination of at least two of a plurality. For example, including at least one of A, B, and C may represent any one or more elements selected from the set composed of A, B, and C.

[0029] As Figure 1 shown, an embodiment of the present application provides a number monitoring method, including:

[0030] S1: Collect positive and negative sample data from outbound call verification data results, user error correction data, and recycled number data, and generate a training sample set.

[0031] In this step, the number monitoring device comprehensively collects samples from multiple data sources, including outbound call verification data results, user error correction data, and recycled number data, and extracts positive and negative samples therefrom.

[0032] Positive samples represent stable numbers, and negative samples represent unstable numbers. These multi-source sample data are cleaned and integrated to form a training sample set with reliable quality and perfect annotation, providing a necessary data basis for subsequent machine learning models. This multi-source data fusion acquisition method ensures the diversity and representativeness of the samples, and helps to improve the generalization ability and prediction accuracy of the final model.

[0033] S2: Based on each number included in the training sample set, extract the historical call detail data of each number from the communication record number monitoring device that stores the historical call records, and preprocess the historical call detail data to generate a number behavior data set.

[0034] Determine an appropriate time window range according to the time period of number verification, and then extract the call detail records within this window from the communication record number monitoring device. These records contain rich communication behavior information, such as call time, call duration, call object, etc. After preprocessing such as unifying the format, handling missing values, and filtering outliers for these original call records, the number monitoring device converts them into a structured number behavior data set, laying a data foundation for subsequent feature extraction. This step focuses on solving the problems of data acquisition and quality assurance, ensuring that subsequent analysis is based on complete and consistent data.

[0035] S3: From three dimensions of number features, call behavior, and call behavior changes, perform feature extraction and processing on the number behavior data set to generate a number feature variable set.

[0036] S3 is a feature variable processing step, which is a process of extracting valuable features from number behavior data. The number monitoring device conducts feature engineering from three dimensions: First is the number feature dimension, extracting basic attributes of the number such as the city of origin, operator type, number segment type, etc.; second is the call behavior dimension, calculating behavior indicators such as the number of calls and the number of call objects in different time periods and types; finally is the call behavior change dimension, calculating statistical features such as mean and variance for the behavior indicators to capture the stability and change patterns of the behavior. After the features of these three dimensions are fused, a feature variable set that comprehensively reflects the characteristics of the number is formed. This multi-dimensional feature extraction method is one of the core innovations of this embodiment of the application, enabling the model to understand and predict the stability state of the number from multiple perspectives.

[0037] S4: Based on the number feature variable set, construct an initial model for predicting number stability through the XGBoost algorithm, and evaluate and optimize the parameters of the prediction results of the initial model for predicting number stability to generate a number stability prediction model.

[0038] First, divide the feature variable set into a training set and a test set according to a preset ratio to ensure the balance and representativeness of the data. Then, use the XGBoost algorithm to construct an initial prediction model, which has excellent performance in dealing with classification problems of structured data. By using the test set to evaluate the model performance and calculating key indicators such as AUC and accuracy, the number monitoring device can comprehensively understand the prediction ability of the model. According to the evaluation results, the number monitoring device performs tuning on model parameters such as tree depth and learning rate, and finally generates an optimized number stability prediction model. This step ensures that the model can accurately capture the characteristics and patterns of number stability and provides a reliable prediction basis for practical applications.

[0039] In this embodiment, the construction of the XGBoost model adopts specific parameter settings and architecture designs. The basic learner of the model uses a decision tree (CART tree), the maximum tree depth is tuned to 6, the number of subtrees is 150, and 8 threads are enabled for parallel computing. The feature column sampling is set to 0.8, that is, 80% of the features are randomly used each time a tree is constructed. In terms of training parameters, the learning rate is set to 0.05, and this relatively low learning rate makes the model training more stable; the minimum number of samples in a leaf node is set to 3 to effectively prevent overfitting; the regularization parameter lambda is set to 1.2, and alpha is set to 0.5, which respectively control the L2 and L1 regularization intensities; the sample sampling ratio is 0.8; the sample weight balance parameter is approximately 5.7, that is, the ratio of the number of negative samples to the number of positive samples. The scale of the training data is about 10,000 number samples, among which 8,500 are stable numbers and 1,500 are unstable numbers, and they are divided into a training set and a test set according to an 8:2 ratio. After feature screening, about 60 key features are retained. The 5-fold cross-validation is adopted during the training process to ensure the model stability, and early_stopping_rounds is set to 50 to avoid overfitting, and it actually converges in about 120 rounds of iteration. The evaluation indicators mainly use AUC, and at the same time, the precision and recall are monitored. The loss function uses logistic regression for binary classification problems. The AUC of the model on the test set reaches 0.92, the accuracy is 87%, the recall is 83%, and the screening accuracy under the high threshold setting (0.8) reaches 95%. The sources of the training data include the call result data recorded by the outbound verification device (about 60%), the user's active error correction feedback data (about 25%), and the number recycling data provided by the operator (about 15%).

[0040] S5: Deploy the number stability prediction model to the business scenario, obtain the yellow page number data in the number library, and use the number stability prediction model to perform stability prediction on the yellow page number data to generate a number stability probability value.

[0041] Convert the model into a persistent format through a serialization method and complete the environment configuration on the business server. Then build a scheduled task scheduler, establish an automated prediction process, and implement continuous monitoring of the prediction process. The number monitoring device obtains the yellow page number data to be monitored from the number library, uses the deployed model to perform stability prediction, and generates a stability probability value for each number. This step transforms the theoretical model into practical applications, establishes an efficient and automated prediction mechanism, and enables the model to continuously support business decisions. The successful implementation of model deployment ensures that the entire number monitoring device can operate stably in the actual environment and continuously output valuable prediction results.

[0042] S6: Based on the number stability probability value, conduct key verification on numbers higher than the preset probability threshold, and update the number library according to the verification results.

[0043] According to the number stability probability value predicted by the model, the number monitoring device sets an appropriate threshold to screen out high-risk changed numbers for key verification. The verification can be carried out by automatically dialing the number monitoring device or manually dialing, with the aim of confirming the actual status of the number. According to the verification results, the number monitoring device updates the number library information, such as removing or marking invalid numbers, and updating the verification time of stable numbers. Importantly, the verification results are also fed back to the model training link to form a mechanism for continuous optimization. By regularly updating the model with new verification data, the number monitoring device can adapt to changes in number usage patterns and maintain long-term effectiveness. This closed-loop optimization mechanism is a significant advantage of the embodiments of this application, which realizes the intelligent management and continuous improvement of the number library.

[0044] Among them, as Figure 2 shown, step S1 specifically includes:

[0045] S1.1: Extract positive samples and negative samples from the results of outbound verification data. Positive samples are numbers with successful verification and no change in the number main body, and negative samples include numbers with failed verification.

[0046] Specifically, by analyzing the verification results recorded by the outbound verification number monitoring device, the numbers are classified according to the verification status. Numbers with successful verification and no change in the number main body are marked as positive samples, representing stable numbers; while numbers with failed verification (such as no answer, number deactivated, number changed, etc.) are marked as negative samples, representing unstable numbers. These data are from the regular outbound verification operations of the number monitoring device and have a high quantity and basic coverage.

[0047] S1.2: Collect user error correction data as a supplementary data source for negative samples. User error correction data refers to the number error or change information feedback by users, and this type of data has high timeliness and accuracy.

[0048] In step S1.2, the number monitoring device collects error correction feedback data from users. This type of data is usually information actively feedback by users when they find errors or changes in the number during use, and it has high timeliness and accuracy. As a supplement to negative samples, user error correction data can timely capture numbers that have changed but have not been verified by the number monitoring device through regular verification, improving the timeliness of the samples.

[0049] S1.3: Integrate the recycled number data as another supplementary data source for negative samples. Recycled numbers refer to numbers that have been recycled by telecom operators. Such numbers are no longer valid or have changed users.

[0050] Integrate the recycled number data provided by telecom operators. This type of data refers to numbers recycled by operators, and these numbers are either no longer valid or have been reassigned to new users. As another important supplement to negative samples, recycled number data can effectively identify the situation of number invalidation caused by number recycling, further enriching the diversity of negative samples.

[0051] S1.4: Clean and integrate the collected sample data, remove duplicate data and abnormal data, and form a well-annotated training sample set to provide a data basis for subsequent model training.

[0052] In step S1.4, clean and integrate the sample data collected in the previous three steps. The number monitoring device removes duplicate data to ensure that the same number does not appear multiple times in the training set; filters abnormal data, such as samples with non-standard formats or incomplete information; and improves the annotation of all samples to ensure that each sample has a clear label (stable or unstable) and necessary attribute information. Through this step, a training sample set with reliable quality and unified structure is formed, laying a data foundation for subsequent model training.

[0053] As Figure 3 shown, step S2 specifically includes:

[0054] S2.1: Determine the preset time window range based on the verification time period of the numbers in the training sample set.

[0055] Specifically, in step S2.1, determine an appropriate time window range based on the verification time period of the numbers in the training sample set. The selection of the time window directly affects the quality of subsequent feature extraction and the accuracy of model prediction.

[0056] Generally, the number monitoring device selects key time periods according to the time cycle of number verification (such as monthly verification, quarterly verification), for example, data of the past three months. The setting of the time window needs to balance the timeliness and sufficiency of data. A too short window may lead to insufficient data and affect the stability of features; a too long window may include outdated information and reduce the accuracy of prediction.

[0057] S2.2: For each number in the training sample set, extract the call detail records within the preset time window from the communication record number monitoring device.

[0058] In step S2.2, the number monitoring device extracts the call detail records within the specified time window from the communication record number monitoring device for each number in the training sample set. These call detail records contain basic communication behavior information of the number, such as call time, call duration, call object (i.e., other numbers that communicate with this number), etc. This step requires efficient data interaction with the communication record number monitoring device, and may involve processing a large number of numbers and a huge amount of historical call records.

[0059] S2.3: Uniform the format, handle missing values, and filter outlier values of the call detail records in sequence to form preprocessed call detail records.

[0060] In step S2.3, the number monitoring device preprocesses the extracted call detail records. First is to uniform the format to ensure that the formats of all data fields are consistent, such as converting time stamps in different formats to a standard format; second is to handle missing values, and adopt appropriate filling strategies for missing data fields, such as mean filling, zero filling, or previous value filling; finally is to filter outlier values, identify and process abnormal data, such as records with extremely long or short call durations. This step ensures the quality and consistency of the data, providing a reliable data basis for subsequent feature extraction.

[0061] S2.4: Perform structured processing on the preprocessed call detail records to generate the number behavior data set.

[0062] Perform structured processing on the preprocessed call detail records to form a number behavior data set that is convenient for feature extraction. Structured processing includes organizing and classifying call details according to numbers, and constructing a data structure suitable for subsequent analysis, such as time series format or aggregated statistical format. Through this step, the original call detail records are converted into a structured number behavior data set, which is convenient for subsequent feature extraction and model training.

[0063] As Figure 4 shown, step S3 specifically includes:

[0064] S3.1: Using the said number behavior dataset, extract the information of the city where the number belongs, the operator type, and the number segment type from the said number feature dimension, and output the number attribute feature set;

[0065] Specifically, in step S3.1, the number monitoring device extracts features from the number feature dimension. Extract the information of the city where the number belongs, which reflects the geographical distribution characteristics of the number; extract the operator type information, such as China Mobile, China Unicom or China Telecom, which reflects the basic network provider of the number; extract the number segment type information, such as whether it is a virtual operator number, whether it is a special number segment, etc. These basic attribute features together constitute the basic classification characteristics of the number, forming the number attribute feature set.

[0066] S3.2: According to the said number attribute feature set, combined with the said number behavior dataset, from the call behavior dimension, count the number of calls and the number of call objects, and form the call behavior feature set.

[0067] In the embodiment of the present application, in the process of S3.2, features are extracted from the call behavior dimension. The number monitoring device calculates the call behavior indicators of each number in different time periods (such as weekly dimension, monthly dimension) and different time types (such as weekdays, rest days) based on the number behavior dataset. The main indicators include the number of calls, that is, the total number of calls initiated or received by the number within a specific time period; the number of call objects (uid number), that is, the number of different numbers that have called this number; and other possible behavior indicators, such as the average call duration, call frequency, etc. These behavior indicators together constitute the usage pattern characteristics of the number, forming the call behavior feature set.

[0068] S3.3: Using the said call behavior feature set, from the call behavior change dimension, calculate the statistical features of each behavior indicator in the call behavior feature set, and construct the call behavior change feature set, where the statistical features include at least one of the mean value or variance.

[0069] In step S3.3, the number monitoring device constructs feature variables from the call behavior change dimension.

[0070] The number monitoring device calculates the statistical features of the behavior indicators extracted in S3.2 to capture the stability and change pattern of the number usage behavior. The key statistical features include the mean value, which reflects the average level of the indicator; the variance, which reflects the degree of fluctuation of the indicator; and other possible statistics, such as trend indicators, seasonal indicators, etc. These statistical features together constitute the change characteristics of the number usage behavior, forming the call behavior change feature set.

[0071] S3.4: Perform feature fusion on the said number attribute feature set, the said call behavior feature set, and the said call behavior change feature set to generate the said number feature variable set.

[0072] In step S3.4, the feature sets generated in the previous three steps are fused to generate a complete set of number feature variables. During the feature fusion process, the number monitoring device ensures the scale consistency of each feature, such as through standardization or normalization processing; at the same time, feature selection or dimensionality reduction processing may be performed to reduce redundant features and lower the model complexity. The finally formed set of number feature variables contains comprehensive features extracted from three dimensions: number features, call behaviors, and changes in call behaviors, providing rich feature inputs for subsequent model training.

[0073] As Figure 5 shown, step S4 specifically includes:

[0074] S4.1: Divide the set of number feature variables into a training set and a test set according to a preset ratio.

[0075] Specifically, in step S4.1, the number monitoring device divides the set of number feature variables generated in S3 into a training set and a test set according to a preset ratio. When dividing the data, two key technical conditions need to be ensured: data balance and data representativeness. Data balance means that the ratio of positive and negative samples (stable numbers and unstable numbers) in the training set and the test set should be kept consistent to prevent the model from being biased towards the majority class. Considering that the number of stable numbers is usually much larger than that of unstable numbers, the number monitoring device uses the stratified sampling technique to divide according to the label ratio, ensuring that the class distribution of the original data is maintained in both the training set and the test set.

[0076] S4.2: Use the test set to evaluate the performance of the initial model for predicting number stability and calculate the performance evaluation indicators.

[0077] Data representativeness means that both the divided training set and test set can fully represent the distribution characteristics of the overall data. For this reason, during the division process, not only the stable / unstable label distribution needs to be considered, but also the distribution of other key features, such as the operator distribution and geographical distribution of the numbers. Through this comprehensive consideration data division method, it is ensured that the data set used for model training has good statistical characteristics.

[0078] In step S4.2, the number monitoring device evaluates the initialized XGBoost model.

[0079] XGBoost (eXtreme Gradient Boosting) is an ensemble learning algorithm, which is particularly suitable for dealing with classification problems of structured data and is used to construct a number stability prediction model in the embodiments of this application. In the model evaluation phase, the number monitoring device uses the test set data to predict the trained model and calculates the key performance indicators. In this embodiment, AUC (Area Under Curve) is mainly used as an indicator to evaluate the discrimination ability of the model. The AUC value ranges from 0.5 to 1, and the larger the value, the stronger the discrimination ability of the model. In addition, indicators such as accuracy, precision, and recall are also calculated to comprehensively evaluate the model performance.

[0080] S4.3: Adjust the parameters of the initial number stability prediction model according to the performance evaluation indicators to generate the number stability prediction model.

[0081] In step S4.3, the number monitoring device tunes the model parameters according to the evaluation results.

[0082] The number monitoring device uses parameter tuning techniques such as GridSearch or Bayesian Optimization to perform performance adjustment on the key parameters of the model.

[0083] The tuned parameters include but are not limited to: the maximum depth of the tree (max_depth), the learning rate (eta), the regularization parameters (lambda, alpha), the minimum number of samples in the leaf node (min_child_weight), etc.

[0084] Through cross-validation, the model is trained and its performance is evaluated multiple times using different parameter combinations to find the optimal parameter combination that can maximize the evaluation indicators. Parameter tuning is a computationally intensive process that requires balancing model performance and computational resource consumption. In the embodiments of this application, an early-stopping mechanism is adopted, that is, when the performance indicators on the validation set do not improve significantly in consecutive multiple rounds of iteration, the training process is terminated early to avoid overfitting and unnecessary waste of computational resources.

[0085] In this embodiment, the construction of the XGBoost model adopts the following specific parameter settings and architecture designs:

[0086] Model architecture:

[0087] Base learner: Decision tree (CART tree);

[0088] Maximum depth of the tree: 6 (optimal value after tuning);

[0089] Number of sub - trees: 150 trees;

[0090] Parallel computing: Enabled, using 8 threads;

[0091] Feature column sampling: 0.8 (randomly use 80% of the features each time a tree is built);

[0092] Training parameters:

[0093] Learning rate (eta): 0.05 (optimal value after tuning, a lower learning rate makes model training more stable);

[0094] Minimum number of samples in a leaf node (min_child_weight): 3 (to prevent overfitting);

[0095] Regularization parameter lambda: 1.2 (L2 regularization parameter, controlling model complexity);

[0096] Regularization parameter alpha: 0.5 (L1 regularization parameter, promoting feature sparsity);

[0097] Sample sampling ratio (subsample): 0.8 (randomly draw 80% of the samples each time a tree is built);

[0098] Sample weight balance parameter (scale_pos_weight): approximately 5.7 (the ratio of the number of negative samples to the number of positive samples);

[0099] Training data scale:

[0100] Total number of samples: approximately 10,000 numbers, of which 8,500 are stable numbers and 1,500 are unstable numbers;

[0101] Ratio of training set / test set division: 8:2;

[0102] Feature dimension: Approximately 60 key features retained after feature screening;

[0103] Training process:

[0104] Adopt k - fold cross - validation (k = 5) to ensure model stability;

[0105] Set early_stopping_rounds to 50 to avoid overfitting;

[0106] Number of training iterations: Actually converges in approximately 120 iterations;

[0107] Evaluation metrics: Mainly use AUC, while monitoring precision and recall;

[0108] Loss function: Logistic regression is used for binary classification problems;

[0109] Feature importance:

[0110] During the training process, the model automatically identified the key features and their importance scores;

[0111] The top five features include: the daily average number of queries in the first week, the variance of the weekly query numbers in the last 4 weeks, the variance of the monthly query numbers in the last 3 months, the longest call duration, and the change rate of the number of call objects;

[0112] Model performance metrics:

[0113] AUC on the test set: 0.92;

[0114] Accuracy rate: 87%;

[0115] Recall rate: 83%;

[0116] Screening accuracy rate under a high threshold setting (0.8): 95%;

[0117] Training data source:

[0118] Call result data recorded by the outbound verification device (about 60%);

[0119] User initiative error correction feedback data (about 25%);

[0120] Number recycling data provided by the operator (about 15%);

[0121] Exemplarily, the number monitoring device divides the set of number feature variables into a training set and a test set according to a preset ratio of 8:2.

[0122] Considering the imbalance in the number of stable numbers and unstable numbers in the number samples (stable numbers usually account for more than 85%), the number monitoring device uses stratified sampling technology to ensure data balance.

[0123] For example, in a set of feature variables containing 10,000 numbers, 8,500 of which are stable numbers and 1,500 are unstable numbers, the number monitoring device will divide them into a training set and a test set according to the same ratio. That is, the training set contains 6,800 stable numbers and 1,200 unstable numbers, and the test set contains 1,700 stable numbers and 300 unstable numbers. This stratified sampling method ensures that both the training set and the test set can maintain the class distribution in the original data. At the same time, the number monitoring device will also ensure that the distributions of the training set and the test set are consistent in other key features, such as the operator distribution (the ratio of mobile, Unicom, and telecom), the geographical distribution (the ratio of different cities), etc., to ensure the representativeness of the data.

[0124] Initialize the model using the XGBoost algorithm and configure the key parameters. The number monitoring device sets the evaluation target (eval_metric) to AUC, which is an indicator to measure the discrimination ability of a binary classification model and is particularly suitable for evaluating the performance of the number stability prediction model. The maximum tree depth (max_depth) of the model is initially set to 5 to control the complexity of the decision tree and avoid overfitting or underfitting. The learning rate (learning_rate) is set to a relatively conservative 0.1, enabling the model to learn gradually and avoid converging to the local optimal solution too quickly. The sample sampling ratio (subsample) is set to 0.8, meaning that 80% of the samples are randomly selected each time a tree is constructed, increasing the randomness of the model to reduce the risk of overfitting. To address the problem of sample imbalance, the number monitoring device also configures the sample weight balancing parameter (scale_pos_weight), whose value is set to the ratio of the number of negative samples to the number of positive samples, approximately 5.7 (8500 / 1500). This parameter configuration enables the model to better learn the features of the minority class (unstable numbers) and mitigate the negative impact brought by sample imbalance.

[0125] Use the configured XGBoost model to fit the training set data and perform feature screening through feature importance evaluation.

[0126] From the initially extracted 143 candidate features, use the "feature_importances_" attribute of XGBoost to screen out the feature variables that have the greatest impact on the prediction results.

[0127] For example, time series features such as "daily average number of queries in the first week", "variance of weekly query numbers in the past 4 weeks", and "variance of monthly query numbers in the past 3 months" are determined as key features.

[0128] After the number monitoring device conducts a preliminary evaluation of the model, it is found that the AUC value of the model on the test set reaches 0.85, the accuracy rate is 83%, but the recall rate is only 76%. To further optimize the model performance, the number monitoring device uses the grid search technique to tune the key parameters. After multiple rounds of tuning, the maximum tree depth is adjusted to 6, the learning rate is reduced to 0.05, the minimum number of samples in leaf nodes (min_child_weight) is set to 3, and the regularization parameter lambda is set to 1.2.

[0129] These adjustments increase the AUC value of the model to 0.92, the accuracy rate to 87%, and the recall rate to 83%. In addition, the number monitoring device also applies an early stopping mechanism, setting early_stopping_rounds to 50, that is, when the performance metrics on the validation set do not improve significantly in 50 consecutive iterations, the training process is terminated early to avoid overfitting and waste of computing resources. The finally generated number stability prediction model can achieve a screening accuracy rate of 95% under a high threshold setting (0.8), successfully balancing accuracy and computational efficiency.

[0130] In addition, the following operations can be further performed after step S4:

[0131] S4.4: Input the prediction result of the number stability prediction model into a Bayesian confidence reconfigurable flow accelerator, and perform Monte Carlo sampling through the accelerator to generate the uncertainty distribution of the prediction result.

[0132] In step S4.4, the number monitoring device inputs the prediction result of the number stability prediction model trained by XGBoost into the Bayesian confidence reconfigurable flow accelerator for further processing. Among them, this accelerator is an innovative computing architecture designed specifically for processing uncertainty estimation in real-time data streams.

[0133] When the number monitoring device generates a stability prediction probability value (such as 0.75) for a number, this single value cannot reflect the reliability of the prediction. The accelerator performs Monte Carlo sampling method, randomly perturbs and resamples the prediction result multiple times, usually performing thousands of sampling operations. For example, for a number with a prediction probability of 0.75, the number monitoring device may generate 10,000 slightly different sample points to form a probability distribution. These sampling results collectively present the "degree of certainty" of the model for the predicted value. The more concentrated the distribution, the more confident the model is in the prediction, and the more dispersed the distribution, the higher the uncertainty. In this way, the number monitoring device no longer only outputs a point estimate value, but generates a complete probability distribution, comprehensively quantifying the uncertainty of the prediction result.

[0134] In one of the embodiments,

[0135] The Bayesian confidence reconfigurable streaming accelerator is an innovative computing architecture, and its detailed technical implementation includes multiple core components and workflows. The core architecture of this accelerator consists of three main functional units: the Bayesian inference unit, the streaming data processing unit, and the dynamic resource allocation unit. The hardware implementation adopts a heterogeneous architecture of FPGA and GPU co-computation, equipped with 32GB of high-bandwidth memory and 512 reconfigurable computing cores, supporting single-precision floating-point operations. In terms of the design of the Bayesian inference unit, the implemented algorithms include Markov Chain Monte Carlo (MCMC) sampling and Hamiltonian Monte Carlo (HMC) sampling. The default configuration is 64 parallel sampling chains, with 1000 sample points per chain. The Bayesian logistic regression model is used to capture the uncertainty of XGBoost predictions, and weak-information priors are used for model parameters, specifically the normal distribution N(0,5). In terms of the design of the streaming data processing unit, the data throughput can reach the prediction results of 10,000 numbers per second. An 8-stage pipeline design is adopted, equipped with a double-buffer mechanism to ensure uninterrupted data flow, and adaptive batch size adjustment (ranging from 32 to 256) is implemented. The design of the dynamic resource allocation unit includes dynamic priority scheduling based on workload and precision requirements, real-time monitoring of the utilization rate of computing units and memory usage, dynamically adjusting the number of sampling chains and the number of samples per chain according to the complexity of prediction results, and intelligently switching between high-performance mode and low-power mode for different types of prediction tasks. In terms of the specific implementation of Monte Carlo sampling, the sample perturbation method adds random noise that conforms to the normal distribution N(0,0.01) to the prediction results of the XGBoost model. The resampling strategy slightly perturbs the feature space to generate multiple similar but slightly different sample points. 10,000 independent samplings are performed for each number prediction, and the Gelman-Rubin statistic is used to monitor the convergence of the sampling chains, usually reaching convergence after 500 - 700 samplings. In terms of the implementation of probability calibration calculation, the calibration method combines quantile calibration and temperature scaling, determines the 90% confidence interval based on the 5% and 95% quantiles of the sampling distribution. For the predicted value p, the calibrated predicted value p' and its confidence interval [p' - δ, p' + δ] are output, and the leave-one-out method is used to optimize the calibration parameters on the validation set. The dynamic weight adjustment mechanism is based on the non-linear function f(w) = a·exp(-b·w) + c of the confidence interval width w, and the parameters are determined by Bayesian optimization as a = 0.85, b = 5.2, c = 0.1. For predictions with too wide confidence intervals, the decision threshold is dynamically adjusted from 0.5 to 0.6 - 0.7, and the effect of each adjustment is recorded for continuous optimization of the adjustment function parameters. Through this complex and precise computing architecture, the Bayesian confidence reconfigurable streaming accelerator can efficiently process a large number of number prediction results, provide reliable uncertainty quantification, and provide more solid technical support for the decision-making of the number monitoring device. For example, the Bayesian confidence reconfigurable streaming accelerator is an innovative computing architecture, and its detailed technical implementation includes the following core components and workflows:

[0136] Core architecture:

[0137] The accelerator consists of three main functional units: a Bayesian inference unit, a streaming data processing unit, and a dynamic resource allocation unit;

[0138] The hardware implementation adopts a heterogeneous architecture of FPGA (Field Programmable Gate Array) and GPU collaborative computing;

[0139] Total memory capacity: 32GB high-bandwidth memory;

[0140] Computing unit: 512 reconfigurable computing cores, supporting single-precision floating-point operations;

[0141] Bayesian inference unit design:

[0142] Implementation algorithms: Markov Chain Monte Carlo (MCMC) sampling and Hamiltonian Monte Carlo (HMC) sampling;

[0143] Number of sampling chains: The default configuration is 64 parallel sampling chains;

[0144] Number of samples per chain: The default setting is 1000 sample points;

[0145] Probability model: Adopt a Bayesian logistic regression model to capture the uncertainty of XGBoost prediction;

[0146] Prior distribution: Use a weakly informative prior for model parameters, specifically the normal distribution N(0,5);

[0147] Streaming data processing unit design:

[0148] Data throughput: Can process the prediction results of up to 10,000 numbers per second;

[0149] Pipeline depth: 8-stage pipeline design;

[0150] Cache mechanism: Adopt a double-buffer design to ensure uninterrupted data flow;

[0151] Batch processing optimization: Adaptive batch size adjustment, ranging from 32 to 256;

[0152] Dynamic resource allocation unit design:

[0153] Task scheduling strategy: Dynamic priority scheduling based on workload and precision requirements;

[0154] Resource monitoring: Real-time monitoring of the utilization rate of computing units and memory usage;

[0155] Adaptive configuration: Dynamically adjust the number of sampling chains and the number of samples per chain according to the complexity of the prediction results;

[0156] Energy efficiency optimization: Intelligently switch between high-performance mode and low-power mode for different types of prediction tasks;

[0157] Specific implementation of Monte Carlo sampling:

[0158] Sample perturbation method: Add random noise that conforms to the normal distribution N(0, 0.01) to the prediction results of the XGBoost model;

[0159] Resampling strategy: Slightly perturb the feature space to generate multiple similar but slightly different sample points;

[0160] Sampling depth: Perform 10,000 independent samplings for each number prediction;

[0161] Convergence determination: Use the Gelman-Rubin statistic to monitor the convergence of the sampling chains, which usually reaches convergence after 500 - 700 samplings;

[0162] Implementation of probability calibration calculation:

[0163] Calibration method: Combine isotonic calibration and temperature scaling;

[0164] Confidence interval calculation: Determine the 90% confidence interval based on the 5% and 95% quantiles of the sampling distribution;

[0165] Mathematical representation: For the predicted value p, output the calibrated predicted value p' and its confidence interval [p' - δ, p' + δ];

[0166] Calibration parameter optimization: Use the hold-out method to optimize the calibration parameters on the validation set;

[0167] Dynamic weight adjustment mechanism:

[0168] Adjustment function: A non-linear function f(w) = a·exp(-b·w) + c based on the width w of the confidence interval;

[0169] Parameter values: Determine a = 0.85, b = 5.2, c = 0.1 through Bayesian optimization;

[0170] Threshold adjustment: For predictions with too wide confidence intervals, dynamically adjust the decision threshold from 0.5 to 0.6 - 0.7;

[0171] Feedback mechanism: Record the effect of each adjustment for continuously optimizing the adjustment function parameters;

[0172] Through this complex and precise computational architecture, the Bayesian confidence reconfigurable streaming accelerator can efficiently process a large number of number prediction results, provide reliable uncertainty quantification, and offer stronger technical support for the decision-making of the number monitoring device. In practical applications, the accelerator has increased the screening accuracy rate of unstable numbers from 87.4% to 92.1%, and the accuracy rate under high threshold settings from 95% to 97.3%, and significantly improved the robustness and reliability of the model.

[0173] S4.5: According to the uncertainty distribution, perform probability calibration calculation to form a 90% confidence interval for the prediction result.

[0174] In step S4.5, the number monitoring device performs probability calibration calculation according to the uncertainty distribution generated in S4.4 to form a 90% confidence interval for the prediction result.

[0175] Specifically, the number monitoring device first conducts statistical analysis on the distribution generated by Monte Carlo sampling, calculates its 5% and 95% quantiles, and determines an interval range that contains 90% of the sample points in the distribution.

[0176] For example, for a number with a predicted probability of 0.75, through probability calibration calculation, a result of "0.75 ± 0.15 (90% confidence interval)" may be obtained, indicating that the number monitoring device is 90% certain that the true stability probability value of this number falls between 0.6 and 0.9, which reflects a relatively high prediction uncertainty.

[0177] For another number with a predicted probability of 0.75, a result of "0.75 ± 0.05 (90% confidence interval)" may be obtained, indicating that the number monitoring device has more confidence in this prediction. The width of the confidence interval directly reflects the certainty degree of the model for the prediction result, providing a richer information basis for subsequent decision-making. This method of uncertainty quantification based on probability theory enables the number monitoring device to distinguish between "knowing" and "not knowing" situations, effectively improving the reliability of the prediction.

[0178] S4.6: Based on the 90% confidence interval value of the prediction result, use a dynamic weight adjustment mechanism to generate a calibrated and optimized prediction result, and update the calibrated and optimized prediction result into the number stability prediction model for the number stability prediction process.

[0179] The number monitoring device generates a calibrated and optimized prediction result based on the 90% confidence interval value calculated in S4.5. This mechanism dynamically adjusts the weight of the prediction result according to the width of the confidence interval.

[0180] For prediction results with a narrow confidence interval (e.g., ±0.05), the number monitoring device assigns a higher weight because the model is more confident in these predictions; for prediction results with a wide confidence interval (e.g., ±0.15), the number monitoring device assigns a lower weight and may conservatively adjust the predicted value towards the median value.

[0181] In terms of specific implementation, the number monitoring device may calculate the adjustment coefficient using a non - linear function based on the interval width. For example, a prediction with an interval width of 0.1 may maintain the original value, while a prediction with an interval width of 0.3 may be adjusted towards the median value by 15%. The calibrated prediction results are updated into the number stability prediction model, replacing or supplementing the original predicted values, and are used for subsequent number stability prediction processes. Through this calibration mechanism, the screening accuracy rate of unstable numbers has increased from the original 87.4% to 92.1%, the accuracy rate under the high - threshold setting has increased from 95% to 97.3%, and the number of identified unstable numbers has increased by approximately 15%, significantly improving the prediction performance of the number monitoring device.

[0182] The reconfigurable streaming accelerator with Bayesian confidence is an innovative computing architecture designed specifically for handling uncertainty estimation in real - time data streams. In the scenario of number stability prediction, simply obtaining a probability value is not enough; it is also necessary to know the credibility of this probability value, that is, how "confident" the model is in its own prediction results. This accelerator uses the Monte Carlo sampling method to perform multiple random samplings on the prediction results, generating a probability distribution of the prediction results, thereby quantifying the uncertainty of the prediction.

[0183] Based on this uncertainty quantification, the number monitoring device performs probability calibration calculations to form a 90% confidence interval for the prediction results, that is, there is a 90% certainty that the true result will fall within this interval. The number monitoring device dynamically adjusts the weight of the prediction results according to the width of the confidence interval, giving a higher weight to prediction results with a narrow confidence interval (low uncertainty) and a lower weight to prediction results with a wide confidence interval (high uncertainty), thereby generating calibrated and optimized prediction results.

[0184] Step S5 specifically includes:

[0185] S5.1: Using the serialization method, convert the number stability prediction model into a model file in a persistent format and complete the environment configuration on the business server.

[0186] Using the serialization method, convert the trained and calibrated number stability prediction model into a model file in a persistent format and complete the environment configuration on the business server.

[0187] In terms of specific implementation, the number monitoring device uses Python's pickle module or the save_model method built into XGBoost to convert the model object into a binary file or a JSON format file for easy storage and subsequent loading. At the same time, the number monitoring device also saves the metadata information related to the model, such as the feature list, normalization parameters, probability calibration parameters, etc., to ensure the integrity of model deployment. In terms of the business server environment configuration, the number monitoring device installs the necessary dependency libraries (such as XGBoost, NumPy, Pandas, etc.), configures appropriate computing resources (such as the number of CPU cores, memory size), sets network access permissions, and tests the model loading and inference functions to ensure that the model can run normally in the production environment. This step is a key link in transforming the theoretical model into an actually usable number monitoring device, which determines whether the model can play a stable role in actual business.

[0188] S5.2: Utilize the configured business server environment to construct a scheduled task scheduler and establish an automated prediction process.

[0189] The number monitoring device usually uses tools such as Cronjobs or Airflow to set prediction tasks that are executed regularly, such as automatically triggering the prediction process at 2 am every day or 8 am every Monday morning.

[0190] The automated process includes multiple links: First is data extraction, obtaining the required data from the number database and the communication record number monitoring device; then is feature calculation, generating the input features required by the model according to the predefined feature engineering process; then is model prediction, loading the persisted model file to execute the prediction calculation; finally is result storage, writing the predicted number stability probability value into the database or the result file. The number monitoring device also configures the dependency relationships between tasks, such as ensuring that the feature calculation is only executed after the data extraction is completed, and a task retry mechanism to handle possible execution failures. This automated scheduling mechanism reduces manual intervention, ensures the regular execution of prediction tasks, and provides technical support for the continuous monitoring of the number database.

[0191] S5.3: According to the automated prediction process, implement a prediction process monitoring mechanism to continuously record and output the prediction status and result distribution data.

[0192] Specifically, the number monitoring device sets up multi-level monitoring points to track the execution status of the prediction task in real time, including: task start time, data loading progress, feature calculation time consumption, model prediction speed, result storage status, etc.

[0193] Meanwhile, the number monitoring device also records the statistical distribution of the prediction results, such as the number of numbers in each probability interval, the proportion of high-risk numbers, the risk distribution of numbers of different operators, etc. These monitoring data are presented to the management personnel through the log number monitoring device, the data panel or the alarm mechanism. In particular, the number monitoring device sets up an anomaly detection mechanism, which automatically triggers an alarm notification when the prediction result distribution shows a significant anomaly (such as a sudden substantial increase in high-risk numbers).

[0194] In addition, the number monitoring device also records the resource consumption situation, such as CPU usage rate, memory occupancy, prediction latency, etc., in order to optimize the performance of the number monitoring device. This comprehensive monitoring mechanism ensures the transparency and visibility of the prediction process, helps to detect and solve potential problems in a timely manner, and guarantees the stable operation of the number monitoring device.

[0195] S5.4: Obtain the yellow page number data in the number library, and use the number stability prediction model to predict the stability of the yellow page number data, and generate a number stability probability value.

[0196] The number monitoring device processes the number library in a batch polling manner, divides the numbers in the library into multiple batches, and processes one batch per day to ensure that all numbers are covered within a complete cycle (such as a month).

[0197] For the retrieved numbers, the number monitoring device first extracts the relevant feature variables, which are consistent with the feature definitions used during model training, including obtaining recent call detail records from the communication record number monitoring device and calculating the feature variables in three dimensions: number features, call behaviors, and changes in call behaviors. The feature calculation process is highly automated to support large-scale data processing. The number monitoring device pays special attention to the consistency of feature processing and missing value processing to ensure consistency with the processing methods in the training stage. Model prediction uses vectorized operations to process multiple numbers at once, improving the calculation efficiency. Finally, the number monitoring device generates a stability probability value between 0 and 1 for each number. The closer the value is to 1, the more likely the number is in an unstable state and the higher the change risk; the closer the value is to 0, the more likely the number is in a stable state and the lower the change risk. These probability values provide a scientific basis for subsequent differential verification.

[0198] The number monitoring device needs to serialize the trained XGBoost model into a persistent format, such as a pickle file or JSON format, for easy storage and loading. Then, configure the necessary dependency environment on the production server to ensure the normal operation of the model. Next, the number monitoring device establishes an automated mechanism for the prediction process, sets up a scheduled task scheduler (such as Cronjobs or Airflow), and automatically triggers the prediction task at a predetermined time interval (such as daily or weekly). The number monitoring device also establishes a monitoring mechanism to monitor the model operation status and prediction result distribution in real time and detect abnormal situations in a timely manner. The number monitoring device obtains the yellow page number data to be monitored from the number library and uses the deployed model to make a stability prediction, generating a stability probability value for each number.

[0199] In addition, after step S5, the following operations can be further performed:

[0200] S5.5: Adopt a large language model based on the active inference framework to deeply analyze the number stability probability value generated by the number stability prediction model and generate an analysis result at the semantic level.

[0201] The number monitoring device adopts a large language model (LLM) based on the active inference framework to deeply analyze the probability value generated by the number stability prediction model. This step establishes a collaborative decision-making framework between the XGBoost model and the large language model.

[0202] When the prediction results of the XGBoost model for certain numbers are in the uncertain interval (such as the probability value is between 0.4 and 0.6) or abnormal patterns appear, the active inference framework will trigger the large language model to conduct a deeper analysis. The large language model generates multi-angle inference chains by retrieving historical similar cases, analyzing the temporal changes of number usage patterns, and integrating specific rules in the industry knowledge base.

[0203] For example, for a merchant number in the catering industry, the LLM may find that the recent call pattern of this number is abnormal and there are discussions on social media about the possible relocation of the restaurant, so it judges that the change risk is higher than the predicted value of the XGBoost model. In addition, the LLM can understand and utilize richer semantic information, such as analyzing the voice text content in the outbound call record to identify semantic clues of number change (such as "This number is no longer XX company's"). Through this in-depth semantic analysis, the number monitoring device generates an analysis result at the semantic level with rich explanatory content, providing supplementary basis or correction suggestions for the prediction decision.

[0204] The large language model of the active inference framework is a deep neural network based on the Transformer architecture, which is specifically adapted and trained for the number monitoring field. The model's basic architecture uses a pre-trained general large language model as the foundation, specifically an autoregressive Transformer model with 175B parameters, including 48 layers of Transformer encoders, 96 attention heads, 128 dimensions per attention head, a word embedding dimension of 12,288, and a context window length of 4,096 tokens. The activation function uses GELU (Gaussian Error Linear Unit), and the layer normalization uses the Pre-LayerNorm structure to improve training stability. In terms of domain adaptation fine-tuning, the parameter-efficient fine-tuning technique LoRA (Low-Rank Adaptation) is used, with the LoRA rank set to 16, and only about 0.1% of the model's parameters are updated. The scale of the fine-tuning data is about 50,000 labeled number analysis cases, and the data sources include historical number verification records, expert analysis reports, and a knowledge base in the telecommunications field. The learning rate is set to 2e-5, and the cosine decay strategy is adopted. The training batch size is 32, and 3 full dataset iterations are performed. The training hardware is 8 × A100 80GB GPUs, and mixed-precision training is used. The active inference framework design adopts the Chain-of-Thought reasoning mechanism, with the number of reasoning steps being 5 - 8 steps, dynamically adjusted according to the complexity of the problem. The reasoning control is based on the selection of reasoning paths through self-consistency verification. Knowledge integration combines external structured knowledge bases with internal parameter knowledge, and the output formatting uses JSON structured output to ensure compatibility with subsequent devices. The integration mechanism with XGBoost adopts a cascaded integration, with XGBoost as the primary predictor and the large language model as the secondary analyzer. The trigger condition is to trigger the large language model analysis when the XGBoost prediction probability is in the interval [0.4, 0.6] or the prediction result significantly deviates from the historical pattern. The XGBoost prediction result, key features, and their importance scores are passed through a structured API call. The decision rule integrates the judgments of the two models based on a weighted voting mechanism, and the weights are dynamically adjusted according to the specific scenario. In solving the "black box" problem, the model not only provides the final judgment but also outputs the complete reasoning chain, detailing every step of the reasoning process from the original data to the conclusion. For example, for a number judged to be of high risk, the model may provide the following reasoning chain: "The outgoing call volume of this number has decreased by 50% in the past 3 weeks (abnormal behavior) → The merchant type is small catering (industry with a high change rate) → The keyword 'change of boss' appears in the outgoing call records (direct change clue) → There is a demolition record in the business district where it is located (environmental risk factor) → Comprehensive judgment: The change risk is high". This transparent reasoning process makes the model's judgment auditable and verifiable, effectively reducing the opacity of AI decision-making.

[0205] The collaborative decision-making framework of the XGBoost model and the large language model is an innovative technical architecture that fully integrates the advantages of two different types of models to form a complementary and collaborative intelligent decision-making device. In terms of architecture design, the framework adopts a "dual-core + multi-level interaction" structure. The XGBoost model, as the first core, is responsible for the efficient quantitative analysis based on structured feature data and is good at extracting stable statistical patterns from historical call records; the large language model, as the second core, is responsible for semantic understanding, knowledge reasoning, and complex scenario judgment, and is good at processing unstructured data and identifying rare patterns. The two core models are connected through multi-level interaction interfaces, including a feature sharing layer, a decision fusion layer, and a feedback optimization layer, to ensure that information can flow bidirectionally and be effectively integrated.

[0206] Specifically, the collaborative decision-making framework of XGBoost and large language models is an innovative technical architecture, and its detailed design adopts a "dual-core + multi-level interaction" structure. The XGBoost model serves as the first core, responsible for efficient quantitative analysis based on structured feature data, and is good at extracting stable statistical patterns from historical call records; the large language model serves as the second core, responsible for semantic understanding, knowledge reasoning, and complex scenario judgment, and is good at processing unstructured data and identifying rare patterns. The two core models are connected through multi-level interaction interfaces, including a feature sharing layer, a decision fusion layer, and a feedback optimization layer, ensuring that information can flow bidirectionally and be effectively integrated. The working principle of this framework is based on a "hierarchical processing + collaborative decision-making" mechanism. First, all number data are initially predicted by the XGBoost model to obtain a basic stability probability value. Then, the device filters out complex cases that need further analysis according to preset rules (such as the prediction result is in the uncertain interval of 0.4 - 0.6, or the prediction result deviates significantly from the historical pattern). For these cases, the device triggers the large language model for in-depth analysis. The large language model receives the structured feature data of the number, the prediction result of XGBoost, and additional unstructured information (such as outbound call record text, merchant information, etc.), and generates supplementary judgments and explanations through multi-step reasoning. Finally, the device integrates the judgment results of the two models through weighted fusion or rule switching to form a final decision. To ensure the controllability and consistency of the output of the large language model, the device interacts with the model through a defined function call interface (Function Calling), defines multiple dedicated functions such as analyzeNumberPattern, retrieveSimilarCases, assessIndustryRisk and other core functions, and ensures the structuring and consistency of input and output through parameter passing in JSON format. This multi-objective optimization strategy of the Framework model framework collects multi-dimensional data including prediction accuracy, recognition coverage rate, and resource consumption efficiency, and establishes a comprehensive optimization index system. The device uses this index system to calculate the gradient direction of each optimization target, constructs a Pareto optimal solution set through mathematical optimization methods, and determines the optimal balance point according to business requirements and resource constraints. Based on this balance point, the device generates updated prediction weight parameters (adjusting the weights of the XGBoost and large language models in decision-making) and threshold parameters for different business scenarios, achieving precise allocation of resources and improvement in the accuracy of number stability prediction.

[0207] The working principle of this framework is based on a "hierarchical processing + collaborative decision-making" mechanism.

[0208] First, all number data are initially predicted by the XGBoost model to obtain the basic stability probability values. Then, the device filters out complex cases that need further analysis according to preset rules (such as the prediction results being in the uncertain interval of 0.4 - 0.6, or the prediction results deviating significantly from the historical patterns). For these cases, the device triggers the large language model for in-depth analysis. The large language model receives the structured feature data of the number, the prediction results of XGBoost, and additional unstructured information (such as outbound call record text, merchant information, etc.), and generates supplementary judgments and explanations through multi-step reasoning. Finally, the device integrates the judgment results of the two models through weighted fusion or rule switching to form the final decision.

[0209] The collaboration mechanism is the core of this framework, which is specifically reflected in three levels: information complementarity, decision fusion, and learning optimization.

[0210] At the information complementarity level, XGBoost provides accurate quantitative analysis results, while the large language model provides semantic understanding and background knowledge; at the decision fusion level, the device dynamically adjusts the weights of the two models according to different scenarios. For example, for regular numbers with rich historical data, it is more inclined to trust the judgment of XGBoost, while for special numbers with sparse data or abnormal behaviors, it attaches more importance to the analysis of the large language model; at the learning optimization level, the two models provide learning feedback to each other. For example, the quantitative features of XGBoost can guide the large language model to focus on key factors, and the new patterns identified by the large language model can be transformed into new features of XGBoost, forming a mutually promoting evolution mechanism.

[0211] To ensure the technical feasibility and transparency of this collaboration framework, a specific technical implementation path is adopted for the application of the large language model. First, the basic large language model (such as deepseek or GPT-4o) is fine-tuned for domain adaptation and further trained using a dataset related to number monitoring. The fine-tuning process adopts parameter-efficient fine-tuning techniques (such as LoRA, P-Tuning), which only update a small number of key parameters in the model, strengthening specific domain knowledge while maintaining the basic capabilities. The fine-tuning data includes labeled number cases, expert analysis reports, and relevant industry knowledge, and guides the model to learn how to analyze number stability problems through instruction fine-tuning.

[0212] In practical applications, the large language model interacts with the device through a defined function calling interface (FunctionCalling).

[0213] For example, the device defines multiple dedicated functions such as analyzeNumberPattern(numberFeatures, historyData) for analyzing number patterns, retrieveSimilarCases (numberInfo, topK) for retrieving similar historical cases, assessIndustryRisk (merchantInfo, industryType) for assessing industry risks, etc. When the large language model is needed for analysis, the device passes relevant data through these structured interfaces, and the model generates a JSON response in a standard format, including analysis conclusions, confidence scores, and reasoning steps. This FunctionCalling mechanism ensures the controllability and consistency of the output of the large language model, avoiding the uncertainties that may be brought about by free text generation.

[0214] In addition, to reduce the "black box" risk, the device implements an interpretability mechanism. When generating analysis results, the large language model not only provides the final judgment but also outputs a complete Chain-of-Thought, detailing every step of the reasoning process from the original data to the conclusion.

[0215] For example, for a number judged to be of high risk, the model may provide the following reasoning chain: "The outgoing call volume of this number has decreased by 50% in the past 3 weeks (abnormal behavior) → The merchant type is small catering (high change rate industry) → The keyword 'change of boss' appears in the outgoing call records (direct change clue) → There are demolition records in the business district where it is located (environmental risk factor) → Comprehensive judgment: The change risk is high". This transparent reasoning process makes the model judgment auditable and verifiable, significantly reducing the opacity of AI decision-making. At the same time, it also provides professional knowledge support for human operators, enabling them to understand and verify the rationality of the device's judgment.

[0216] In a specific example, the large language model of the active inference framework adopted in this embodiment is a deep neural network based on the Transformer architecture, which is specifically adapted and trained for the number monitoring field. It includes:

[0217] Model basic architecture:

[0218] Basic model: A pre-trained general large language model is used as the basis, specifically an autoregressive Transformer model with a parameter scale of 175B;

[0219] Number of hidden layers: 48-layer Transformer encoder;

[0220] Number of attention heads: 96 attention heads;

[0221] Dimension of each attention head: 128;

[0222] Word embedding dimension: 12,288;

[0223] Context window length: 4,096 tokens;

[0224] Activation function: GELU (Gaussian Error Linear Unit);

[0225] Layer normalization: Using the Pre-LayerNorm structure to improve training stability;

[0226] Domain adaptation fine-tuning:

[0227] Fine-tuning method: Parameter-efficient fine-tuning technique LoRA (Low-Rank Adaptation);

[0228] LoRA rank: 16;

[0229] Proportion of fine-tuned parameters: Only update about 0.1% of the model parameters;

[0230] Scale of fine-tuning data: Approximately 50,000 labeled number analysis cases;

[0231] Data sources: Historical number verification records, expert analysis reports, telecommunications domain knowledge bases;

[0232] Learning rate: 2e-5, using the cosine decay strategy;

[0233] Training batch size: 32;

[0234] Number of training epochs: 3 full dataset iterations;

[0235] Training hardware: 8 × A100 80GB GPUs, mixed-precision training;

[0236] Active inference framework design:

[0237] Inference chain structure: Using the Chain-of-Thought inference mechanism;

[0238] Number of inference steps: 5 - 8 steps, dynamically adjusted according to the problem complexity;

[0239] Inference control: Inference path selection based on self-consistency verification;

[0240] Knowledge integration: Combining external structured knowledge bases and internal parameter knowledge;

[0241] Output formatting: JSON structured output to ensure compatibility with subsequent devices;

[0242] Model and XGBoost integration mechanism:

[0243] Integrated architecture: Cascade integration is adopted, with XGBoost as the primary predictor and the large language model as the secondary analyzer; Trigger condition: When the prediction probability of XGBoost is in the interval [0.4, 0.6] or the prediction result deviates significantly from the historical pattern, the large language model analysis is triggered; Information transmission: The XGBoost prediction result, key features, and their importance scores are transmitted through structured API calls; Decision rule: Based on the weighted voting mechanism, the judgments of the two models are integrated, and the weights are dynamically adjusted according to the specific scenario.

[0244] FunctionCalling implementation:

[0245] Define a function set: 8 core functions such as analyzeNumberPattern, retrieveSimilarCases, assessIndustryRisk, etc.;

[0246] Calling method: Parameter passing in JSON format to ensure the structuring and consistency of input and output;

[0247] Execution control: Support for chained function calls to implement complex analysis processes;

[0248] Error handling: Set a timeout mechanism and exception capture to ensure the stability of the device;

[0249] Model evaluation metrics:

[0250] Explanation accuracy: 94.7% (the proportion of explanations provided by the model that are consistent with human experts);

[0251] Decision assistance ability: In boundary cases where XGBoost is difficult to judge, the decision accuracy is increased by 17.3 percentage points;

[0252] Inference latency: The average time taken to analyze each number is 0.8 seconds;

[0253] Resource consumption: The peak memory usage is 16GB, supporting batch processing mode;

[0254] S5.6: Combining the analysis results at the semantic level, adopt a multi-objective optimization strategy to output optimized prediction parameters and optimized verification strategies.

[0255] The number monitoring device first constructs a multi-dimensional optimization index system that includes dimensions such as prediction accuracy rate, recognition coverage rate, and resource consumption efficiency. This multi-objective perspective breaks the traditional single-accuracy optimization idea, enabling the number monitoring device to make more balanced decisions in complex business environments. Due to conflicts among different optimization objectives (such as increasing accuracy may reduce coverage or increase resource consumption), the number monitoring device calculates the gradient directions of each optimization objective and constructs a Pareto optimal solution set. The Pareto optimal solution refers to the set of solutions where it is impossible to improve all other objectives simultaneously without harming at least one objective.

[0256] Through this multi-objective optimization, the number monitoring device can increase the overall working efficiency by approximately 35% with the same resource input. The number monitoring device combines business requirements and resource constraints, selects the best balance point from the Pareto front, generates optimized prediction weight parameters (adjusting the weights of XGBoost and LLM in decision-making) and threshold parameters for different scenarios, and formulates differentiated verification strategies for different types of numbers to achieve precise resource allocation.

[0257] S5.7: Based on the optimized prediction parameters and the optimized verification strategy, iteratively optimize the number stability prediction model to improve the accuracy of number stability prediction.

[0258] This step realizes a closed-loop from analysis and decision-making to model optimization. The number monitoring device applies the parameters generated by multi-objective optimization to the prediction process and adjusts the decision-making mechanism of the model.

[0259] For example, for numbers in industries identified by LLM as high-risk industries (such as industries with frequent store replacements like catering and retail), the number monitoring device will automatically lower the unstable determination threshold to increase monitoring sensitivity; while for numbers in stable industries (such as large enterprises), the threshold will be increased to reduce unnecessary verification.

[0260] The number monitoring device also establishes a feedback and iteration mechanism, feeds the verification results back to the model training session, continuously updates the training samples, optimizes the feature weights of the XGBoost model, and enriches the case knowledge base of the LLM. Through continuous learning and adaptation, the number monitoring device can continuously improve itself as the business environment changes.

[0261] This iterative optimization mechanism increases the prediction accuracy rate from 95% of the original plan to 98.2%. It performs particularly well in dealing with boundary cases and atypical patterns. At the same time, it improves the average recognition lead time from 7 days to 15 days, leaving more sufficient time for business adjustment. This closed-loop optimization reflects the adaptive learning ability of the number monitoring device, enabling the number stability prediction to continuously improve and adapt to changing business requirements.

[0262] The large language model of the active inference framework is an artificial intelligence number monitoring device capable of performing in-depth semantic analysis and inference. In the number monitoring scenario, when the prediction results of the XGBoost model for certain numbers are in the uncertain interval (such as the probability value is between 0.4 - 0.6) or abnormal patterns occur, the active inference framework will trigger the large language model to conduct deeper analysis. The large language model generates multi-angle inference chains by retrieving similar case histories, analyzing the temporal changes in number usage patterns, and integrating specific rules in the industry knowledge base (such as the number update cycle characteristics of specific industries), providing supplementary evidence or correction suggestions for the predictions of XGBoost.

[0263] Specifically, the number monitoring device first obtains the numbers with probability values in a specific interval (such as 0.4 - 0.6), performs in-depth feature analysis on these numbers to form preliminary analysis results. Then, the number monitoring device retrieves the historical case database, finds historical cases similar to the current number's characteristics, and analyzes the temporal changes in number usage patterns to generate number change risk characteristics. Based on the active inference framework, the number monitoring device integrates all analysis results, and at the same time considers the outbound call record text and merchant status information to comprehensively evaluate the change risk of the number.

[0264] The number monitoring device also adopts a multi-objective optimization strategy, considering multiple objectives such as prediction accuracy, recognition coverage, and resource consumption efficiency at the same time. Through the Pareto optimality principle, the number monitoring device finds the best balance among these potentially conflicting objectives, generating optimized prediction parameters and verification strategies. These optimization results are applied to the number stability prediction process to continuously improve the model performance.

[0265] Among them, in S5.5, a large language model based on the active inference framework is used to conduct in-depth analysis on the number stability probability value generated by the number stability prediction model, generating analysis results at the semantic level, including:

[0266] A1: Obtain the numbers with the number stability probability value within the preset probability interval, and perform in-depth feature analysis on this number to form preliminary analysis results.

[0267] First, screen out the numbers with the number stability probability value in a specific uncertain interval (such as between 0.4 and 0.6). These numbers are boundary cases that are difficult for the XGBoost model to determine and need further in-depth analysis.

[0268] For these numbers, the number monitoring device performs in-depth feature analysis, including a detailed analysis of number attributes, historical call behaviors, and change patterns. The analysis process not only considers existing feature variables but also further explores potential associated features, such as the time distribution feature of number usage and the group feature of interaction objects. Through this in-depth feature exploration, the number monitoring device forms a preliminary analysis result for each boundary number, laying a foundation for subsequent historical case matching and semantic analysis. This step reflects the special attention of the number monitoring device to boundary situations that are difficult for the model to judge and improves the overall prediction accuracy.

[0269] A2: Using the preliminary analysis result, retrieve the historical case database and analyze the temporal changes in the number usage pattern to generate number change risk features.

[0270] The number monitoring device uses the preliminary analysis result to retrieve similar historical cases in the historical case database. These historical cases contain number samples with known results in the past and their complete feature change trajectories. Through the case matching algorithm, the number monitoring device finds the historical case most similar to the currently analyzed number and analyzes the change results and reasons for the number stability in these cases.

[0271] At the same time, the number monitoring device performs a temporal analysis of the usage pattern of the current number to detect the change trend of behavioral features over time, such as sudden changes in call frequency and changes in the composition of call objects. This analysis method based on historical experience and temporal changes can reveal dynamic risk signals that may be ignored by static features and generate more comprehensive number change risk features. This step makes full use of historical experience and time dimension information, enhancing the perception ability of number change risks.

[0272] A3: Based on the active inference framework, fuse the number change risk features, and at the same time extract and process the outbound call record text and merchant status information related to the number, and output a comprehensive evaluation result of the number change risk as the analysis result at the semantic level.

[0273] Based on the active inference framework, perform intelligent fusion of the number change risk features generated in the previous two steps.

[0274] The active inference framework evaluates the risk status of the number from multiple perspectives and dimensions by simulating the thinking process of human experts. Particularly importantly, in this step, the number monitoring device also extracts and processes the outbound call record text related to the number, such as the voice-to-text content in the call record, and identifies semantic clues that may indicate a number change (such as "this number is no longer XX company's").

[0275] Meanwhile, the number monitoring device analyzes the merchant status information associated with the number, such as the business operation of the merchant, online reviews, news reports, etc., and discovers external factors that may affect the stability of the number. By comprehensively considering this multi-dimensional information, the number monitoring device outputs a comprehensive risk assessment result with rich semantic explanations. This step introduces semantic understanding and knowledge reasoning capabilities, significantly enhancing the analysis depth and prediction ability of the number monitoring device.

[0276] It should be noted that in S5.6, a multi-objective optimization strategy is adopted to output optimized prediction parameters and optimized verification strategies, including:

[0277] B1. Collect data on prediction accuracy, recognition coverage rate, and resource consumption efficiency, and establish a multi-dimensional optimization index system.

[0278] A comprehensive multi-dimensional optimization index system is constructed, no longer only focusing on a single prediction accuracy index.

[0279] Specifically, the number monitoring device simultaneously collects three types of key data: prediction accuracy data, which reflects the correctness of the model's judgment; recognition coverage rate data, which represents the proportion of actually unstable numbers that can be recognized; and resource consumption efficiency data, which measures the verification resource consumption required to achieve a specific accuracy and coverage rate.

[0280] In addition, the number monitoring device may also consider other auxiliary indicators, such as prediction interpretability, prediction lead time, etc. Through the unified collection and standardized processing of these multi-dimensional indicators, the number monitoring device establishes an index system that can comprehensively evaluate the performance of the model and strategy. This multi-dimensional perspective enables the number monitoring device to make more balanced and comprehensive optimization decisions in a complex business environment, rather than simply pursuing the maximization of a single indicator.

[0281] B2. Utilize the multi-dimensional optimization index system to calculate the gradient directions of each optimization objective, construct a Pareto optimal solution set, and determine the optimal balance point.

[0282] Based on the established multi-dimensional index system, Pareto optimality analysis is carried out. The number monitoring device calculates the gradient direction of each optimization objective, that is, determines the influence direction and degree of changing a certain parameter on each index. Since there may be conflicts between different optimization objectives (for example, improving accuracy may reduce coverage), the number monitoring device constructs a Pareto optimal solution set through mathematical optimization methods. The Pareto optimal solution refers to the set of solutions that cannot improve all other objectives simultaneously without damaging at least one objective. On this basis, the number monitoring device combines business requirements and resource constraints to determine the optimal balance point, that is, selects the most suitable solution for the current business scenario on the Pareto frontier. This decision-making method based on multi-objective optimization theory enables the number monitoring device to find a scientific and reasonable balance among multiple conflicting objectives, avoiding the one-sidedness problem that may be brought about by single-index optimization.

[0283] B3. According to the optimal balance point, generate updated prediction weights and threshold parameters, and apply the updated parameters to the number stability prediction process.

[0284] Generate specific operation parameters and strategies according to the determined optimal balance point.

[0285] There are mainly two types of parameters: one is the prediction weight parameter, which adjusts the weight distribution of the XGBoost model and the large language model in the final decision-making, and different weight combinations may be adopted for different types of numbers; the other is the threshold parameter, which sets different judgment thresholds for different business scenarios and number types. For example, a lower threshold may be adopted for high-risk industries to improve sensitivity. These updated parameters are directly applied to the number stability prediction process, affecting subsequent judgment results and verification strategies. The number monitoring device also establishes a real-time monitoring mechanism for the parameter effects, continuously evaluates the effects after parameter adjustment, and provides feedback for the next round of optimization. This step transforms the abstract optimization theory into specific execution parameters, realizes the transition from theoretical analysis to practical application, and enables the achievements of multi-objective optimization to effectively improve the actual performance of the number monitoring device.

[0286] Step S6 specifically includes:

[0287] S6.1: Based on the number stability probability value, set an appropriate unstable probability threshold to screen out high-risk changed numbers.

[0288] In step S6.1, the number monitoring device sets an appropriate unstable probability threshold according to business requirements and resource conditions. By analyzing the ROC curve (Receiver Operating Characteristic Curve) and the PR curve (Precision-Recall Curve), the number monitoring device determines the model performance at different thresholds.

[0289] For example, when the probability threshold is selected as 0.5, the screening accuracy rate of unstable numbers can reach 87.4%; when the threshold is increased to 0.8, the screening accuracy rate can be increased to 95%, but the number of identified unstable numbers will decrease accordingly. The selection of the threshold needs to balance the accuracy rate and the coverage rate, and consider the available verification resources.

[0290] S6.2: Conduct key verification on the screened high-risk numbers to confirm their actual stable status.

[0291] In step S6.2, the number monitoring device conducts key verification on the screened high-risk numbers. The verification can be carried out through the automatic outbound call number monitoring device for basic validity verification, or through manual dialing for in-depth information confirmation. The purpose of the verification is to confirm the actual status of the numbers, including whether they are in use, whether the affiliated entity has changed, and whether they are still associated with the original recorded merchant, etc.

[0292] S6.3: Update the number library information according to the verification results, and feedback the verification results to the model training link to achieve continuous optimization of the model.

[0293] In step S6.3, the number monitoring device updates the number library information according to the verification results. For the numbers confirmed to have changed, the number monitoring device removes them from the library or marks them as invalid; for the numbers confirmed to be stable, the number monitoring device updates their last verification time and adjusts their priority in subsequent monitoring.

[0294] In addition, the verification results are also feedback to the model training link to form a closed-loop optimization mechanism. The number monitoring device compares the verification results with the model prediction results, generates new marked data, and uses these new data to update the model regularly (such as monthly or quarterly) to improve the prediction accuracy of the model. This continuous learning mechanism enables the model to adapt to changes in number usage patterns and newly emerging features, maintaining long-term effectiveness.

[0295] The embodiment of the present application further provides a number monitoring device, and the device includes:

[0296] A sample collection module, configured to collect positive and negative sample data in the outbound call verification data results, user correction data, and recycled number data, and generate a training sample set;

[0297] A data processing module, configured to extract the historical call detail data of each number from the communication record device storing the historical call records of the numbers based on each number included in the training sample set, and preprocess the historical call detail data to generate a number behavior data set;

[0298] A feature extraction module, configured to extract and process features from the number behavior data set in three dimensions of number features, call behavior, and call behavior changes, and generate a set of number feature variables;

[0299] A model training module, configured to construct an initial model for predicting number stability based on the set of number feature variables through the XGBoost algorithm, and evaluate and optimize the parameters of the prediction results of the initial model for predicting number stability, and generate a number stability prediction model;

[0300] A prediction module, configured to deploy the number stability prediction model to a business scenario, obtain yellow page number data in a number library, and use the number stability prediction model to predict the stability of the yellow page number data, and generate a number stability probability value;

[0301] A verification and update module, configured to conduct key verification on numbers higher than a preset probability threshold based on the number stability probability value, and update the number library according to the verification results.

[0302] An embodiment of the present application further provides a computer device, where the computer device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method of the above number monitoring method.

[0303] An embodiment of the present application further provides a computer-readable storage medium, which stores computer instructions for causing a computer to execute the method of the above number monitoring method. An embodiment of the present application further provides a computer program product, including computer instructions, which implement the steps of the above number monitoring method when executed by a processor.

[0304] The present application has the following technical effects:

[0305] Efficient resource utilization: Through a differential verification strategy based on model prediction, the number monitoring device only conducts key verification on high-risk numbers, greatly reducing the resource waste caused by indiscriminate verification and improving the resource utilization efficiency. Experimental results show that compared with traditional full-scale verification, the present number monitoring device can save about 70% of verification resources while maintaining comparable data accuracy.

[0306] Improve prediction accuracy: By comprehensively extracting features from three dimensions of number features, call behavior, and call behavior changes, and combining the powerful learning ability of the XGBoost algorithm, the number monitoring device can accurately identify the change risk of numbers. Under a high threshold setting, the number monitoring device can achieve a screening accuracy rate of 95%, far exceeding the accuracy of traditional methods.

[0307] Intelligent management: The number monitoring device constructs a complete closed-loop for number library maintenance. Through a cyclic mechanism of model prediction, key verification, library update, and feedback optimization, it realizes the automated and intelligent management of number library maintenance. In particular, by introducing the Bayesian confidence accelerator and the active inference framework of large language models, the number monitoring device can continuously learn and optimize to adapt to the changing business environment.

[0308] Early warning ability: By deeply analyzing the number behavior patterns, the number monitoring device can identify numbers with high change risks before the actual number changes, providing an early warning for business adjustment. Compared with the traditional ex-post verification method, this number monitoring device has increased the average identification lead time from 0 days to 7 - 15 days, leaving sufficient time for business adjustment.

[0309] In summary, the number monitoring method and device provided in this application, by deeply mining the number behavior characteristics, combined with advanced machine learning technologies and intelligent optimization strategies, realize the efficient and accurate maintenance of the commercial number library, solve the technical problems such as resource waste, lagging update, and low accuracy in traditional methods, greatly improve the maintenance efficiency and quality of the number library, and provide reliable data guarantee for the application of the commercial number library.

[0310] The embodiments of the present disclosure also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of a number monitoring method described in the above method embodiments. Among them, the storage medium can be a volatile or non-volatile computer-readable storage medium.

[0311] In addition, the embodiments of the present disclosure also provide a computer program product, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of a number monitoring method provided in any one of the above embodiments of the present disclosure. For details, please refer to the above method embodiments and will not be elaborated here.

[0312] Among them, the above computer program product can be specifically implemented in a manner of hardware, software, or a combination thereof. In an optional embodiment, the computer program product is specifically embodied as a computer storage medium, which can be a volatile or non-volatile computer-readable storage medium. In another optional embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.

[0313] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices and apparatuses can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein. In several embodiments provided by the present disclosure, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some communication interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0314] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0315] In addition, in each embodiment of the present disclosure, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0316] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present disclosure. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.

[0317] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present disclosure, used to illustrate the technical solutions of the present disclosure, rather than limiting them. The protection scope of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present disclosure can still modify the technical solutions described in the foregoing embodiments or easily conceive of changes, or perform equivalent replacements for some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. A number monitoring method, characterized in that, Including: Collect positive and negative sample data from the outbound call verification data results, user correction data, and recycled number data to generate a training sample set; Based on each number included in the training sample set, extract the historical call detail data of each number from the communication record device storing the number historical call records, and preprocess the historical call detail data to generate a number behavior data set; Extract and process features from the number behavior data set from three dimensions: number features, call behavior, and call behavior changes to generate a number feature variable set; Based on the number feature variable set, construct an initial model for predicting number stability through the XGBoost algorithm, and evaluate and optimize the parameters of the prediction results of the initial model for predicting number stability to generate a number stability prediction model; Deploy the number stability prediction model to the business scenario, obtain the yellow page number data in the number library, and use the number stability prediction model to predict the stability of the yellow page number data to generate a number stability probability value; Based on the number stability probability value, conduct key verification on numbers higher than the preset probability threshold, and update the number library according to the verification results; After using the number stability prediction model to predict the stability of the yellow page number data, it further includes: Adopt a large language model based on the active inference framework to deeply analyze the number stability probability value generated by the number stability prediction model to generate an analysis result at the semantic level; Combined with the analysis result at the semantic level, adopt a multi-objective optimization strategy to output optimized prediction parameters and optimized verification strategies; Based on the optimized prediction parameters and the optimized verification strategy, iteratively optimize the number stability prediction model to improve the accuracy of number stability prediction; The adoption of a large language model based on the active inference framework to deeply analyze the number stability probability value generated by the number stability prediction model to generate an analysis result at the semantic level includes: Obtain the numbers within the preset probability interval of the number stability probability value, and perform in-depth feature analysis on the number to form a preliminary analysis result; Utilize the preliminary analysis result to retrieve the historical case database and analyze the temporal changes in the number usage pattern to generate number change risk features; Based on the active inference framework, fuse the number change risk features, and at the same time extract and process the outbound call record text and merchant status information related to the number, and output a comprehensive evaluation result of the number change risk as the analysis result at the semantic level; The adoption of a multi-objective optimization strategy to output optimized prediction parameters and optimized verification strategies includes: Collect data on prediction accuracy, recognition coverage rate, and resource consumption efficiency, and establish a multi-dimensional optimization index system; Utilize the multi-dimensional optimization index system to calculate the gradient direction of each optimization objective, construct a Pareto optimal solution set, and determine the optimal balance point; According to the optimal balance point, generate updated prediction weights and threshold parameters, and apply the updated parameters to the number stability prediction process.

2. The method according to claim 1, wherein Based on each number included in the training sample set, extract the historical call detail data of each number from the communication record device storing the historical call records of the numbers, and preprocess the historical call detail data to generate a number behavior data set, including: Determine a preset time window range based on the verification time period of the numbers in the training sample set; For each number in the training sample set, extract the call detail records within the preset time window from the communication record device; Perform format standardization, missing value processing, and outlier filtering on the call detail records in sequence to form preprocessed call detail records; Perform structured processing on the preprocessed call detail records to generate the number behavior data set.

3. The method according to claim 1, wherein Extract and process features from the number behavior data set in three dimensions: number features, call behavior, and call behavior changes to generate a number feature variable set, including: Using the number behavior data set, extract information on the city of number ownership, operator type, and number segment type from the number feature dimension, and output a number attribute feature set; According to the number attribute feature set, combined with the number behavior data set, count the number of calls and the number of call objects from the call behavior dimension to form a call behavior feature set; Using the call behavior feature set, calculate the statistical features of each behavior index in the call behavior feature set from the call behavior change dimension, and construct a call behavior change feature set, where the statistical features include at least one of mean or variance; Fuse the number attribute feature set, the call behavior feature set, and the call behavior change feature set to generate the number feature variable set.

4. The method according to claim 1, wherein Evaluate and optimize the parameters of the prediction result of the initial number stability prediction model to generate a number stability prediction model, including: Divide the number feature variable set into a training set and a test set according to a preset ratio; Use the test set to evaluate the performance of the initial number stability prediction model and calculate the performance evaluation index; Adjust the parameters of the initial number stability prediction model according to the performance evaluation index to generate the number stability prediction model.

5. The method according to claim 4, characterized in that After generating the number stability prediction model, it further includes: Input the prediction result of the number stability prediction model into a Bayesian confidence reconfigurable flow accelerator, and perform Monte Carlo sampling through the accelerator to generate the uncertainty distribution of the prediction result; Perform probability calibration calculation according to the uncertainty distribution to form a 90% confidence interval of the prediction result; Based on the 90% confidence interval value of the prediction result, use a dynamic weight adjustment mechanism to generate a calibrated and optimized prediction result, and update the calibrated and optimized prediction result to the number stability prediction model for the number stability prediction process.

6. The method according to claim 1, wherein Deploy the number stability prediction model to the business scenario, including: Adopt a serialization method to convert the number stability prediction model into a model file in a persistent format and complete the environment configuration on the business server; Using the configured business server environment, construct a timed task scheduler and establish an automated prediction process; According to the automated prediction process, implement a prediction process monitoring mechanism, continuously record and output prediction status and result distribution data.

7. A number monitoring device, characterized in that, Including: A sample collection module, used to collect positive and negative sample data from outbound call verification data results, user error correction data, and recycled number data, and generate a training sample set; A data processing module, used to extract the historical call detail data of each number from the communication record device storing the historical call records of numbers based on each number included in the training sample set, and preprocess the historical call detail data to generate a number behavior data set; A feature extraction module, used to extract and process features from three dimensions of number features, call behaviors, and call behavior changes of the number behavior data set to generate a number feature variable set; A model training module, used to construct an initial model for predicting number stability based on the number feature variable set through the XGBoost algorithm, and evaluate and optimize the parameters of the prediction result of the initial model for predicting number stability to generate a number stability prediction model; A prediction module, used to deploy the number stability prediction model to a business scenario, obtain the yellow page number data in the number library, and use the number stability prediction model to predict the stability of the yellow page number data to generate a number stability probability value; A verification and update module, used to conduct key verification on numbers with a probability value higher than a preset probability threshold based on the number stability probability value, and update the number library according to the verification results; After using the number stability prediction model to predict the stability of the yellow page number data, the prediction module is further used for: Adopt a large language model based on the active inference framework to deeply analyze the number stability probability value generated by the number stability prediction model to generate an analysis result at the semantic level; Combined with the analysis result at the semantic level, adopt a multi-objective optimization strategy to output optimized prediction parameters and optimized verification strategies; Based on the optimized prediction parameters and the optimized verification strategy, iteratively optimize the number stability prediction model to improve the accuracy of predicting number stability; The adoption of a large language model based on the active inference framework to deeply analyze the number stability probability value generated by the number stability prediction model to generate an analysis result at the semantic level includes: Obtain the numbers within a preset probability interval of the number stability probability value, and perform in-depth feature analysis on the number to form a preliminary analysis result; Use the preliminary analysis result to retrieve the historical case database and analyze the temporal changes of the number usage pattern to generate number change risk features; Based on the active inference framework, fuse the number change risk features, and simultaneously extract and process the outbound call record text and merchant status information related to the number, and output a comprehensive evaluation result of the number change risk as the analysis result at the semantic level; The adoption of a multi-objective optimization strategy to output optimized prediction parameters and optimized verification strategies includes: Collect data on prediction accuracy, recognition coverage rate, and resource consumption efficiency, and establish a multi-dimensional optimization index system; Using the multi-dimensional optimization index system, calculate the gradient directions of each optimization objective, construct a Pareto optimal solution set, and determine the optimal balance point; According to the optimal balance point, generate updated prediction weights and threshold parameters, and apply the updated parameters to the process of predicting the stability of numbers.

Citation Information

Patent Citations

  • Yellow page number analysis control method and device based on deep learning

    CN119520680A