Number correctness evaluation method and system based on machine learning
Through the machine learning-based number accuracy evaluation method, using XGBoost model and feature engineering technology, the problem that number accuracy evaluation in the existing technology relies on subjective experience and cannot dynamically adapt to data changes is solved, and a more scientific and efficient number accuracy evaluation is achieved.
Patent Information
- Application Number
- CN202510415152.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-03
AI Technical Summary
The number accuracy evaluation method based on fixed rules and subjective experience in the prior art has the problem that data weight assignment is too dependent on subjective experience and cannot dynamically adapt to data changes and capture new features.
A machine learning-based number accuracy evaluation method is used to obtain sample data corresponding to the numbers and names in the number library, and filter and annotate processing are performed to form a training data set. Then, using XGBoost model and feature engineering technology, a number accuracy evaluation model is built, feature importance is dynamically adjusted, model parameters are optimized, and number accuracy is achieved to achieve multi-dimensional evaluation.
It improves the objectivity and scientificity of number accuracy assessment, and can more comprehensively capture the complex feature patterns of number accuracy, improves the comprehensiveness and accuracy of evaluation, and maintains the timeliness of evaluation results through dynamic threshold setting and model update mechanisms.
Smart Images

Figure CN119939193A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data quality management, and in particular to a number correctness assessment method and system based on machine learning. Background Art
[0002] Telephone number databases are an indispensable basic data resource in modern communications and business activities, and the accuracy of numbers has a significant impact on the quality of communication services, corporate image, and user experience. In application scenarios such as corporate address books, customer management systems, and number query services, the accuracy assessment of number data is the core work of data quality management.
[0003] Traditional number accuracy assessment methods mainly rely on manual verification and simple rule judgment. These methods usually verify the validity of numbers by manually dialing, or set fixed weight values for scoring based on single factors such as number update time and source channel. These methods can play a certain role in the management of small-scale number databases, but they lack systematicity and scientificity and are difficult to apply to the management of large-scale number databases.
[0004] The most advanced number accuracy assessment technology currently uses a custom ranking value to identify number accuracy. This method assigns different weights to number data from different sources based on preset rules, and obtains the final accuracy score through a simple weighted calculation. This method is an improvement over pure manual assessment, can handle a larger number database, and provides basic accuracy ranking functions.
[0005] However, this ranking value evaluation method based on fixed rules and subjective experience has obvious defects: its data weight assignment relies too much on subjective experience and lacks objective data support; at the same time, as the scale of the number database expands and the data characteristics become increasingly complex, fixed rules cannot dynamically adapt to data changes and capture new features, resulting in a large deviation between the accuracy evaluation results and the actual situation, and cannot meet the high requirements of modern communications and commercial applications for number accuracy. Summary of the invention
[0006] In view of this, the present application provides a number correctness evaluation method and system based on machine learning, which solves the problem that the data weight assignment of the ranking value evaluation method based on fixed rules and subjective experience in the prior art is too dependent on subjective experience and cannot dynamically adapt to data changes and capture new features.
[0007] The embodiment of the present application provides a number correctness assessment method based on machine learning, including: obtaining sample data corresponding to numbers and names in a number library, screening and annotating the sample data to obtain a training data set; forming a standard feature vector set based on the training data set; using the standard feature vector set to set the initial parameter configuration of an XGBoost model; based on feature importance analysis, screening key feature variables from the standard feature vector set, using the screened feature variables to train the XGBoost model to obtain a number correctness assessment model; using the number correctness assessment model to set a probability threshold; using the number correctness assessment model to predict the numbers in the number library to obtain a predicted probability value; comparing the predicted probability value with the set probability threshold, and outputting a number correctness assessment result.
[0008] According to one embodiment of the present application, a feature source definition list is determined based on the training data set to form a standard feature vector set, including: extracting the conventional attribute features, call behavior fluctuation features, source association features and historical verification data features of the number from the training data set to form a feature source definition list; based on the feature source definition list, extracting the original feature data of the number, encoding and standardizing the original feature data to obtain a feature original value set.
[0009] According to one embodiment of the present application, the initial parameter configuration of the XGBoost model is set using the canonical feature vector set, including: based on the data distribution characteristics of the canonical feature vector set, setting the objective function type of the model to a binary classification function, the maximum depth parameter of the tree to a preset depth value, and the learning rate parameter to a preset learning rate value, to form a basic parameter configuration; according to the basic parameter configuration, setting the sample sampling ratio parameter and the sample balancing processing parameter to obtain the initial parameter configuration.
[0010] According to one embodiment of the present application, key feature variables are screened from the standard feature vector set based on feature importance analysis, including: based on the configured initial parameters, using the standard feature vector set, training the XGBoost model, and calculating the feature importance scores through the feature importance evaluation mechanism of the XGBoost model; screening and sorting the feature variables according to the feature importance scores, determining a preset number of key feature variables, and constructing an optimized feature variable set.
[0011] According to one embodiment of the present application, the number correctness evaluation model is used to set a probability threshold, including: inputting a test data set into the number correctness evaluation model for verification, calculating accuracy, precision and recall indicators, and forming a model performance evaluation report, wherein the test data set is divided from the training data set according to a preset ratio during the training of the XGBoost model; using the model performance evaluation report to analyze the impact of different probability thresholds on model performance, and drawing a threshold-performance correspondence curve; according to the threshold-performance correspondence curve, combined with application requirements, setting the probability threshold.
[0012] According to one embodiment of the present application, after outputting the number correctness evaluation result, it also includes: based on the number correctness evaluation result, performing scenario processing for the application requirements of number library management, data cleaning and recommendation system, and outputting scenario application results; collecting feedback data and newly added number samples in the scenario application results, updating the number correctness evaluation model, and forming an optimized number correctness evaluation model.
[0013] According to one embodiment of the present application, updating the number correctness evaluation model includes: based on the newly added feedback number samples, using a stochastic gradient descent algorithm with a constant learning rate to perform online parameter updates on the number correctness evaluation model; inputting text description information, industry background knowledge, and historical verification records associated with the number into a large language model, extracting semantic features, and generating an unstructured feature vector; fusing the unstructured feature vector with the original features of the number to form a fused vector; and using the fused vector to update the number correctness evaluation model to achieve iterative optimization of the number correctness evaluation model.
[0014] According to one embodiment of the present application, after the output number correctness assessment result, it also includes: for continuous features and categorical features, respectively constructing feature distance calculation matrices for the number, and designing corresponding kernel functions to obtain a hybrid feature processing model; using the hybrid feature processing model, processing the training data set, optimizing the kernel function parameters, and forming a Gaussian process model; from the prediction results of the number correctness assessment model, selecting number samples whose prediction probability values are between a first preset threshold and a second preset threshold, and inputting them into the Gaussian process model, calculating the prediction uncertainty index, so as to identify number samples that require manual verification.
[0015] According to one embodiment of the present application, feature distance calculation matrices for the number are constructed for continuous features and categorical features, respectively, including: according to the continuous features of the number, using a radial basis function kernel to calculate the distance between features to form a continuous feature distance matrix; according to the categorical features of the number, using a Hamming distance kernel to calculate the feature dissimilarity to form a categorical feature distance matrix; and performing a weighted combination of the continuous feature distance matrix and the categorical feature distance matrix to obtain the hybrid feature processing model.
[0016] The embodiment of the present application also provides a number correctness assessment device based on machine learning, including: a training data generation module, used to obtain sample data corresponding to numbers and names in a number library, and screen and annotate the sample data to obtain a training data set; a feature engineering processing module, used to form a standard feature vector set based on the training data set; a model training module, used to use the normalized feature vector set to set the initial parameter configuration of the XGBoost model, and screen key feature variables based on feature importance analysis, and use the screened feature variables to train the XGBoost model to obtain a number correctness assessment model; a result output module, used to use the number correctness assessment model, set a probability threshold, and predict the correctness of the numbers in the number library, and output the number correctness assessment results.
[0017] An embodiment of the present application also provides a computer device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above-mentioned machine learning-based number correctness assessment method.
[0018] An embodiment of the present application also provides a computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable a computer to execute the above-mentioned number correctness assessment method based on machine learning.
[0019] An embodiment of the present application also provides a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the above-mentioned number correctness assessment method based on machine learning.
[0020] This application has the following technical effects: by using the XGBoost algorithm to replace the traditional fixed rule evaluation, a machine learning evaluation model based on multi-dimensional features is realized, which greatly improves the objectivity and scientificity of the number correctness evaluation; by comprehensively considering multi-dimensional features such as the number's general attributes, call behavior fluctuations, source associations and historical verification, the complex feature patterns of the number correctness are fully captured, and the comprehensiveness and accuracy of the evaluation are improved; through the dynamic threshold setting mechanism, the probability threshold can be flexibly adjusted according to the needs of different application scenarios, thereby achieving accurate screening of incorrect numbers; through the model update and maintenance mechanism, the evaluation model is ensured to continuously adapt to data changes and new features, and the timeliness and accuracy of the evaluation results are maintained; through scenario-based application processing, the number correctness evaluation results are seamlessly applied to various business scenarios such as number library management, data cleaning and recommendation systems, thereby improving the overall data quality and user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following is a brief introduction to the drawings required for use in the embodiments. The drawings herein are incorporated into the specification and constitute a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and are used together with the specification to illustrate the technical solutions of the present disclosure. It should be understood that the following drawings only illustrate certain embodiments of the present disclosure and should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can also be obtained based on these drawings without creative work.
[0022] Figure 1 is a flowchart of a method for evaluating number correctness based on machine learning according to an embodiment of the present application; Figure 2 is a schematic diagram of a training data set preparation process according to an embodiment of the present application; Figure 3 is a schematic diagram of a feature engineering process according to an embodiment of the present application; Figure 4 This is a schematic diagram of the XGBoost model training and optimization process according to an embodiment of the present application; Figure 5 It is a schematic diagram of the model evaluation and threshold setting process according to an embodiment of the present application. DETAILED DESCRIPTION
[0023] In order to make the purpose, technical scheme and advantages of the embodiments of the present disclosure clearer, the technical scheme in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all of the embodiments. The components of the embodiments of the present disclosure generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure provided in the drawings is not intended to limit the scope of the present disclosure for protection, but merely represents the selected embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present disclosure.
[0024] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.
[0025] The term "and / or" herein only describes an association relationship, indicating that three relationships may exist. For example, A and / or B may represent the following three situations: A exists alone, A and B exist at the same time, and B exists alone. In addition, the term "at least one" herein represents any combination of at least two of any one or more of a plurality of. For example, including at least one of A, B, and C may represent including any one or more elements selected from the set consisting of A, B, and C.
[0026] The specific implementation of the present application is described in detail below in conjunction with the accompanying drawings. It should be understood that the following description is only exemplary and is not intended to limit the scope of protection of the present application.
[0027] like Figure 1 As shown, the embodiment of the present application provides a method for evaluating number correctness based on machine learning, including: S1: Obtain sample data corresponding to numbers and names in a number database, filter and annotate the sample data, and obtain a training data set.
[0028] In S1, sample data corresponding to numbers and names in the number database are obtained, and the sample data are screened and annotated to obtain a training data set. This step is the basis of the entire number correctness evaluation method and is intended to construct a high-quality training data set.
[0029] First, extract the number data and its corresponding name information from the existing number database as the original data source. Then, according to the correspondence between the number and the name, screen out accurate number (the number matches the name) and inaccurate number (the number does not match the name) samples to form an initial sample set. Next, perform binary labeling on these samples, usually using 1 for accurate numbers and 0 for inaccurate numbers, so that the problem is converted into a binary classification task. Finally, divide the data set into a training set and a test set according to a certain ratio (such as 7:3) to ensure that the distribution of each type of sample in the two data sets is basically the same. Through these processes, a structured and standardized training data set is obtained, laying a solid foundation for subsequent model training.
[0030] Among them, S1 obtains sample data corresponding to numbers and names in the number library, screens and annotates the sample data to obtain a training data set. When implemented specifically, the following sub-steps are included: S1.1: Data sample screening: Based on the historical data in the number database, the accurate number (number and name match) and inaccurate number (number and name do not match) samples are screened according to the correspondence between the number and the name to build the initial sample set. First, the existing number data is selected from the number database as the original data source. By analyzing these number data, the records where the number and name match are identified as accurate number samples, and the records where the number and name do not match are identified as inaccurate number samples. In this step, the accuracy status of the number can be determined using existing verification records, user feedback data, or business rules.
[0031] S1.2: Data sample labeling: Based on the initial sample set selected in step S1.1, perform binary labeling on each number sample, usually using 1 to represent an accurate number (positive sample) and 0 to represent an inaccurate number (negative sample). This binary labeling method transforms the problem into a binary classification task, which is suitable for modeling using machine learning algorithms. The labeling process needs to ensure the accuracy of the label, as this will directly affect the learning effect and prediction performance of the model.
[0032] S1.3: Dataset division: Based on the sample data annotated in step S1.2, divide the dataset into a training set and a test set according to a certain ratio (e.g., 7:3) to ensure that the distribution of each type of sample in the two datasets is basically the same. Stratified sampling technology is used during division to ensure that the ratio of positive and negative samples in the training set and the test set is consistent with the original dataset to avoid model bias caused by uneven sample distribution. The training set is used for model learning and parameter optimization, while the test set is used to evaluate the generalization ability and prediction performance of the model.
[0033] S2: Form a set of standard feature vectors based on the training data set.
[0034] First, based on the training data set, the main dimensions of number characteristics are determined, including general attribute characteristics of the number (such as number type, location, operator, update time, etc.), call behavior fluctuation characteristics (such as recent call frequency, call duration changes, etc.), source association characteristics (such as data source channel, collection method, etc.) and historical verification data characteristics (such as historical verification results, verification time, etc.).
[0035] Then, the raw data of these features are extracted from the number library and related systems, and necessary calculations and conversions are performed.
[0036] Finally, the categorical features are encoded (such as one-hot encoding or label encoding), and the numerical features are standardized (such as Z-score standardization or Min-Max normalization) to ensure that all feature data formats are unified and the range is appropriate. Through these processes, a standardized feature vector set is formed, providing high-quality input data for model training.
[0037] S3: Using the canonical feature vector set, set the initial parameter configuration of the XGBoost model.
[0038] As a powerful ensemble learning algorithm, the performance of XGBoost depends largely on parameter configuration.
[0039] First, based on the data distribution characteristics of the normalized feature vector set, the objective function type of the model is set to a binary classification function (such as "binary:logistic"), because number accuracy evaluation is essentially a binary classification problem. Then, the maximum depth parameter of the tree is set to a preset depth value (such as 8). This parameter controls the complexity of each decision tree and needs to find a balance between model complexity and generalization ability. At the same time, the learning rate parameter is set to a preset learning rate value (such as 0.1). Although a smaller learning rate slows down the model training speed, it can improve the stability and generalization ability of the model.
[0040] In addition, it is necessary to set the sample sampling ratio parameter (such as 0.8) and sample balancing processing parameters to deal with the sample imbalance problem and improve the robustness of the model. These initial parameter configurations provide a good starting point for subsequent model training.
[0041] The specific construction process of the XGBoost model is as follows: Model structure: The XGBoost model in this application adopts a tree ensemble structure, which contains 100 decision trees. The maximum depth of each tree is set to 8 to balance the complexity and generalization ability of the model.
[0042] Learning parameters: The learning rate is set to 0.1, the regularization parameters α and λ are set to 1 and 2 to control the overfitting risk of the model.
[0043] Training process: k=5-fold cross-validation method was used, 500 rounds of training iterations were performed each time, and an early stopping strategy (early_stopping_rounds=50) was used to avoid overfitting.
[0044] Feature selection: Through the feature importance scoring mechanism, feature variables with importance scores greater than 0.01 were screened out, and finally the top 141 most important features were selected.
[0045] Parameter optimization: A grid search method is used to find the optimal combination in the following parameter space: max_depth[4,6,8,10], min_child_weight[1,3,5], gamma[0,0.1,0.2], subsample[0.6,0.8,1.0].
[0046] S4: Based on the feature importance analysis, key feature variables are screened from the standard feature vector set, and the XGBoost model is trained using the screened feature variables to obtain a number correctness evaluation model.
[0047] This step first trains a basic version of the XGBoost model based on the initial parameter configuration using a set of canonical feature vectors, with the goal of evaluating the contribution of each feature to the prediction result. By analyzing the feature importance evaluation mechanism of the XGBoost model (such as the "feature_importances_" attribute), the importance score of each feature is calculated, and the most relevant feature variables (such as "data source type", "number of months since the last manual data", etc.) are selected accordingly. Feature screening can not only improve the training efficiency of the model, but also reduce the risk of overfitting. Then, based on the screened feature set and optimized parameter configuration, the complete model training process is carried out. Cross-validation techniques are usually used, and the optimal parameter combination is found through grid search or random search. Finally, the final model is trained on the entire training set using the optimal parameters to obtain an evaluation model that can effectively distinguish between accurate and inaccurate numbers.
[0048] S5: Using the number correctness assessment model, set a probability threshold.
[0049] This step aims to determine the optimal decision threshold in the model application. First, the test data set is input into the trained model to obtain the model's accuracy prediction probability for each number, and the performance indicators such as accuracy, precision, and recall are calculated to form a model performance evaluation report. Then, the impact of different probability thresholds on the model screening performance is analyzed, and multiple different thresholds (such as from 0.1 to 0.9) are tried. For each threshold, the corresponding performance indicators are calculated, and the threshold-performance correspondence curve (such as threshold-precision curve, threshold-recall curve, etc.) is drawn. Finally, based on these curves and combined with actual application requirements (such as giving priority to ensuring the screening accuracy or pursuing balanced performance), the optimal probability threshold is determined. For example, in the application scenario of screening incorrect numbers, a higher threshold (such as 0.75) may need to be set to achieve a 95% screening accuracy rate. In this way, a scientific basis is provided for the decision-making of the model in practical applications.
[0050] S6: Use the number correctness assessment model to predict the numbers in the number database to obtain a prediction probability value.
[0051] S6: using the number correctness evaluation model to predict the numbers in the number database to obtain a prediction probability value. This step applies the trained model to actual data.
[0052] First, the number data in the number database is preprocessed, including feature extraction, encoding, and standardization, to ensure that the same processing method is used as the training data to ensure the consistency and accuracy of the prediction. Then, the processed feature vector is input into the trained XGBoost model to obtain the accuracy prediction probability value of each number. This probability value represents the possibility that the model considers the number to be an accurate number, and the range is usually between 0 and 1. Through batch processing, all numbers in a large-scale number database can be efficiently evaluated to achieve fast and automated number accuracy prediction.
[0053] S7: Compare the predicted probability value with the set probability threshold, and output the number correctness evaluation result.
[0054] S7 compares the predicted probability value with the set probability threshold and outputs the number correctness evaluation result. According to the probability threshold set in S5, the predicted probability value obtained in S6 is compared with the threshold. Generally, if the predicted probability value is greater than or equal to the threshold, it is determined to be an accurate number; if the predicted probability value is less than the threshold, it is determined to be an inaccurate number. In this way, a clear accuracy evaluation result can be obtained for each number in the number library.
[0055] The predicted probability value is compared with the set probability threshold, and the number correctness evaluation result is output. This is the last step of the method, which converts the prediction result of the model into a specific decision.
[0056] Specifically, the predicted probability value obtained in step S6 is compared with the probability threshold set in step S5. Generally, if the predicted probability value is greater than or equal to the threshold, it is determined to be an accurate number; if the predicted probability value is less than the threshold, it is determined to be an inaccurate number.
[0057] For example, if the threshold is set to 0.75, numbers with a predicted probability greater than or equal to 0.75 are determined to be accurate, and numbers with a predicted probability less than 0.75 are determined to be inaccurate.
[0058] In addition to the binary classification results, the evaluation results can also include the original predicted probability value to provide more fine-grained accuracy evaluation information. The final evaluation results can be output in different formats, such as lists, reports, or visual charts, to meet the needs of different application scenarios. In this way, the accuracy evaluation of all numbers in the number database is achieved, providing strong support for number database management, data cleaning, and business decision-making.
[0059] like Figure 2 As shown, in S2, a set of standard feature vectors is formed according to the training data set, including: S2.1: Extracting conventional attribute features, call behavior fluctuation features, source association features, and historical verification data features of the number from the training data set to form a feature source definition list.
[0060] Based on the training data set generated in step S1.3, the main dimensions of the number features are determined, including the number's general attribute features, call behavior fluctuation features, source association features, and historical verification data features, etc., to form a feature source definition list.
[0061] Specifically, the general attribute characteristics of a number include basic information such as number type, location, operator, and update time; the call behavior fluctuation characteristics include behavioral data such as recent call frequency and changes in call duration; the source association characteristics include related information such as data source channels, collection methods, and provider credibility; the historical verification data characteristics include historical verification results, verification time, verification methods, and other historical data.
[0062] S2.2: Extracting original feature data of the number based on the feature source definition list, encoding and standardizing the original feature data to obtain a set of feature original values.
[0063] Based on the feature source list defined in step S2.1, extract the original feature data from the number library and related systems, perform necessary calculations and conversions, and generate a set of feature original values. Different extraction and calculation methods are used for different types of features.
[0064] For example, for time-related features, the number of months since the last manual verification can be calculated; for behavioral features, the variance of the number of weekly queries in the past four weeks can be calculated; for source features, the codes of different source channels can be extracted, etc.
[0065] In the subsequent process, it is also necessary to encode the categorical features and standardize the numerical features to ensure that all feature data formats are unified and the range is appropriate. For categorical features (such as number type, source channel, etc.), use the One-Hot Encoding or Label Encoding method to convert them into numerical form; for numerical features, use standardization (Z-score standardization) or normalization (Min-Max normalization) to scale the values to the appropriate range to avoid the impact of dimensional differences on model training.
[0066] like Figure 3 As shown, in S3, the standard feature vector set is used to set the initial parameter configuration of the XGBoost model, including: S3.1: Based on the data distribution characteristics of the standard feature vector set, the objective function type of the model is set to a binary classification function, the maximum depth parameter of the tree is set to a preset depth value, and the learning rate parameter is set to a preset learning rate value, thereby forming a basic parameter configuration.
[0067] Based on the normalized feature vector set output from step S2.3, determine the initial parameter configuration of the XGBoost model. For the objective function type, select "binary:logistic" as the objective function because number accuracy evaluation is essentially a binary classification problem. For the maximum depth parameter "max_depth" of the tree, set it to the preset depth value (such as 8). This parameter controls the complexity of each decision tree. Too large a depth may lead to overfitting, while too small a depth may lead to underfitting. For the learning rate parameter "eta", set it to the preset learning rate value (such as 0.1). This is a relatively conservative choice. Although a smaller learning rate slows down the model training speed, it can improve the stability and generalization ability of the model.
[0068] Based on the data distribution characteristics of the standard feature vector set, the objective function type of the model is set to a binary classification function, the maximum depth parameter of the tree is set to a preset depth value, and the learning rate parameter is set to a preset learning rate value, thereby forming a basic parameter configuration.
[0069] In this sub-step, technicians need to set the initial core parameter configuration of the XGBoost model according to the characteristics of the number accuracy prediction task. First, for the objective function type, select "binary:logistic" as the objective function. This is because number accuracy evaluation is essentially a binary classification problem that requires predicting whether the number is accurate.
[0070] Secondly, the maximum depth parameter "max_depth" of the tree is a key parameter for controlling the complexity of the decision tree, which is usually set to 8. This parameter determines the maximum number of layers that each decision tree can split. Too large a depth may cause the model to overfit the training data and lose generalization ability, while too small a depth may cause the model to be too simple and unable to capture complex patterns in the data.
[0071] Again, the learning rate parameter "eta" is usually set to 0.1, which is a relatively conservative choice. The smaller the learning rate, the slower the model training speed, but the stability and generalization ability are usually improved; too large a learning rate may cause the model training to be unstable or fall into a local optimal solution. The reasonable setting of these basic parameters provides a scientific starting point for subsequent model training, ensuring that the model can effectively learn the relationship between number features and accuracy.
[0072] S3.2: According to the basic parameter configuration, set the sample sampling ratio parameter and the sample equalization processing parameter to obtain the initial parameter configuration.
[0073] According to the basic parameter configuration formed in step S3.1, the sample sampling ratio parameter "subsample" (such as 0.8) is set, which means that only 80% of the data is used each time the tree is built. This randomness helps the model better cope with different data distributions. At the same time, considering the difference in the number of accurate numbers and inaccurate numbers in the sample, the sample balancing processing parameter "scale_pos_weight" is set for sample balancing. The parameter value is usually set to the ratio of the number of negative samples to the number of positive samples to ensure that the model does not favor the category with a larger sample size.
[0074] In addition to the basic parameters, some additional parameters need to be set to improve the robustness of the model and handle sample imbalance. First, the sample sampling ratio parameter "subsample" is usually set to 0.8, which means that only 80% of the data samples are randomly used each time the tree is built. This randomness helps reduce the risk of overfitting and makes the model more robust to changes in data distribution.
[0075] Secondly, since there is usually a large difference between the number of accurate numbers and inaccurate numbers in the actual number database (sample imbalance problem), it is necessary to set the sample balance processing parameter "scale_pos_weight" to balance it. This parameter value is usually set to the ratio of the number of negative samples to the number of positive samples to ensure that the model does not favor the category with a larger sample size.
[0076] For example, if the number of inaccurate numbers (negative samples) is three times the number of accurate numbers (positive samples), "scale_pos_weight" can be set to 3. In addition, other parameters such as regularization parameters "alpha" and "lambda" can be set to control the complexity of the model and prevent overfitting. Through the comprehensive configuration of these parameters, a complete initial parameter configuration scheme is formed, laying a solid foundation for model training.
[0077] like Figure 4 As shown, in S4, based on feature importance analysis, key feature variables are screened from the standard feature vector set, including: S4.1: Based on the configured initial parameters, the XGBoost model is trained using the standard feature vector set, and the feature importance score is calculated through the feature importance evaluation mechanism of the XGBoost model.
[0078] Based on the configured initial parameters, a basic XGBoost model is trained using a set of canonical feature vectors. The importance score of each feature is calculated through the "feature_importances_" attribute of the XGBoost model's feature importance evaluation mechanism, which reflects the number of times the feature is used as a split point in all trees and the degree to which these splits improve the model performance.
[0079] The purpose of this sub-step is not to obtain the final prediction model, but to evaluate the contribution of each feature to the prediction result.
[0080] First, a basic XGBoost model is trained using the initial parameter configuration and the normalized feature vector set set in step S3. In the XGBoost model, each split point of the decision tree is based on a certain feature, and the degree to which these splits improve the model performance directly reflects the importance of the feature. The importance score of each feature can be obtained through the "feature_importances_" attribute built into the XGBoost model. This score comprehensively considers the number of times the feature is used as a split point in all trees, and the degree to which these splits improve the model's objective function (such as the log-likelihood function). The feature importance score not only provides insight into the model's decision-making mechanism, but also provides an objective basis for subsequent feature screening. In this way, we can deeply understand the impact of different features on the number accuracy assessment, so as to optimize the model in a more targeted manner.
[0081] S4.2: Screen and sort the feature variables according to the feature importance scores, determine a preset number of key feature variables, and construct an optimized feature variable set.
[0082] The feature variables are screened and sorted according to the feature importance scores, a preset number of key feature variables are determined, and an optimized feature variable set is constructed.
[0083] In this sub-step, based on the feature importance scores calculated in step S4.1, all feature variables are sorted in descending order, and then a preset number (e.g., 141) of feature variables with high importance scores are selected to construct an optimized feature variable set. These selected key features may include "data source type", "number of months from the last manual data", "last outbound call status", "weekly query number variance in the past 4 weeks", etc. They cover multiple dimensional information of the number and can fully reflect the accuracy characteristics of the number. The feature screening process can not only improve the training efficiency of the model (reduce computing resource consumption and training time), but also reduce the risk of overfitting, because irrelevant or redundant features may introduce noise and interfere with the learning process of the model. In addition, feature screening also improves the interpretability of the model, and can more clearly understand which factors have the greatest impact on number accuracy. Through this step, a streamlined and effective feature subset is obtained, which is ready for subsequent full model training.
[0084] like Figure 5 As shown, in S5, the number correctness evaluation model is adopted to set the probability threshold, including: S5.1: Input the test data set into the number correctness evaluation model for verification, calculate the accuracy, precision and recall rate indicators, and form a model performance evaluation report, wherein the test data set is divided from the training data set according to a preset ratio during the training of the XGBoost model.
[0085] In S5.1, based on the number correctness evaluation model completed in S4, a performance test is performed using a test data set.
[0086] The test data set is divided from the training data set in S1.3 according to a preset ratio and is independent of the training data set. The test data set is input into the model to obtain the accuracy prediction probability of the model for each number, and to calculate performance indicators such as accuracy, precision and recall to form a model performance evaluation report.
[0087] The test data set is input into the number correctness evaluation model for verification, and the accuracy, precision and recall rate indicators are calculated to form a model performance evaluation report, wherein the test data set is divided from the training data set according to a preset ratio during the training of the XGBoost model.
[0088] After the model training is completed, its performance needs to be fully evaluated to verify the effectiveness and accuracy of the model.
[0089] First, input the test data set (the independent data set divided in step S1.3) into the trained model to obtain the model's accuracy prediction probability for each number. Then, according to the preset initial threshold (usually 0.5), the prediction probability is converted into a binary classification result (accurate / inaccurate). By comparing the prediction results with the actual labels, a series of performance indicators are calculated: Accuracy reflects the overall proportion of correct classifications by the model; Precision reflects the proportion of numbers predicted as "accurate" by the model that are actually accurate; Recall reflects the proportion of numbers that are actually accurate that are correctly predicted by the model; F1 score is the harmonic average of precision and recall, which comprehensively considers the balance between the two.
[0090] In addition, you can draw ROC curves (receiver operating characteristic curves) and calculate AUC values (area under the curve). These indicators can more comprehensively evaluate the performance of the model at different thresholds. Through the comprehensive analysis of these evaluation indicators, a model performance evaluation report is formed, which provides a data basis for subsequent threshold sensitivity analysis.
[0091] S5.2: Use the model performance evaluation report to analyze the impact of different probability thresholds on model performance and draw a threshold-performance correspondence curve.
[0092] In S5.2, based on the model performance evaluation report of step S5.1, analyze the impact of different probability thresholds on the model screening performance. Try multiple different probability thresholds (such as from 0.1 to 0.9, with an interval of 0.05), and for each threshold, calculate the corresponding precision, recall and other indicators, and draw threshold-precision curves, threshold-recall curves, etc., to form threshold-performance correspondence curves.
[0093] The model outputs the probability value (between 0 and 1) that the sample belongs to the positive class. A threshold needs to be set to convert the probability value into a classification result. Different thresholds will lead to different classification results, thus affecting the model's performance indicators such as precision and recall. In order to find the threshold that best suits practical applications, a threshold sensitivity analysis is required.
[0094] In specific implementation, multiple probability thresholds are selected within a reasonable range (such as 0.1 to 0.9, with an interval of 0.05). For each threshold, the trained model is used to predict the test data set and calculate the corresponding performance indicators. For example, when the threshold is set to 0.5, the model's screening accuracy for incorrect numbers may be 92.3%; when the threshold is increased to 0.75, the screening accuracy can reach 95%, but it may cause a decrease in the recall rate, that is, some incorrect numbers are not screened out. By plotting the performance indicators under different thresholds into curves, threshold-performance correspondence curves are formed, such as threshold-precision curves, threshold-recall curves, threshold-F1 score curves, etc. This visual display can intuitively reflect the impact of threshold changes on various performance indicators, helping technicians to deeply understand the behavioral characteristics of the model under different decision boundaries.
[0095] S5.3: According to the threshold-performance correspondence curve and in combination with application requirements, a probability threshold is set.
[0096] In S5.3, based on the threshold-performance correspondence curve generated in step S5.2, the optimal probability threshold is determined in combination with actual application requirements (such as giving priority to ensuring the screening accuracy). For example, in the application scenario of screening incorrect numbers, more attention may be paid to the accuracy rate to avoid mistakenly judging the correct number as incorrect; when the probability threshold is selected as 0.75, a 95% incorrect number screening accuracy rate can be achieved.
[0097] In different application scenarios, the emphasis on precision and recall may be different, so the optimal probability threshold needs to be determined according to specific needs. For example, in the number library management scenario, more emphasis may be placed on the screening accuracy of incorrect numbers (i.e., precision) to avoid mistakenly judging correct numbers as incorrect; while in the recommendation system, more emphasis may be placed on recall to ensure that as many accurate numbers as possible are recommended.
[0098] In the application of screening incorrect numbers, when the probability threshold is selected as 0.75, a 95% screening accuracy rate can be achieved. This shows that by increasing the threshold, the precision of the model can be further improved, although it may be at the expense of a certain recall rate. The specific threshold selection should be based on business needs and acceptable risk levels to find the best balance between precision and recall.
[0099] In addition, more complex methods can be used to determine the optimal threshold, such as those based on cost functions. For example, the business costs of different types of errors (false positives and false negatives) can be defined, and then the threshold that minimizes the overall cost can be selected. The final threshold, together with the trained model, will constitute a complete number correctness assessment solution, ensuring that the model can provide the assessment results that best meet business needs in actual applications.
[0100] The specific algorithm for threshold setting is as follows: Performance index calculation: For each candidate threshold T (from 0.1 to 0.9, step size 0.05), the following performance index is calculated: Precision P(T)=TP(T) / (TP(T)+FP(T)), Recall rate R(T)=TP(T) / (TP(T)+FN(T)), F1 score F1(T)=2·P(T)·R(T) / (P(T)+R(T)), Among them, TP(T), FP(T), and FN(T) are the number of true positives, false positives, and false negatives under the threshold T, respectively.
[0101] Cost function: Define a cost function C(T) that takes into account business needs: C(T)=w1·(1-P(T))+w2·(1-R(T)), Among them, w1 and w2 are weight coefficients set according to business requirements, indicating the relative importance of precision and recall.
[0102] Optimal threshold selection: Select the threshold that minimizes the cost function C(T) as the optimal threshold: T_opt=argminC(T), Threshold Validation: The performance of the selected threshold is validated on an independent validation set to ensure its stability on different datasets.
[0103] In addition, after outputting the number correctness evaluation result in S7, the embodiment of the present application further includes: S8: Based on the number correctness evaluation result, scenario processing is performed according to the application requirements of number library management, data cleaning and recommendation system, and scenario application results are output.
[0104] S8 performs scenario-based processing according to the number correctness evaluation results, targeting the application requirements of number library management, data cleaning and recommendation systems, and outputs scenario application results: for the number library management scenario, the number library can be sorted, graded or marked based on the evaluation results, so that management personnel can process number data of different quality levels in a targeted manner; for the data cleaning scenario, low-accuracy numbers can be automatically screened out for verification or update based on the evaluation results; for the recommendation system scenario, the number accuracy probability value can be used as a weight factor of the recommendation algorithm to ensure that numbers with high accuracy are recommended first.
[0105] S9: Collect feedback data and newly added number samples in the scenario application results, update the number correctness evaluation model, and form an optimized number correctness evaluation model.
[0106] S9 collects feedback data and new number samples in the scenario application results, updates the number correctness evaluation model, and forms an optimized number correctness evaluation model: In the actual application process, user feedback, new number verification results and other data can be used as new training samples for regular model update and optimization. A fixed update cycle (such as monthly or quarterly) can be set, or the update process can be triggered based on the number of new samples or changes in model performance.
[0107] Wherein, S9 updates the number correctness assessment model, including: S9.1: Based on the newly added feedback number samples, the stochastic gradient descent algorithm with a constant learning rate is used to perform online parameter update on the number correctness evaluation model.
[0108] Specifically, based on the newly added feedback number samples, the stochastic gradient descent algorithm with a constant learning rate is used to perform online parameter updates on the number correctness assessment model. Online learning allows the model to be updated instantly when new data arrives, without the need to retrain the entire model. When the system receives a new number verification result, it is immediately input into the model as a training sample, and the model parameters are adjusted through a gradient update. The design of a constant learning rate ensures the stability of the model parameter update.
[0109] According to the newly added feedback number samples, the stochastic gradient descent algorithm with a constant learning rate is used to update the parameters of the number correctness evaluation model online. Traditional model updates usually require batch retraining after collecting enough new samples, which is time-consuming and consumes a lot of computing resources. Online learning technology allows the model to be updated instantly when new data arrives, without the need to retrain the entire model, greatly improving the system's response speed and adaptability.
[0110] In the specific implementation process, when the system receives a new number verification result (such as confirmation of number accuracy obtained through user feedback or external verification), it is immediately input into the model as a training sample, and the model parameters are adjusted through a gradient update. Unlike the traditional learning rate decay strategy, a constant learning rate design is adopted here to ensure the stability of model parameter updates and continuous adaptability to new patterns. The constant learning rate avoids the problem of reduced model adaptability to new data due to excessive learning rate decay, and is particularly suitable for scenarios such as number accuracy assessment where the data distribution may change over time. In addition, online updates can also be combined with window-based strategies to give higher weights to samples in the recent period, so that the model pays more attention to recent data patterns and further improves the response speed to changes in data distribution.
[0111] S9.2: Input text description information, industry background knowledge, and historical verification records associated with the number into a large language model to extract semantic features and generate an unstructured feature vector.
[0112] Input the text description information, industry background knowledge and historical verification records associated with the number into the large language model, extract semantic features and generate unstructured feature vectors. The large language model can extract implicit features from unstructured data such as text descriptions, industry background and historical records, such as identifying the industry type from the company description associated with the number, summarizing the verification mode from historical records, etc.
[0113] Input the text description information, industry background knowledge and historical verification records associated with the number into the large language model, extract semantic features, and generate unstructured feature vectors. The XGBoost model mainly processes structured feature data, but in actual number management, there is also a large amount of unstructured text information related to the number, such as descriptions of companies associated with the number, industry background knowledge, notes on historical verification records, etc. These text data may contain rich semantic information, which is of great value in judging the accuracy of the number. With its powerful semantic understanding and reasoning capabilities, the large language model (LLM) can extract implicit features from these unstructured texts.
[0114] During the implementation process, we first collect various text description information related to the number, and then input these texts into the pre-trained large language model. LLM can understand the semantic content of the text, such as identifying industry type, business status and other information from the company description, and analyzing verification mode and change rules from historical records. Through the processing of LLM, these text information are converted into fixed-dimensional feature vectors. These vectors capture the semantic information in the text and can reflect the consistency between the number and the name, the activity of the company, the timeliness of the information and other characteristics of multiple dimensions, providing a new source of information for subsequent feature fusion and model optimization.
[0115] The large language model used in this application is specifically implemented as follows: Model architecture: A Transformer-based pre-trained language model is used, which includes a 12-layer attention mechanism, a hidden layer dimension of 768, and 12 attention heads.
[0116] Input processing: Tokenize the text information related to the number (such as company description and industry background knowledge) and input it into the model. The maximum sequence length is set to 256.
[0117] Feature extraction: The [CLS] token output of the last layer of the model is used to obtain a 768-dimensional text representation vector as a semantic feature vector.
[0118] Fine-tuning strategy: Perform domain-adaptive fine-tuning on text corpora related to the number field, use the Adam optimizer with a learning rate of 2e-5, and train for 10 epochs.
[0119] Reasoning process: batch processing (batch_size=32) of the relevant text of each number, extracting semantic features, and reducing the dimension to 64 through dimensionality reduction technology (PCA) to maintain a balance with the structured features.
[0120] S9.3: Fusing the unstructured feature vector with the original features of the number to form a fused vector.
[0121] The unstructured feature vector is fused with the original features of the number to form a fused vector. It is necessary to design a suitable feature fusion strategy, such as feature splicing, weighted combination or hierarchical fusion, to organically combine semantic features with structured features. The fused feature vector contains more comprehensive and richer number information.
[0122] In order to fully utilize the complementary advantages of structured features and unstructured features, it is necessary to effectively fuse these two types of features. Feature fusion is a key step that directly affects the quality and expression ability of the fused features.
[0123] In the implementation process, three main fusion strategies are usually adopted: feature concatenation (directly connecting two types of feature vectors by dimension into a longer vector), weighted combination (setting weights for different types of features and then performing weighted summation), or hierarchical fusion (first processing in each feature space and then fusing at the high-level semantic level). For the task of number accuracy assessment, feature concatenation is usually the most direct and effective method, but it is necessary to pay attention to the problem of unbalanced feature dimensions. High-dimensional features can be compressed through dimensionality reduction techniques (such as PCA or t-SNE).
[0124] In addition, the attention mechanism can also be introduced to dynamically adjust the weights of different features according to the characteristics of the current sample, making the fusion process more intelligent and adaptive. The fused feature vector contains multi-dimensional information of the number, including both precise structural attributes (such as time, frequency, source, etc.) and rich semantic-level descriptions (such as industry background, corporate status, etc.), which can provide the model with a more comprehensive and in-depth basis for judgment and improve the accuracy and robustness of the evaluation.
[0125] The specific implementation method of feature fusion is as follows: Feature standardization: Z-score standardization is performed on structured features and unstructured features respectively, so that their mean is 0 and standard deviation is 1.
[0126] Dimension matching: The unstructured feature vector (768 dimensions) is reduced to a dimension close to the structured feature (64 dimensions) through principal component analysis (PCA).
[0127] Feature splicing: directly splice the standardized structured feature vector and the dimensionally reduced unstructured feature vector to form a fused feature vector.
[0128] Feature weight: The attention mechanism is introduced to dynamically adjust the weights of the two types of features. The specific formula is: F 融合 =α·F 结构化 +(1-α)·F 非结构化 ; Among them, α is the weight coefficient calculated by the attention network, which is adaptively adjusted according to the characteristics of the current sample.
[0129] Feature selection: The fused feature vectors are evaluated again for feature importance to select the feature subset that is most valuable for prediction.
[0130] S9.4: Using the fusion vector, update the number correctness evaluation model to achieve iterative optimization of the number correctness evaluation model.
[0131] Use the fused feature vector to retrain or fine-tune the evaluation model so that the model can make more accurate evaluation judgments using the newly added semantic features.
[0132] The following is a detailed description of steps S9.1-9.4 and S10.1-10.3, formatted in natural paragraphs: S9.1: Based on the newly added feedback number samples, the stochastic gradient descent algorithm with a constant learning rate is used to perform online parameter update on the number correctness evaluation model.
[0133] Traditional model updates usually require batch retraining after collecting enough new samples, which is time-consuming and consumes a lot of computing resources. Online learning technology allows the model to be updated instantly when new data arrives, without the need to retrain the entire model, greatly improving the system's response speed and adaptability.
[0134] In the specific implementation process, when the system receives a new number verification result (such as confirmation of number accuracy obtained through user feedback or external verification), it is immediately input into the model as a training sample, and the model parameters are adjusted through a gradient update. Unlike the traditional learning rate decay strategy, a constant learning rate design is adopted here to ensure the stability of model parameter updates and continuous adaptability to new patterns. The constant learning rate avoids the problem of reduced model adaptability to new data due to excessive learning rate decay, and is particularly suitable for scenarios such as number accuracy assessment where the data distribution may change over time. In addition, online updates can also be combined with window-based strategies to give higher weights to samples in the recent period, so that the model pays more attention to recent data patterns and further improves the response speed to changes in data distribution.
[0135] S9.2: Input text description information, industry background knowledge, and historical verification records associated with the number into a large language model to extract semantic features and generate an unstructured feature vector.
[0136] The XGBoost model mainly processes structured feature data, but in actual number management, there is also a large amount of unstructured text information related to the number, such as descriptions of the number-related companies, industry background knowledge, notes on historical verification records, etc. These text data may contain rich semantic information, which is of great value in judging the accuracy of the number.
[0137] Large language models (LLMs) can extract implicit features from these unstructured texts with their powerful semantic understanding and reasoning capabilities. During the implementation process, various types of text description information related to the number are first collected, and then these texts are input into the pre-trained large language model. LLM can understand the semantic content of the text, such as identifying industry type, operating status and other information from the company description, and analyzing verification patterns and change rules from historical records. Through the processing of LLM, these text information are converted into feature vectors of fixed dimensions. These vectors capture the semantic information in the text and can reflect the consistency of the number and name, the activity of the company, the timeliness of the information and other characteristics in multiple dimensions, providing a new source of information for subsequent feature fusion and model optimization.
[0138] S9.3: Fusing the unstructured feature vector with the original features of the number to form a fused vector.
[0139] In order to fully utilize the complementary advantages of structured features and unstructured features, it is necessary to effectively fuse these two types of features. Feature fusion is a key step that directly affects the quality and expression ability of the fused features.
[0140] In the implementation process, three main fusion strategies are usually adopted: feature concatenation (directly connecting two types of feature vectors by dimension into a longer vector), weighted combination (setting weights for different types of features and then performing weighted summation), or hierarchical fusion (first processing in each feature space and then fusing at the high-level semantic level). For the task of number accuracy assessment, feature concatenation is usually the most direct and effective method, but it is necessary to pay attention to the problem of unbalanced feature dimensions. High-dimensional features can be compressed through dimensionality reduction techniques (such as PCA or t-SNE).
[0141] In addition, the attention mechanism can also be introduced to dynamically adjust the weights of different features according to the characteristics of the current sample, making the fusion process more intelligent and adaptive. The fused feature vector contains multi-dimensional information of the number, including both precise structural attributes (such as time, frequency, source, etc.) and rich semantic-level descriptions (such as industry background, corporate status, etc.), which can provide the model with a more comprehensive and in-depth basis for judgment and improve the accuracy and robustness of the evaluation.
[0142] S9.4: Using the fusion vector, update the number correctness evaluation model to achieve iterative optimization of the number correctness evaluation model.
[0143] After obtaining the fused feature vector, it is necessary to effectively apply it to the model update. Different update strategies can be adopted according to the scale and nature of the fused features. If the fused features are similar to the original features in dimension, the fused features can be directly used to replace the original features and retrain or fine-tune the XGBoost model. If the fused feature dimension increases significantly, the model structure or parameters may need to be adjusted to adapt to the new feature space.
[0144] In the implementation process, an effective method is to adopt an integrated learning strategy, retain the original XGBoost model, train a new model to specifically handle the fusion features, and then combine the advantages of the two models through model fusion (such as weighted averaging or stacking). In addition, transfer learning technology can be used to transfer the knowledge of the original model to the new feature space to accelerate the model's adaptation and learning on the new features.
[0145] Model updates should not be a one-off process, but a closed-loop iterative optimization process: as new data continues to accumulate and the model continues to be applied, the system can automatically collect feedback, extract new features, optimize the model, and then redeploy the updated model to actual applications. Through this continuous iteration, the number correctness assessment model can continuously improve itself, adapt to changes in data distribution and the emergence of new features, and provide increasingly accurate assessment results.
[0146] In another embodiment, after S7 outputs the number correctness evaluation result, it also includes: S10: For continuous features and categorical features, construct feature distance calculation matrices for the numbers respectively, and design corresponding kernel functions to obtain a hybrid feature processing model.
[0147] Among them, S10 constructs a feature distance calculation matrix for the number for the continuous feature and the categorical feature respectively, including: S10.1: Based on the continuous features of the number, radial basis function kernel is used to calculate the distance between features to form a continuous feature distance matrix.
[0148] According to the continuous features of the number, the radial basis function kernel is used to calculate the distance between features to form a continuous feature distance matrix. For the continuous features of the number (such as call frequency, update time, etc.), the radial basis function kernel (RBF kernel) can effectively calculate the similarity of samples in the continuous feature space.
[0149] According to the continuous characteristics of the number, the radial basis function kernel is used to calculate the distance between the characteristics to form a continuous characteristic distance matrix. The Gaussian process model is a powerful non-parametric Bayesian method, which is particularly suitable for handling uncertainty estimation.
[0150] When building a Gaussian process model, the selection and design of the kernel function is a key step, which determines how the model measures the similarity of samples. For continuous features of numbers (such as call frequency, update time, query times, and other numerical features), the radial basis function kernel (RBF kernel, also known as the Gaussian kernel) is a widely used choice.
[0151] The form of the RBF kernel function is k(x,x')=exp(-||x-x'||² / 2l²), where ||x-x'|| is the Euclidean distance between samples and l is the length scale parameter that controls the rate at which similarity changes with distance.
[0152] In the implementation process, all continuous features are first separated from the number features, and these features are preprocessed (such as standardization) to eliminate dimensional differences. Then, the Euclidean distance between any two samples on these continuous features is calculated, and the distance is converted into a similarity measure through the RBF kernel function. In this way, for all sample pairs in the data set, their similarity in the continuous feature space can be calculated to form an n×n continuous feature distance matrix (n is the number of samples). This matrix captures the distribution structure of samples in the continuous feature space and provides a basic metric for the subsequent Gaussian process model.
[0153] S10.2: Based on the classification features of the number, the feature dissimilarity is calculated using the Hamming distance kernel to form a classification feature distance matrix.
[0154] According to the classification characteristics of the number, the Hamming distance kernel is used to calculate the feature dissimilarity to form a classification feature distance matrix. For classification characteristics (such as source type, number type, etc.), the Hamming distance can measure the degree of difference between two samples in classification characteristics.
[0155] According to the classification characteristics of the number, the Hamming distance kernel is used to calculate the feature dissimilarity to form a classification feature distance matrix. In addition to continuous features, numbers also have many classification characteristics (such as number type, source channel, industry, etc.), which cannot be directly calculated using Euclidean distance.
[0156] For categorical features, Hamming distance is a suitable metric that counts the number of elements that differ in corresponding positions in two vectors of equal length.
[0157] In the implementation process, we first need to convert the categorical features into an appropriate representation. Usually, we use One-Hot Encoding to convert the categorical variables into binary vectors. Then, for the encoded vectors of any two samples, we calculate their Hamming distance, that is, the number of different bits. The Hamming distance kernel function can be expressed as k(x,x')=exp(-γ·H(x,x')), where H(x,x') is the Hamming distance and γ is a scaling parameter that controls the degree of influence of distance on similarity.
[0158] In this way, the similarity of all sample pairs in the data set in terms of classification features is calculated to form a classification feature distance matrix. This matrix reflects the similarity relationship between samples in the classification feature space and captures the impact of category information on number accuracy. Unlike the continuous feature matrix, the classification feature matrix focuses more on the measurement of category consistency, providing another angle for similarity evaluation.
[0159] S10.3: Performing a weighted combination of the continuous feature distance matrix and the classification feature distance matrix to obtain the mixed feature processing model.
[0160] The continuous feature distance matrix and the classification feature distance matrix are weightedly combined to obtain the mixed feature processing model. By setting appropriate weight coefficients, the continuous feature distance matrix and the classification feature distance matrix are combined into a composite kernel function matrix as the basis of the Gaussian process model.
[0161] The continuous feature distance matrix and the categorical feature distance matrix are weightedly combined to obtain the hybrid feature processing model. Continuous features and categorical features capture different aspects of number characteristics, and they need to be effectively combined to form a unified similarity measure. In the Gaussian process model, this combination is usually achieved through a linear combination or product of kernel functions.
[0162] This application adopts a weighted linear combination method, that is, k(x,x')=α·k_cont(x,x')+β·k_cat(x,x'), where k_cont is a continuous feature kernel function, k_cat is a classification feature kernel function, and α and β are weight coefficients that control the relative importance of the two types of features.
[0163] During implementation, these weight coefficients can be optimized through methods such as cross-validation or maximizing marginal likelihood to find the combination that best suits the current data. The weighted combination results in a new kernel function matrix that integrates the information of continuous features and categorical features, and can more comprehensively measure the similarity between samples. This mixed feature processing model is the core component for building a Gaussian process model, which determines how the model understands the data distribution and makes predictions. By reasonably designing and optimizing this mixed kernel function, the Gaussian process model can better adapt to the characteristics of number data, provide more accurate predictions and more reasonable uncertainty estimates, especially for those boundary cases that are difficult to clearly judge through the XGBoost model.
[0164] S11: Utilize the hybrid feature processing model to process the training data set, optimize kernel function parameters, and form a Gaussian process model.
[0165] S11 uses the hybrid feature processing model to process the training data set, optimizes the kernel function parameters, and forms a Gaussian process model: by maximizing the edge likelihood function, the various parameters in the composite kernel function are optimized, such as the length scale parameter of the RBF kernel, the weight coefficient of the Hamming distance kernel, etc. This optimization process is usually implemented using numerical optimization techniques such as the conjugate gradient method or the L-BFGS algorithm. The optimized Gaussian process model can not only predict the accuracy probability value of the number, but also estimate the uncertainty of this prediction.
[0166] S12: From the prediction results of the number correctness assessment model, select number samples whose prediction probability values are between the first preset threshold and the second preset threshold, input them into the Gaussian process model, calculate the prediction uncertainty index, and identify number samples that require manual verification.
[0167] S12 selects number samples whose prediction probability values are between a first preset threshold and a second preset threshold from the prediction results of the number correctness evaluation model, inputs them into the Gaussian process model, calculates the prediction uncertainty index, and identifies number samples that require manual verification.
[0168] When the XGBoost model predicts some numbers in the "gray area" (such as probability values between 0.4-0.6), these numbers are input into the Gaussian process model for further analysis. In addition to outputting the mean accuracy probability, the Gaussian process model also provides the variance or standard deviation, which indicates the uncertainty of the prediction. By setting an appropriate uncertainty threshold, it is possible to identify those numbers whose accuracy is still difficult to determine even after evaluation by two models and mark them as requiring manual verification.
[0169] The embodiment of the present application also provides a number correctness assessment device based on machine learning, including: a training data generation module, used to obtain sample data corresponding to numbers and names in a number library, and screen and annotate the sample data to obtain a training data set; a feature engineering processing module, used to form a standard feature vector set based on the training data set; a model training module, used to use the normalized feature vector set to set the initial parameter configuration of the XGBoost model, and screen key feature variables based on feature importance analysis, and use the screened feature variables to train the XGBoost model to obtain a number correctness assessment model; a result output module, used to use the number correctness assessment model, set a probability threshold, and predict the correctness of the numbers in the number library, and output the number correctness assessment results.
[0170] The machine learning-based number correctness assessment method and system in the embodiments of the present application achieve a scientific assessment of number accuracy by utilizing the XGBoost algorithm and multi-dimensional feature engineering, effectively solving the problems of strong subjectivity and poor adaptability of traditional assessment methods, and providing an advanced and reliable technical solution for number data quality management in modern communications and commercial applications.
[0171] The present disclosure also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the machine learning-based number correctness assessment method and system described in the above method embodiment are executed. The storage medium can be a volatile or non-volatile computer-readable storage medium.
[0172] In addition, an embodiment of the present disclosure also provides a computer program product, which stores a computer program. When the computer program is executed by a processor, the steps of the machine learning-based number correctness assessment method and system provided in any of the above embodiments of the present disclosure are executed. For details, please refer to the above method embodiments, which will not be repeated here.
[0173] The computer program product may be implemented in hardware, software or a combination thereof. In an optional embodiment, the computer program product is embodied as a computer storage medium, which may be a volatile or non-volatile computer-readable storage medium. In another optional embodiment, the computer program product is embodied as a software product, such as a software development kit (SDK) and the like.
[0174] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, the specific working process of the above-described equipment and devices can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here. In the several embodiments provided in the present disclosure, it should be understood that the disclosed equipment, devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0175] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0176] In addition, each functional unit in each embodiment of the present disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0177] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium that is executable by a processor. Based on this understanding, the technical solution of the present disclosure, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present disclosure. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0178] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present disclosure, which are used to illustrate the technical solutions of the present disclosure, rather than to limit them. The protection scope of the present disclosure is not limited thereto. Although the present disclosure is described in detail with reference to the aforementioned embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the aforementioned embodiments within the technical scope disclosed in the present disclosure, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should be included in the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure shall be based on the protection scope of the claims.
Claims
1. A method for evaluating number correctness based on machine learning, characterized in that: include: Obtain sample data corresponding to numbers and names in the number library, screen and annotate the sample data to obtain a training data set; According to the training data set, a set of standard feature vectors is formed; Using the canonical feature vector set, setting the initial parameter configuration of the XGBoost model; According to the feature importance analysis, key feature variables are screened from the standard feature vector set, and the XGBoost model is trained using the screened feature variables to obtain a number correctness evaluation model; Using the number correctness assessment model, setting a probability threshold; Using the number correctness assessment model to predict the numbers in the number database to obtain a prediction probability value; The predicted probability value is compared with the set probability threshold, and the number correctness evaluation result is output.
2. The method according to claim 1, characterized in that Determine a feature source definition list according to the training data set to form a standard feature vector set, including: Extracting conventional attribute features, call behavior fluctuation features, source association features, and historical verification data features of the number from the training data set to form a feature source definition list; According to the feature source definition list, the original feature data of the number is extracted, and the original feature data is encoded and standardized to obtain a feature original value set.
3. The method according to claim 1, characterized in that Using the canonical feature vector set, the initial parameter configuration of the XGBoost model is set, including: Based on the data distribution characteristics of the canonical feature vector set, the objective function type of the model is set to a binary classification function, the maximum depth parameter of the tree is set to a preset depth value, and the learning rate parameter is set to a preset learning rate value, thereby forming a basic parameter configuration; According to the basic parameter configuration, the sample sampling ratio parameter and the sample equalization processing parameter are set to obtain the initial parameter configuration.
4. The method according to claim 1, characterized in that: According to the feature importance analysis, key feature variables are screened from the standard feature vector set, including: Based on the configured initial parameters, the XGBoost model is trained using the set of canonical feature vectors, and a feature importance score is calculated using a feature importance evaluation mechanism of the XGBoost model; The feature variables are screened and sorted according to the feature importance scores, a preset number of key feature variables are determined, and an optimized feature variable set is constructed.
5. The method according to claim 1, characterized in that The adopting the number correctness evaluation model to set the probability threshold comprises: Input the test data set into the number correctness evaluation model for verification, calculate the accuracy, precision and recall rate indicators, and form a model performance evaluation report, wherein the test data set is divided from the training data set according to a preset ratio during the training of the XGBoost model; Using the model performance evaluation report, analyze the impact of different probability thresholds on model performance, and draw a threshold-performance correspondence curve; The probability threshold is set according to the threshold-performance correspondence curve and in combination with application requirements.
6. The method according to claim 1, characterized in that After outputting the number correctness evaluation result, the method further includes: According to the number correctness evaluation results, scenario processing is performed for the application requirements of number library management, data cleaning and recommendation system, and scenario application results are output; The feedback data and newly added number samples in the scenario application results are collected, the number correctness evaluation model is updated, and an optimized number correctness evaluation model is formed.
7. The method according to claim 6, characterized in that The updating of the number correctness evaluation model comprises: According to the newly added feedback number samples, the random gradient descent algorithm with a constant learning rate is used to perform online parameter update on the number correctness evaluation model; Input text description information, industry background knowledge, and historical verification records associated with the number into a large language model to extract semantic features and generate an unstructured feature vector. The unstructured feature vector is fused with the original feature of the number to form a fused vector; The fusion vector is used to update the number correctness evaluation model, thereby achieving iterative optimization of the number correctness evaluation model.
8. The method according to claim 1, characterized in that After outputting the number correctness evaluation result, the method further includes: For continuous features and categorical features, feature distance calculation matrices for the numbers are constructed respectively, and corresponding kernel functions are designed to obtain a hybrid feature processing model; Using the hybrid feature processing model, processing the training data set, optimizing kernel function parameters, and forming a Gaussian process model; From the prediction results of the number correctness evaluation model, select number samples whose prediction probability values are between the first preset threshold and the second preset threshold, input them into the Gaussian process model, calculate the prediction uncertainty index, and identify number samples that require manual verification.
9. The method according to claim 8, characterized in that A feature distance calculation matrix for the number is constructed for the continuous feature and the categorical feature, including: According to the continuous features of the number, radial basis function kernel is used to calculate the distance between features to form a continuous feature distance matrix; According to the classification characteristics of the number, the characteristic dissimilarity is calculated using the Hamming distance kernel to form a classification characteristic distance matrix; The continuous feature distance matrix and the classification feature distance matrix are weightedly combined to obtain the mixed feature processing model.
10. A device for evaluating number correctness based on machine learning, characterized in that: include: A training data generation module is used to obtain sample data corresponding to numbers and names in the number library, screen and annotate the sample data, and obtain a training data set; A feature engineering processing module, used to form a standard feature vector set according to the training data set; A model training module is used to screen key feature variables from a standard feature vector set based on feature importance analysis, and train the XGBoost model using the screened feature variables to obtain a number correctness evaluation model; The result output module is used to adopt the number correctness evaluation model, set the probability threshold, and perform correctness prediction on the numbers in the number database, and output the number correctness evaluation result.
Citation Information
Patent Citations
Method for predicting number-carrying transfer-out of telecommunication user based on machine learning
CN112153636A
Network fraud number detection method and system, storage medium and terminal equipment
CN113591924A
Number classification verification method and related device
CN119484700A
Enterprise credit scoring model modeling method and system
CN119722282A
Identity information risk assessment method and apparatus, and computer device and storage medium
WO2020015089A1
Cited By
Coal bed gas productivity prediction method and system combined with artificial intelligence
CN120258328A
Sandy soil liquefaction discrimination method based on CPTU and weighted nonlinear similarity
CN120974283A
Large code model-oriented multi-stage core code data selection method and system
CN122364841A
A Multi-Stage Core Code Data Selection Method and System for Large Code Models
CN122364841B