Method, apparatus, electronic device, and storage medium for express number recognition

The method uses data preprocessing and a stacked ensemble model with XGboost, NGboost, Catboost, and LightGBM algorithms to accurately and stably identify courier numbers, enhancing call connection rates and communication efficiency.

CN114547001BActive Publication Date: 2025-07-15SHANDONG BRANCH OF BEST TONE INFORMATION
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210106569.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-28
Publication Date
2025-07-15
Estimated Expiration
2042-01-28

AI Technical Summary

Technical Problem

The prior art cannot identify express numbers with high accuracy and stability, resulting in low communication user connectivity and poor communication efficiency. The existing models are prone to falling into local optimal solutions, and performance deteriorates when processing high-dimensional feature variables.

Method used

The SMOTE TomeK algorithm is used for data sampling, combined with XGboost, NGboost and Catboost models for primary training, used LightGBM models as secondary models, and fused multiple models through Stacking strategies to optimize hyperparameters to improve accuracy and stability, forming an XNCLBoost model.

Benefits of technology

It improves the accuracy and stability of express number identification, improves the connection rate and communication efficiency, provides high-accurate incoming number identification, and creates a good communication network environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure HDA0003493661270000011
    Figure HDA0003493661270000011
Patent Text Reader

Abstract

The present invention relates to a method, device, electronic device and storage medium for express number recognition. The method for express number recognition includes the steps of: S1, inputting express number black and white list data and signaling call record data, cleaning the data, and obtaining the original data set required for the model through data association and fusion; S2, using the SMOTE TomeK algorithm to perform comprehensive sampling on the original data set to form a model sample data set; S3, dividing the model sample data set into a training set and a test set, and respectively using the XGboost model, NGboost model, and Catboost model for model training to form a primary model; S4, using the five-fold cross-validation method to train the XGboost model, NGboost model, and Catboost model, using the test set for verification, and outputting prediction values; S5, respectively using the output values of the XGboost model, NGboost model, and Catboost model as input features of the LightGBM model to form a secondary model, performing training on the secondary model, and outputting a model that meets the pre-set model accuracy after training, thereby forming an XNCLBoost model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to network communication technology, and in particular to a method, device, electronic device and storage medium for express number identification. Background Art

[0002] With the continuous development of the new generation of 5G communication technology, communication users are experiencing the convenience brought by communication technology in life and work; at the same time, harassing calls have brought troubles to people, not only disrupting the normal life and work order of communication users, but also greatly endangering the calls of some normal numbers. For example, if express delivery calls are identified as harassing calls, it will bring harm to the vital interests of the people. How to identify express delivery numbers from the existing network, so as to provide users with accurate identification and prompts of incoming call numbers, has become a technical issue of common concern to communication operators. However, because the problem of express delivery number identification is too complicated, it has not yet been completely solved.

[0003] In real life, both express delivery numbers and fraud numbers have characteristics such as high call frequency, large call volume, and short call time. At the same time, express delivery numbers have characteristics such as rapid location changes in a short period of time.

[0004] The prior art does not involve a solution for efficient identification of express numbers. Patent application title: A method for identifying express numbers based on call behavior (CN109274834A), based on constructing a blacklist and whitelist call record table and express feature identification rules, and then obtaining the threshold of each communication indicator to identify the express number. This method does not fully consider the content of communication indicator information, and the blacklist and whitelist call record table has certain limitations. Patent application title: A method, device and computer storage medium for identifying express numbers (CN110519466A), wherein the XGB model is used to train communication indicator information data for express number identification. Because the XGB model has many hyperparameters, it is easy to fall into a local optimal solution. When processing ultra-high dimensional feature variables, the performance will be reduced.

[0005] Therefore, effectively identifying express numbers from signaling call record data with high accuracy and high stability has become a technical problem that needs to be solved urgently. Summary of the invention

[0006] The technical problem to be solved by the present invention is to provide a highly accurate and stable automated intelligent express number identification method, which can provide correct identification and display of incoming call numbers for the majority of mobile users, so as to improve the connection rate and communication efficiency of express numbers and create a good communication network environment.

[0007] To solve the above technical problems, according to one aspect of the present invention, a method for express number recognition is provided. The method includes the following steps: S1. Input the black and white list data of express numbers and signaling call detail records data, and clean the data in the way of data ETL project. The data ETL project includes data extraction (Extract), transformation (Transform), and loading (Load), and then obtain the original data set required for the model through data association and fusion; S2. Use the SMOTE TomeK (synthetic sampling) algorithm to perform synthetic sampling on the original data set to form a model sample data set; S3. Divide the model sample data set into a training set and a test set according to the ratio of x:y for the training samples, where x ∈ [0,1], y ∈ [0,1], and x + y = 1. Respectively use the XGboost (Extreme gradient boosting) model, NGboost (Natural gradient boosting) model, and Catboost (Categorical gradient boosting) model for model training to form a basic classifier model as the primary model; S4. Use the five-fold cross-validation method to train the XGboost model, NGboost model, and Catboost model, use the test set for verification, and output the predicted value; S5. Respectively use the output value XG_Score of the XGboost model, the output value NG_Score of the NGboost model, and the output value Cat_Score of the Catboost model as the input features of the LightGBM (Light Gradient Boosting Machine) model to form a secondary model, perform training on the secondary model, and output a model that meets the pre-set model accuracy after training, thereby forming the XNCLBoost model (the acronym of the above four algorithms of XGboost, NGboost, Catboost, and LightGBM). Among them, the XNCLBoost model is fused based on the Stacking strategy. When there are many hyperparameters in the LightGBM model, the Gaussian Bayesian algorithm is used to optimize the value range of the hyperparameters to improve the prediction accuracy and robustness of the model. Among them, based on the Gaussian process improved Bayesian optimization algorithm, a covariance function in the form of a multi-strategy combination is used to simultaneously capture the smoothness and amplitude of the objective function.

[0008] According to an embodiment of the present invention, the method for express number recognition may further include the following steps: S6. Input the signaling call detail records data to be measured into the XNCLBoost model. After model prediction, output the prediction result, and eliminate the objection data according to the prediction result.

[0009] Further, the method for identifying the express delivery number may further include the following steps: S7. Apply the output model prediction data to the call business card and number identification service scenario to identify and label the incoming call number for the user; S8. Collect the data and complaint data feedback by the express delivery marking platform in the service scenario application, and update the express delivery black and white list data set with the data and complaint data feedback by the express delivery platform to form a closed-loop modeling process.

[0010] According to an embodiment of the present invention, the input express delivery black and white list data may be the existing number identification data of the platform, and can be optimized and updated by the prediction data calculated by the model and the data verified after platform application.

[0011] According to an embodiment of the present invention, the input signaling call record data may be the call record data to be measured (number encrypted). The call record is based on the call record data set corresponding to the number, and the call record data collection period is the recent one month, two months or three months. The specific data collection period is determined according to actual needs and is not limited thereto.

[0012] According to an embodiment of the present invention, call record data includes calling number call record data and called number call record data, and the call record data has characteristic variables, which may include: the calling number call area heat value, used to record the total number of cell / building positions in the call signaling when the number initiates an active call; the calling number outgoing frequency, used to record the total number of calls in the call signaling when the number initiates an active call; the calling number connection frequency, used to record the total number of successful connection times among the active call times in the call signaling when the number initiates an active call; the calling number average ringing duration, used to record the total ringing duration of calls in the call signaling when the number initiates an active call; the calling number average call duration, used to record the total normal call duration of calls in the call signaling when the number initiates an active call; the called number call area heat value, used to record the total number of cell / building positions in the call signaling when the number receives a passive call; the called number incoming frequency, used to record the total number of outgoing times in the call signaling when the number receives a passive call. The specific characteristic variables are not limited to this. For example, they can be: the calling contact dispersion (the total number of distinct called numbers in the call signaling when the number initiates an active call), the calling number active hours per day (counting the number of natural hours when the number initiates an active call daily based on call bill signaling data), the calling number active days per week (counting the number of natural days when the number initiates an active call weekly based on call bill signaling data), the calling number active days per month (counting the number of days when the number initiates an active call monthly based on call bill signaling data), the calling number monthly voice call fee consumption amount (counting the voice call fee amount of the number in the current month based on call bill signaling data), the calling number monthly data usage cost (counting the data usage cost amount of the number in the current month based on call bill signaling data), etc.; they can also be: the called number connection frequency (the total number of successful connection times among the passive received times in the call signaling when the number receives a passive call), the called number ringing duration (the total ringing duration of calls in the call signaling when the number receives a passive call), the called number call duration (the total normal call duration of calls in the call signaling when the number receives a passive call), the called contact dispersion (the total number of distinct calling party numbers in the call signaling when the number receives a passive call), etc.

[0013] According to a second aspect of the present invention, there is provided an express number recognition device, including: a model sample generation unit, which is used to process the input express number black and white list data and signaling call detail record data, and generate and output a model sample data set. The data processing methods include data cleaning through an ETL project and comprehensive sampling of the original data set using the SMOTETomeK algorithm; a primary model, which includes an XGboost model, an NGboost model, and a Catboost model. The first-layer model divides the model sample data set into a training set and a test set according to the ratio of x:y, where x ∈ [0,1], y ∈ [0,1], and x + y = 1. The XGboost model, the NGboost model, and the Catboost model are trained using the five-fold cross-validation method, and the trained data is then verified using the test set to output predicted values; a secondary model, which is constructed based on the LightGBM algorithm. After being trained by the primary model, the output values of the XGboost model XG_Score, the NGboost model NG_Score, and the Catboost model Cat_Score are respectively used as input features of the LightGBM model to form a secondary model, and the secondary model is trained. After training, a model that meets the pre-set model accuracy is output, thereby forming an XNCLBoost model. Among them, the XNCLBoost model is fused based on the Stacking strategy. When there are many hyperparameters in the LightGBM model, the Gaussian Bayesian algorithm is used to optimize the value range of the hyperparameters to improve the prediction accuracy and robustness of the model. Among them, based on the Gaussian process improved Bayesian optimization algorithm, a covariance function in the form of a multi-strategy combination is used to simultaneously capture the smoothness and amplitude of the objective function.

[0014] According to a third aspect of the present invention, there is provided an electronic device, including: a memory, a processor, and an express number recognition program stored on the memory and executable on the processor. When the express number recognition program is executed by the processor, the steps of the above-mentioned express number recognition method are implemented.

[0015] According to a fourth aspect of the present invention, there is provided a computer storage medium, on which an express number recognition program is stored. When the express number recognition program is executed by the processor, the steps of the above-mentioned express number recognition method are implemented.

[0016] Compared with the prior art, the technical solutions provided by the embodiments of the present invention can at least achieve the following beneficial effects:

[0017] 1. In the sampling of the XNCLBoost model sample data, the SMOTE TomeK algorithm is used to comprehensively sample the original data, effectively solving the problem of data imbalance.

[0018] 2. In the XNCLBoost model framework, tree models such as the XGboost model, the NGboost model, and the Catboost model are used as the first-layer base classifiers. At the same time, the advantages of each model are retained, effectively overcoming problems such as slow convergence of model parameters and overfitting of the model.

[0019] 3. In the XNCLBoost model framework, the LightGBM model is used as the second-layer meta-classifier to achieve the fusion of multiple tree models. This model effectively reduces the data dimension, removes redundant data, has good non-linear mapping ability, and has a relatively high classification prediction accuracy.

[0020] 4. In the optimization of LightGBM model parameters, the covariance function in the Gaussian process is used to simultaneously characterize the smoothness and amplitude of the objective function, forming an improved Gaussian Bayesian optimization algorithm.

[0021] 5. This solution is based on the Stacking strategy fusion model for express number recognition. The Stacking algorithm uses multi-fold cross-validation, which is more robust than using the hold-out set method.

[0022] 6. The present invention provides an automated intelligent recognition method for express numbers, which improves the connection rate and communication efficiency of express numbers, creates a good communication network environment, provides a high-accuracy and high-stability express number recognition method, and provides correct recognition and display of incoming call numbers for the majority of mobile users. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings of the embodiments will be briefly introduced below. Obviously, the drawings in the following description only relate to some embodiments of the present invention and do not limit the present invention.

[0024] Figure 1 It is a flowchart showing the express number recognition method according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings of the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0026] Unless otherwise defined, the technical terms or scientific terms used herein shall have the ordinary meanings as understood by those of ordinary skill in the art to which this invention pertains. The terms "first", "second" and similar terms used in the specification and claims of this patent application for invention do not denote any order, quantity or importance, but are merely used to distinguish different components. Similarly, terms such as "a" or "an" do not denote a quantity limitation, but mean that there is at least one.

[0027] Figure 1 is a flowchart showing the express number recognition method according to an embodiment of the present invention.

[0028] As Figure 1 shown, the method for recognizing an express number includes the following steps:

[0029] S1. Input the express number black and white list data and signaling call detail record data, and clean the data in the data ETL engineering manner. The data ETL engineering includes data extraction, transformation, and loading, and then obtain the original data set required for the model through data association and fusion;

[0030] S2. Use the SMOTE TomeK algorithm to perform comprehensive sampling on the original data set to form a model sample data set;

[0031] S3. Divide the model sample data set into a training set and a test set according to the ratio of x:y for the training samples, where x ∈ [0,1], y ∈ [0,1], and x + y = 1. Respectively use the XGboost model, NGboost model, and Catboost model for model training to form a basic classifier model as the primary model;

[0032] S4. Use the five-fold cross-validation method to train the XGboost model, NGboost model, and Catboost model, and use the test set for verification to output prediction values;

[0033] S5. Respectively use the output value XG_Score of the XGboost model, the output value NG_Score of the NGboost model, and the output value Cat_Score of the Catboost model as the input features of the LightGBM model to form a secondary model, and perform training on the secondary model. After training, output a model that meets the pre-set model accuracy, thereby forming the XNCLBoost model. Among them, the XNCLBoost model is based on the Stacking strategy for fusion. When there are many hyperparameters in the LightGBM model, the Gaussian Bayesian algorithm is used to optimize the value range of the hyperparameters to improve the prediction accuracy and robustness of the model. Among them, based on the Gaussian process improved Bayesian optimization algorithm, a covariance function in the form of a multi-strategy combination is used to simultaneously capture the smoothness and amplitude of the objective function.

[0034] The technical solution of the present invention adopts the SMOTE TomeK algorithm to perform comprehensive sampling of the original data in the sampling of sample data of the XNCLBoost model, which effectively solves the problem of data imbalance. In the XNCLBoost model framework, tree models such as the XGboost model, the NGboost model and the Catboost model are used as the first-layer base classifiers, while retaining the advantages of each model, effectively overcoming the problems of slow convergence of model parameters and model overfitting. In the XNCLBoost model framework, the LightGBM model is used as the second-layer meta-classifier to realize the fusion of multiple tree models. This model effectively reduces the data dimension, removes redundant data, has good nonlinear mapping capabilities, and has high classification prediction accuracy. In the optimization of LightGBM model parameters, the covariance function in the Gaussian process is used to simultaneously characterize the smoothness and amplitude of the objective function to form an improved Gaussian Bayesian optimization algorithm.

[0035] According to one or some embodiments of the present invention, the method for express number identification further includes the following steps:

[0036] S6. The signaling call record data to be tested is input into the XNCLBoost model. After prediction by the model, the prediction result is output, and objectionable data is eliminated based on the prediction result.

[0037] According to one or some embodiments of the present invention, the method for express number identification further includes the following steps:

[0038] S7, applying the output model prediction data to the caller name card and number recognition business scenario to identify and mark the caller number for the user;

[0039] S8. Collect the feedback data and complaint data from the express marking platform in business scenario applications, update the express blacklist and whitelist dataset with the feedback data and complaint data from the express platform, and form a closed loop of the modeling process.

[0040] This solution uses the Stacking strategy fusion model for express number recognition. The Stacking algorithm uses multi-fold cross-validation, which is more robust than the holdout set method. The Stacking algorithm method for model fusion refers to a technology that uses a model to combine other base models. First, multiple different primary learners need to be trained, and then the outputs of these models are used as input to train a new model to obtain the final model.

[0041] The basic idea of the Stacking algorithm is as follows:

[0042] Input: training set

[0043] D={(x1, y1), (x2, y2), ....., (xm , y m )}

[0044] Primary machine learning classification algorithms A1, A2,..., A n

[0045] Secondary machine learning classification algorithm A

[0046] Initialization parameters: various machine learning hyperparameters such as the K-fold cross-validation parameter K of the dataset, learning rate α, learning step size l, tree depth n, etc.

[0047] Steps:

[0048] Step1: for i = 1, 2,....., n do

[0049] Step 2: Out_A i -A i (D);

[0050] Step 3: end for

[0051] Step 4:

[0052] Step 5: for j = 1, 2,....., m do

[0053] Step 6: for i - 1, 2,....., n do

[0054] Step 7: z ji = Out_A i (x j );

[0055] Step 8: end for

[0056] Step 9: D″ = D′ ∪ ((z j1 , z j2 ,......, z jn ), y j );

[0057] Step 10: end for

[0058] Step 11: Out_A′A(D″);

[0059] Output:

[0060]

[0061] According to one or some embodiments of the present invention, the input express black and white list data is the existing number identification data of the platform, and can be optimized and updated by the predicted data calculated by the model and the data verified after the platform application.

[0062] According to one or some embodiments of the present invention, the input signaling call detail record data is the call detail record data to be measured. The call detail record is based on the call record data set corresponding to the number. The call record data collection period is the recent one month, the recent two months or the recent three months. The specific data collection period depends on actual needs and is not limited thereto.

[0063] According to one or some embodiments of the present invention, the call record data includes the calling number call record data and the called number call record data. The call record data has characteristic variables, and the characteristic variables may include: the calling number call area heat value, which is used to record the total number of cell / building positions in the call signaling when the number initiates an active call; the calling number outgoing frequency, which is used to record the total number of calls in the call signaling when the number initiates an active call; the calling number connection frequency, which is used to record the total number of successful connection times in the active calls in the call signaling when the number initiates an active call; the calling number average ringing duration, which is used to record the total ringing duration of the calls in the call signaling when the number initiates an active call; the calling number average call duration, which is used to record the total normal call duration of the calls in the call signaling when the number initiates an active call; the called number call area heat value, which is used to record the total number of cell / building positions in the call signaling when the number receives a passive call; the called number incoming frequency, which is used to record the total number of outgoing calls in the call signaling when the number receives a passive call. The specific characteristic variables are not limited thereto. For example, it can also be: the calling contact dispersion degree (the total number of distinct called numbers in the call signaling when the number initiates an active call), the calling number daily active hours (counting the number of natural hours when the number initiates an active call per day based on the call detail record signaling data), the calling number weekly active days (counting the number of natural days when the number initiates an active call per week based on the call detail record signaling data), the calling number monthly active days (counting the number of days when the number initiates an active call per month based on the call detail record signaling data), the calling number monthly voice call fee consumption amount (counting the voice call fee amount of the number in the current month based on the call detail record signaling data), the calling number monthly traffic cost (counting the traffic cost amount of the number in the current month based on the call detail record signaling data), etc.; it can also be: the called number connection frequency (the total number of successful connection times in the passive calls in the call signaling when the number receives a passive call), the called number ringing duration (the total ringing duration of the calls in the call signaling when the number receives a passive call), the called number call duration (the total normal call duration of the calls in the call signaling when the number receives a passive call), the called contact dispersion degree (the total number of distinct calling party numbers in the call signaling when the number receives a passive call), etc.

[0064] The call behavior of express delivery phone numbers often shows clustering in the spatial dimension. The called numbers dialed by express delivery numbers have a large dispersion, while the urban dispersion of the called numbers is low.

[0065] The call behavior of express delivery phone numbers often shows a dense call distribution within a short period in the time dimension. Generally, there are 5 - 6 dense calls within 5 minutes.

[0066] When the express delivery number A calls the called number B and the called number B is not answered, the express delivery number A will repeat the call to the called number B within a short period, with an average interval of nearly 75 seconds.

[0067] When the express delivery number A calls the called number B and the called number B is not answered, the express delivery number A will send a short message to the called number B within a short period.

[0068] Calling number repeated call index 1: When the express delivery number A calls the called number B and the called number B is not answered, the express delivery number A will repeat the call to the called number B within a short period, and it is the total number of calls between the express delivery number A and the called number B before connection.

[0069] Calling number repeated call index 2: When the express delivery number A calls the called number B and the called number B is not answered, the express delivery number A will repeat the call to the called number B within a short period, and it is the total number of calls between the express delivery number A and the called number B before a successful short message is sent.

[0070] The present invention provides an automated intelligent recognition method for express delivery numbers, which improves the connection rate and communication efficiency of express delivery numbers, creates a good communication network environment, provides a highly accurate and stable express delivery number recognition method, and provides correct identification and display of incoming call numbers for the majority of mobile users.

[0071] According to another aspect of the present invention, there is provided an express number recognition device, including: a model sample generation unit, which is used to process the input express number black and white list data and signaling call detail records data, and generate and output a model sample data set. Among them, the data processing methods include data ETL project for data cleaning and SMOTETomeK algorithm for comprehensive sampling of the original data set; a primary model, which has an XGboost model, an NGboost model, and a Catboost model. The first-layer model divides the model sample data set into a training set and a test set according to the ratio of x:y of the training samples, where x ∈ [0,1], y ∈ [0,1], and x + y = 1. Among them, the XGboost model, the NGboost model, and the Catboost model are trained using the five-fold cross-validation method, and the trained data is then verified using the test set to output predicted values; a secondary model, which is constructed based on the LightGBM algorithm. After being trained by the primary model, the output value XG_Score of the XGboost model, the output value NG_Score of the NGboost model, and the output value Cat_Score of the Catboost model are used as the input features of the LightGBM model respectively to form a secondary model for training the secondary model. After training, a model that meets the pre-set model accuracy is output, thereby forming an XNCLBoost model. Among them, the XNCLBoost model is fused based on the Stacking strategy. When there are many hyperparameters in the LightGBM model, the Gaussian Bayesian algorithm is used to optimize the value range of the hyperparameters to improve the prediction accuracy and robustness of the model. Among them, based on the Gaussian process improved Bayesian optimization algorithm, a covariance function in the form of a multi-strategy combination is used to capture the smoothness and amplitude of the objective function at the same time.

[0072] According to still another aspect of the present invention, there is provided an express number recognition device, including: a memory, a processor, and an express number recognition program stored on the memory and executable on the processor. When the express number recognition program is executed by the processor, the steps of the above-mentioned express number recognition method are implemented.

[0073] The present invention also provides a computer storage medium.

[0074] An express number recognition program is stored on the computer storage medium. When the express number recognition program is executed by the processor, the steps of the above-mentioned express number recognition method are implemented.

[0075] Among them, the method implemented when the express number recognition program running on the processor is executed can refer to the respective embodiments of the express number recognition method of the present invention, which will not be elaborated here.

[0076] The present invention also provides a computer program product.

[0077] The computer program product of the present invention includes an express number recognition program, and when the express number recognition program is executed by a processor, the steps of the express number recognition method described above are implemented.

[0078] Among them, the method implemented when the express number recognition program running on the processor is executed can refer to each embodiment of the express number recognition method of the present invention, which will not be elaborated here.

[0079] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) as described above, and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present invention.

[0080] The above is only an exemplary embodiment of the present invention and is not used to limit the protection scope of the present invention. The protection scope of the present invention is determined by the appended claims.

Claims

1. A method for identifying express delivery numbers, the method comprising the following steps: S1. Input the express delivery number black and white list data and signaling call detail record data, and clean the data in the data ETL engineering mode. The data ETL engineering includes data extraction, transformation, and loading, and then obtain the original data set required by the model through data association and fusion; S2. Use the SMOTE TomeK algorithm to comprehensively sample the original data set to form a model sample data set; S3. Divide the model sample dataset into a training set and a test set according to the x:y ratio of the training samples, where x∈[0,1], y∈[0,1], x + y = 1. Respectively use the XGboost model, NGboost model, and Catboost model for model training to form a basic classifier model as the primary model; S4. Use the five-fold cross-validation method to train the XGboost model, NGboost model, and Catboost model, use the test set for verification, and output the prediction value; S5. Respectively use the output value XG_Score of the XGboost model, the output value NG_Score of the NGboost model, and the output value Cat_Score of the Catboost model as the input features of the LightGBM model to form a secondary model for training the secondary model. After training, output a model that meets the pre-set model accuracy, thereby forming the XNCLBoost model, wherein, the XNCLBoost model is fused based on the Stacking strategy, wherein, when there are many hyperparameters in the LightGBM model, the Gaussian Bayesian algorithm is used to optimize the value range of the hyperparameters to improve the prediction accuracy and robustness of the model, wherein, based on the Gaussian process improved Bayesian optimization algorithm, a covariance function in the form of a multi-strategy combination is used to simultaneously capture the smoothness and amplitude of the objective function; S6. Input the signaling call detail record data to be measured into the XNCLBoost model. After model prediction, output the prediction result, and eliminate the objection data according to the prediction result; S7. Apply the output model prediction data to the incoming call business card and number identification service scenarios to identify and label the incoming call numbers for users.

2. The method according to claim 1, the method further comprising the following steps: S8. Collect the data and complaint data feedback by the express delivery marking platform in the application of the service scenario, and update the express delivery black and white list data set with the express delivery platform feedback data and complaint data to form a closed-loop modeling process.

3. The method according to claim 2, wherein The input express delivery black and white list data is the existing number identification data of the platform, and can be optimized and updated by the prediction data calculated by the model and the data verified after platform application.

4. The method according to claim 1, wherein The input signaling call detail record data is the call detail record data to be measured. The call detail record is based on the call record data set corresponding to the number, and the call record data collection period is the recent one month, recent two months, or recent three months.

5. The method according to claim 4, wherein, The call record data includes the calling number call record data and the called number call record data, and the call record data has characteristic variables, and the characteristic variables include: The heat value of the calling number call area is used to record the total number of cell / building positions in the call signaling when the number initiates an active call; The calling number outgoing call frequency is used to record the total number of calls in the call signaling when the number initiates an active call; The calling number connection frequency is used to record the total number of successful connections among the number of active calls in the call signaling when the number initiates an active call; The average ringing duration of the calling number is used to record the total ringing duration of the call in the call signaling when the number initiates an active call; The average call duration of the calling number is used to record the total normal call duration in the call signaling when the number initiates an active call; The called number call area thermal value is used to record the total number of cell / building locations in the call signaling when the number receives a passive call; The called number's incoming call frequency is used to record the total number of outgoing calls in the call signaling when the number accepts passive calls.

6. A device for identifying a courier number, comprising: A model sample generating unit, the model sample generating unit is used to process the input express number blacklist and whitelist data and signaling call record data, generate and output a model sample data set, wherein the data processing method includes data ETL engineering cleaning data and SMOTE TomeK algorithm for comprehensive sampling of the original data set; The primary model includes an XGboost model, an NGboost model, and a Catboost model. The first layer model divides the model sample data set into a training set and a test set according to the ratio of x:y, where x∈[0,1], y∈[0,1], and x+y=1. The XGboost model, the NGboost model, and the Catboost model are trained using a five-fold cross-validation method. The trained data is then validated using a test set to output a predicted value. The secondary model is constructed based on the LightGBM algorithm. After the primary model is trained, the output value XG_Score of the XGboost model, the output value NG_Score of the NGboost model, and the output value Cat_Score of the Catboost model are used as the input features of the LightGBM model to form a secondary model. The secondary model is trained and a model that meets the pre-set model accuracy is output after training, thereby forming an XNCLBoost model. The XNCLBoost model is based on Stacking strategy fusion. When the LightGBM model has many hyperparameters, the Gaussian Bayesian algorithm is used to optimize the value range of the hyperparameters to improve the prediction accuracy and robustness of the model. Among them, the Bayesian optimization algorithm is improved based on Gaussian process, and the covariance function in the form of multi-strategy combination is used to capture the smoothness and amplitude of the objective function at the same time; Among them, the signaling call record data to be tested is input into the XNCLBoost model, and after model prediction, the prediction result is output, and objectionable data is eliminated based on the prediction result; the output model prediction data is applied to the caller name card and code number recognition business scenario to identify and mark the incoming call number for the user.

7. An electronic device, comprising: A memory, a processor, and an express number identification program stored on the memory and executable on the processor. When the express number identification program is executed by the processor, the steps of the method for identifying an express number as described in any one of claims 1 to 5 are implemented.

8. A computer storage medium, wherein, An express number identification program is stored on the computer storage medium. When the express number identification program is executed by a processor, the steps of the method for identifying an express number as described in any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • An express number recognition method based on call behavior

    CN109274834A

  • Express number identification method and device, and computer storage medium

    CN110519466A

  • Prediction device for personalized COH (controlled ovarian hyperstimulation) scheme based on machine learning

    CN111145912A

  • Fan spindle fault prediction method based on sliding window characteristics

    CN113392575A