A method, device, system and electronic device for automatically updating a model

By constructing the initial model and judging its evaluation indicators, the update features suitable for new sample data are solved, and the prediction accuracy of the model is improved.

CN113011596BActive Publication Date: 2025-07-08阳光保险集团股份有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110193968.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-20
Publication Date
2025-07-08
Estimated Expiration
2041-02-20

AI Technical Summary

Technical Problem

In the prior art, the feature variables suitable for new sample data cannot be selected during the model update, resulting in inaccurate prediction results of the updated model.

Method used

By constructing an initial model, it is determined whether its evaluation indicators meet the predetermined requirements, obtain the second time series data, determine multiple update features, and update the initial model based on these features.

Benefits of technology

Ensure that the updated model feature variable is suitable for new sample data, improves the prediction accuracy of the model and solves the problem of inaccurate prediction results caused by the inapplicability of feature variables.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113011596B_ABST
    Figure CN113011596B_ABST
Patent Text Reader

Abstract

The present application discloses a method, apparatus, system and electronic device for automatically updating a model. The method includes: constructing an initial model based on multiple initial features of first time-series data within a first predetermined time period; determining whether a model evaluation index of the constructed initial model meets a predetermined requirement; if the model evaluation index of the initial model meets the predetermined requirement, acquiring second time-series data within a second predetermined time period; determining multiple updated features of the acquired second time-series data; and updating the initial model based on the multiple updated features. By adopting the above method, apparatus, system and electronic device for automatically updating a model, the problem that the model prediction accuracy is reduced due to the change of feature distribution in the case of a long time span is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of big data technology, and in particular, to a method, device, system and electronic device for automatically updating a model. Background Art

[0002] With the continuous development of big data technology, machine learning models have been widely used in various industries. Comprehensively understanding the overview of data and exploring the characteristics of each variable are important links in building a machine learning model. In a machine learning task, the selection of features usually determines the upper limit of the model effect. Good feature selection can not only prevent the curse of dimensionality, reduce the training time, but also enhance the generalization ability of the model and reduce overfitting. Therefore, optimizing the selection of feature variables is of great significance for building a machine learning model. In many business scenarios, due to the limited amount of data used by the machine learning model, it is usually necessary to select data with a time span of more than two years. However, as time goes by, the new sample data has changed greatly compared with the sample data used when building the model, which leads to the gradual deterioration of the model prediction effect. Therefore, the model needs to be updated regularly. In the prior art, the model is usually updated by directly inputting the newly selected sample data into the model.

[0003] In the above-mentioned existing model update method, the updated model still uses the feature variables used when building the model, but the feature variables used when building the model are no longer applicable to the new sample data, which leads to the problem that the prediction result of the updated model is inaccurate. Summary of the Invention

[0004] In view of this, the present application provides a method, device, system and electronic device for automatically updating a model, the purpose of which is to select feature variables applicable to new sample data when updating the model, and avoid the problem that the prediction result of the updated model is inaccurate due to the inapplicability of the feature variables to the new sample data.

[0005] In a first aspect, an embodiment of the present application provides a method for automatically updating a model, including:

[0006] Constructing an initial model based on multiple initial features of first time-series data within a first predetermined time period;

[0007] Judging whether the model evaluation index of the constructed initial model meets a predetermined requirement;

[0008] If the model evaluation index of the initial model meets the predetermined requirement, obtaining second time-series data within a second predetermined time period;

[0009] Determining multiple updated features of the obtained second time-series data;

[0010] Update the initial model based on multiple update features.

[0011] Optionally, determining multiple update features of the obtained second time-series data may include: (A) determining multiple preset time points within a second predetermined time period; (B) for each preset time point, determining a sample set corresponding to the preset time point, where the sample set includes training data before the preset time point and test data after the preset time point obtained by dividing the second time-series data according to the preset time point; (C) determining multiple candidate feature groups of the second time-series data, with each candidate feature group including at least one candidate feature; (D) for each candidate feature group, constructing multiple classifiers for the candidate feature using the determined sample set; (E) for each candidate feature group, determining a feature evaluation index of the candidate feature group under each classifier; (F) based on the determined feature evaluation indexes, determining multiple update features from multiple candidate feature groups.

[0012] Optionally, determining multiple update features from multiple candidate feature groups based on the determined feature evaluation indexes may include: (F1) for each candidate feature group, determining a statistical value of the feature evaluation index of the candidate feature group under each classifier; (F2) based on the statistical values corresponding to all candidate feature groups, selecting a target candidate feature group from all candidate feature groups; (F3) determining whether the number of candidate features in the target candidate feature group reaches a predetermined value; (F4) if the number of candidate features in the target candidate feature group does not reach the predetermined value, constructing a new candidate feature group based on the candidate features in the target candidate feature group, updating the candidate feature group with the constructed new candidate feature group, and returning to execute step (D); (F5) if the number of candidate features in the target candidate feature group reaches the predetermined value, determining the candidate features in the target candidate feature group as multiple update features.

[0013] Optionally, the new candidate feature group may include multiple new candidate feature groups, each new candidate feature group may include the candidate features in the target candidate feature group and one other candidate feature, and the other candidate feature is a feature among the multiple candidate features of the second time-series data other than the candidate features in the target candidate feature group.

[0014] Optionally, the model evaluation metrics include a model stability metric and a model prediction effect metric. Among them, determining whether the model evaluation metrics of the constructed initial model meet the predetermined requirements may include: determining whether the model stability metric of the initial model is greater than a first set threshold, and determining whether the model prediction effect metric of the initial model is less than a second set threshold; if the model stability metric of the initial model is greater than the first set threshold, and / or, the model prediction effect metric of the initial model is less than the second set threshold, then it is determined that the model evaluation metrics of the initial model meet the predetermined requirements; if the model stability metric of the initial model is not greater than the first set threshold, and the model prediction effect metric of the initial model is not less than the second set threshold, then it is determined that the model evaluation metrics of the initial model do not meet the predetermined requirements.

[0015] In a second aspect, an embodiment of the present application provides a model automatic update device, including:

[0016] A construction module, which constructs an initial model based on multiple initial features of first time-series data within a first predetermined time period;

[0017] A judgment module, which judges whether the model evaluation metrics of the constructed initial model meet the predetermined requirements;

[0018] An acquisition module, if the model evaluation metrics of the initial model meet the predetermined requirements, then acquires second time-series data within a second predetermined time period;

[0019] A determination module, which determines multiple update features of the acquired second time-series data;

[0020] An update module, which updates the initial model based on the multiple update features.

[0021] Optionally, the determination module may determine the multiple update features in the following manner: determine multiple preset time points within the second predetermined time period; for each preset time point, determine a sample set corresponding to the preset time point, where the sample set includes training data before the preset time point and test data after the preset time point obtained by dividing the second time-series data according to the preset time point; determine multiple candidate feature groups of the second time-series data, and each candidate feature group includes at least one candidate feature; for each candidate feature group, construct multiple classifiers for the candidate feature by using the determined sample set; for each candidate feature group, determine the feature evaluation metrics of the candidate feature group under each classifier; based on the determined feature evaluation metrics, determine multiple update features from the multiple candidate feature groups.

[0022] In a third aspect, an embodiment of the present application provides a model automatic update system, including: a background management server, a remote server;

[0023] The background management server obtains the first time-series data within the first predetermined time period from the remote server, and constructs an initial model based on multiple initial features of the obtained first time-series data;

[0024] The background management server determines whether the model evaluation index of the constructed initial model meets the predetermined requirements. If the model evaluation index of the initial model meets the predetermined requirements, it obtains the second time-series data within the second predetermined time period from the remote server;

[0025] The background management server determines multiple updated features of the obtained second time-series data, and updates the initial model based on the multiple updated features.

[0026] In a fourth aspect, an embodiment of the present application provides an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps of the above-mentioned model automatic update method are executed.

[0027] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is run by a processor, the steps of the above-mentioned model automatic update method are executed.

[0028] The embodiments of the present application bring the following beneficial effects:

[0029] The embodiments of the present application provide a model automatic update method, including: constructing an initial model based on multiple initial features of the first time-series data within the first predetermined time period; determining whether the model evaluation index of the constructed initial model meets the predetermined requirements; if the model evaluation index of the initial model meets the predetermined requirements, obtaining the second time-series data within the second predetermined time period; determining multiple updated features of the obtained second time-series data; and updating the initial model based on the multiple updated features. In the present application, when updating the model, feature variables applicable to the new sample data are selected, avoiding the problem that the prediction result of the updated model is inaccurate due to the feature variables not being applicable to the new sample data.

[0030] To make the above objects, features, and advantages of the present application more obvious and understandable, the following preferred embodiments are specifically described below in conjunction with the accompanying drawings. Description of the Drawings

[0031] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the embodiments. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other relevant drawings can also be obtained based on these drawings.

[0032] Figure 1 It is a schematic flowchart of the model automatic update method provided by the embodiments of the present application.

[0033] Figure 2 It is a schematic flowchart of the steps for determining multiple update features provided by the embodiments of the present application;

[0034] Figure 3 It is a schematic structural diagram of the model automatic update device provided by the embodiments of the present application;

[0035] Figure 4 It is a schematic structural diagram of the model automatic update system provided by the embodiments of the present application;

[0036] Figure 5 It is a schematic structural diagram of an electronic device provided by the embodiments of the present application. Specific embodiments

[0037] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions of the present application in conjunction with the accompanying drawings. Obviously, the described embodiments are some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0038] In the prior art, the model is usually updated by directly inputting newly selected sample data into the model. In the above-mentioned existing model update method, the updated model still uses the feature variables used when initially constructing the model. As time goes by, these feature variables are no longer applicable to the new sample data. If the model is still updated based on these feature variables, the prediction results of the updated model will be inaccurate. For example, in the model for studying the surrender rate of an insurance institution, due to the non-standard operation of a certain institution in the early stage, the surrender rate is relatively high. At this time, it is appropriate to select the variable "institution" as an important feature. However, after a period of rectification by the institution, the surrender rate has significantly decreased. If the variable "institution" is still used as an important feature at this time, then these important features do not match the new sample data, resulting in inaccurate prediction results of the updated model.

[0039] Based on this, the embodiments of the present application provide a model automatic update method, apparatus, system and electronic device, which selects feature variables applicable to new sample data when updating the model, avoiding the problem that the prediction result of the updated model is inaccurate due to the inapplicability of the feature variables to the new sample data.

[0040] For the convenience of understanding this embodiment, first, a model automatic update method disclosed in the embodiments of the present application will be introduced in detail. Figure 1 The flowchart of the model automatic update method provided by the embodiments of the present application is shown as Figure 1 shown, and the method includes the following steps:

[0041] Step 101, construct an initial model based on multiple initial features of the first time-series data within the first predetermined time period.

[0042] Specifically, the first predetermined time period can be a time interval selected from the time range corresponding to all the original data, and the original data corresponding to the selected time interval is used as the first time-series data. In one example, the first predetermined time period may refer to a time interval that is before the current time and adjacent to the current time. For example, the most recent 6 months. At this time, the first time-series data may refer to the time-series data of the most recent 6 months.

[0043] In this step, the first time-series data can be obtained from multiple data owners. The multiple data owners may include, but are not limited to, at least one of the following: Internet companies, e-commerce companies, express delivery companies, banks, insurance institutions. As an example, the first time-series data may include, but is not limited to, at least one of the following: the Internet surfing behavior data of users on the Internet (such as, website browsing records, search records), the electronic transaction data of users on the e-commerce platform (such as, order data, product browsing information, payment data), the express delivery logistics data of users, the business processing data of users in banks, the insurance policy data of users.

[0044] Here, the time-series data may refer to time series data, that is, a data column recorded in chronological order. As time goes by, the time-series data is constantly changing. Correspondingly, the features of the time-series data are also changing. This results in a large difference between the new time-series data and the original time-series data after a period of time, and there are also differences in their respective corresponding features. Therefore, it is necessary to update the initial model constructed based on multiple initial features to ensure the prediction effect of the model.

[0045] Constructing the initial model may include the following steps: data preprocessing, feature selection, model training, model deployment.

[0046] Data preprocessing may include the following steps: data integration, target variable annotation, feature variable derivation, and data cleaning. Among them, data integration requires collecting and integrating data from different data sources; target variable annotation and feature variable derivation are to process the integrated original data, which may include operations such as transforming and aggregating the original data to achieve feature data extraction; data cleaning may include deduplicating the original data, filling in missing values, and deleting outliers.

[0047] In addition, multiple initial features of the first time-series data can be determined by performing feature extraction on the first time-series data.

[0048] In the embodiments of the present application, feature selection can adopt the feature acquisition method used in step 104 below, that is, the same feature acquisition method as used to determine multiple updated features of the second time-series data can be used to determine multiple initial features of the first time-series data. The specific feature acquisition process will be described in step 104 later, and the present invention will not elaborate on this. In addition, the present application is not limited to this, and other feature extraction methods can also be used in step 101 to determine multiple initial features of the first time-series data.

[0049] Model training may include dividing the training set and the test set, and model selection. Among them, dividing the training set and the test set can adopt an experimental evaluation method of dividing the training set and the test set according to the time dimension; the selectable models include but are not limited to the logistic regression model, the LightGBM model, the decision tree model, the GDBT model, and the neural network model. Those skilled in the art can determine the selected model according to actual needs, and the present application does not make any limitations on this.

[0050] Model deployment includes but is not limited to exporting the model as a file in a specified format, deploying the model file to the management platform, and managing the model version.

[0051] In the present application, after the initial model is constructed, the initial model can be put into a test environment (such as the Bate environment) for testing based on the test set. After the test passes, the initial model can be launched into the production environment for formal operation, and the test environment can be isolated from the production environment.

[0052] In one example, assume that the first time-series data is the user's insurance policy data, the initial model is a model for predicting the user's renewal rate (i.e., the additional insurance rate), the first predetermined time period is 24 months from T - 24 months to T months, and T month refers to the month corresponding to the current time. Then the original data corresponding to these 24 months is the first time-series data. In the above business scenario, the first time-series data may include insurance policy data, user behavior data, claim data, and complaint data. The target variable (i.e., the output of the initial model) may include marking customers who have purchased an insurance policy before a certain time point and then purchased other insurance policies as "yes", and marking customers who have purchased an insurance policy before this time point but have not purchased other insurance policies so far as "no". The multiple initial features may include the number of insurance policies purchased in the past and the total premium paid in the past.

[0053] Taking the above example as an example, the training set and test set can be divided as follows. From the first time-series data, select the data for three months during the period from T month to T - 3 months as the test set of the initial model to evaluate the prediction effect of the initial model. From the first time-series data, select the data for twenty-one months during the period from T - 24 months to T - 3 months as the training set of the model to construct the initial model.

[0054] In one embodiment, assume that the initial model selected is the LightGBM model. Then the process of model deployment can be as follows: first, export the model in json format, parse it into rules and model files, and place them together in the automatic deployment management platform. At the same time, store the multiple initial features developed in the database. Pass the sample to be predicted in json format into the automatic deployment management platform, and running the model can return the prediction result.

[0055] Step 102, determine whether the model evaluation index of the constructed initial model meets the predetermined requirements.

[0056] For example, after the initial model is launched into the production environment for operation, the above judgment can be periodically executed at a predetermined time interval. For example, the duration of the model monitoring period can be determined as the duration of the predetermined time interval. When the online operation time of the initial model reaches the duration of the model monitoring period, determine the model evaluation index of the initial model, and determine whether the model evaluation index of the initial model meets the predetermined requirements. When the online operation time of the initial model does not reach the duration of the model monitoring period, do not execute step 102. Here, those skilled in the art can set the duration of the model monitoring period based on experience. In addition, the duration of the model monitoring period can also be adjusted according to the update requirements of the model. For example, if a higher prediction accuracy of the model is required, the duration of the model monitoring period can be shortened; if a lower prediction accuracy of the model is required, the duration of the model monitoring period can be extended.

[0057] In this step, the time-series data within the model monitoring period can be obtained, and the obtained time-series data can be used to determine the model evaluation indicators of the initial model. Here, the data producer of the time-series data within the model monitoring period is the same as that of the first time-series data. For example, taking the first time-series data as the Internet access behavior data of different users within the first predetermined time period, the time-series data within the model monitoring period can refer to the Internet access behavior data generated by different users within the model monitoring period.

[0058] Specifically, after the model is deployed, it can be monitored to be used as the basis for judging whether to update the model. Among them, the monitoring results can be evaluated through the model evaluation indicators, and the model evaluation indicators can include the model stability index and the model prediction effect index. In one example, it is judged whether the model stability index of the initial model is greater than the first set threshold, and it is judged whether the model prediction effect index of the initial model is less than the second set threshold. Here, the sizes of the first set threshold and the second set threshold can be set based on experience or other means.

[0059] If the model stability index of the initial model is greater than the first set threshold, and / or, the model prediction effect index of the initial model is less than the second set threshold, it is determined that the model evaluation indicators of the initial model meet the predetermined requirements; if the model stability index of the initial model is not greater than the first set threshold, and the model prediction effect index of the initial model is not less than the second set threshold, it is determined that the model evaluation indicators of the initial model do not meet the predetermined requirements. Among them, the first set threshold and the second set threshold can be set according to the business scenario, the index type or experience.

[0060] In the embodiments of the present application, the model stability index may include, but is not limited to, PSI (population stability index); the model prediction effect index may include, but is not limited to, accuracy, precision, and recall. Those skilled in the art can select appropriate evaluation indicators according to actual needs, and the present application does not make any limitations in this regard.

[0061] In a specific example, the monitoring of the model stability index is based on the prediction results corresponding to the training set data, and compares with the prediction results corresponding to the test set data within the model stability monitoring period. When the PSI value is greater than the first set threshold (such as 25%), it indicates that there is a significant change in the predicted score (such as the score of the probability of policy purchase) or the distribution of features, and then the model is updated. The monitoring of the model prediction effect index is to label the prediction samples after the end of the model effect monitoring period, and calculate the precision rate of the model effect index. If the precision rate is lower than the first set threshold (such as 70%), the model is updated. Among them, both the model stability monitoring period and the model effect monitoring period are time intervals preset according to experience, such as 3 months. The model stability monitoring period and the model effect monitoring period can take the same time interval or different time intervals. If one of the above model stability index and model prediction effect index meets the predetermined requirements, it is regarded that the model evaluation index meets the predetermined requirements to update the initial model; when both the model stability index and the model effect index do not meet the predetermined requirements, it is regarded that the model evaluation index does not meet the predetermined requirements, and there is no need to update the initial model.

[0062] If the model evaluation index of the initial model does not meet the predetermined requirements, there is no need to update the initial model. At this time, return to execute step 102, and continue to judge the initial model based on the first time-series data within the first predetermined time period. Here, taking the first predetermined time period as the most recent 6 months as an example, as time goes by, the first time-series data corresponding to the first predetermined time period is also constantly updated. When it is determined that the model evaluation index of the initial model does not meet the predetermined requirements, use the updated first time-series data to continue the judgment.

[0063] If the model evaluation index of the initial model meets the predetermined requirements, execute step 103: Obtain the second time-series data within the second predetermined time period.

[0064] Step 103, if the model evaluation index of the initial model meets the predetermined requirements, obtain the second time-series data within the second predetermined time period.

[0065] Specifically, the second predetermined time period can be a preset time interval, and the length of its time interval can be the same as that of the first predetermined time period or different from that of the first predetermined time period. The second predetermined time period and the first predetermined time period can be completely non-overlapping or partially overlapping. In an example, the second predetermined time period can refer to a time interval that is before the current time and adjacent to the current time. It should be understood that the current time here refers to the moment when it is determined that the model evaluation index of the initial model meets the predetermined requirements, while the current time in step 101 refers to the moment when the initial model is constructed.

[0066] In one example, for the case where the second predetermined time period partially overlaps with the first predetermined time period, the second predetermined time period may be composed of two sub-time periods, namely the model monitoring cycle time period and the overlapping time period. The model monitoring cycle is the model monitoring cycle determined according to experience in step 102 above, and the overlapping time period is the time period during which the first predetermined time period overlaps with the second predetermined time period. The data formed by combining the data corresponding to the model monitoring cycle time period and the overlapping time period in chronological order is determined as the second time series data.

[0067] Taking the example in step 102 above, assume that the model monitoring cycle is 3 months, the first predetermined time period is from T - 24 months to T months, and the duration of the second predetermined time period is 24 months. Then, the model monitoring cycle time period included in the second predetermined time period is from T months to T + 3 months, and the overlapping time period is from T - 21 months to T months. The data corresponding to the time period from T - 21 months to T + 3 months is determined as the second time series data.

[0068] Step 104, determine multiple update features of the obtained second time series data.

[0069] Figure 2 As shown in the flowchart of the step for determining multiple update features provided in the embodiments of the present application, when executing step 104, Figure 2 as shown, it specifically includes the following steps:

[0070] Step 201, determine multiple preset time points within the second predetermined time period.

[0071] Specifically, take multiple preset time points within the second predetermined time period. For each preset time point, the second predetermined time period can be divided into two time intervals according to this preset time point. Multiple preset time points correspond to multiple second predetermined time periods each divided into two time intervals.

[0072] In the above example, assume that the number of preset time points is 4. In the time period from T - 21 months to T + 3 months in step 103 above, the preset time points are T - 6 months, T - 4 months, T - 2 months, and T months respectively. Then, 4 second predetermined time periods with different segmentation points can be formed, namely the second predetermined time period composed of two time intervals [T - 21 months, T - 6 months] and [T - 6 months, T + 3 months]; the second predetermined time period composed of two time intervals [T - 21 months, T - 4 months] and [T - 4 months, T + 3 months]; the second predetermined time period composed of two time intervals [T - 21 months, T - 2 months] and [T - 2 months, T + 3 months]; the second predetermined time period composed of two time intervals [T - 21 months, T months] and [T months, T + 3 months].

[0073] Step 202: For each preset time point, determine the sample set corresponding to this preset time point.

[0074] Specifically, for each preset time point, divide the second time series data according to this preset time point into training data before this preset time point and test data after this preset time point, and the sample set consists of this training data and this test data. That is to say, after dividing the second predetermined time period into two time intervals according to this preset time point, determine the second time series data corresponding to the time interval before this preset time point as training data, and determine the second time series data corresponding to the time interval after this preset time point as test data. Since the position of each preset time point in the second predetermined time period is different, multiple different combinations of training data and test data will be obtained.

[0075] Taking the example of Step 201 above, the following 4 different sample sets can be determined: Sample set 1 includes the training data corresponding to the time period [T - 21 months, T - 6 months] and the test data corresponding to the time period [T - 6 months, T + 3 months]; Sample set 2 includes the training data corresponding to the time period [T - 21 months, T - 4 months] and the test data corresponding to the time period [T - 4 months, T + 3 months]; Sample set 3 includes the training data corresponding to the time period [T - 21 months, T - 2 months] and the test data corresponding to the time period [T - 2 months, T + 3 months]; Sample set 4 includes the training data corresponding to the time period [T - 21 months, T months] and the test data corresponding to the time period [T months, T + 3 months].

[0076] Step 203: Determine multiple candidate feature groups of the second time series data. Here, each candidate feature group may include at least one candidate feature.

[0077] Specifically, after determining the second time series data, multiple candidate feature groups can be determined from the second time series data. Here, various feature extraction methods can be used to determine multiple candidate features corresponding to the second time series data, and the candidate features obtained by feature extraction are used to form multiple candidate feature groups. It should be understood that various combination methods can be used to combine the candidate features obtained by feature extraction to obtain candidate feature groups. In a preferred example, each candidate feature can be used as a candidate feature group, that is, a candidate feature group contains one candidate feature.

[0078] Taking the above example as an example, assuming that 10 candidate features are determined for the second time series data, each of the 10 determined candidate features can be independently formed into 1 candidate feature group, that is, a total of 10 candidate feature groups are formed, which are candidate feature group 1 to candidate feature group 10 respectively.

[0079] Step 204: For each candidate feature group, construct multiple classifiers for the candidate features using the determined sample set.

[0080] Specifically, for each candidate feature group, use the determined sample set to construct multiple classifiers using the construction method of the initial model. The multiple classifiers constructed can be of the same type as the initial model. For example, they can all be GDBT models. However, this application is not limited to this, and multiple classifiers can also be constructed in other ways, and the types of the constructed classifiers can also be different from the type of the initial model. Here, the model parameters of each classifier are different, and the sample sets for each classifier are also different.

[0081] Taking the example in Step 102 above, for each candidate feature group, 4 LightGBM classifiers can be constructed using all 4 sample sets. That is to say, one candidate feature group corresponds to multiple classifiers, and one sample set is used to train one classifier for this candidate feature group.

[0082] Step 205: For each candidate feature group, determine the feature evaluation metrics of the candidate feature group under each classifier.

[0083] Specifically, for each candidate feature group, the candidate features in the candidate feature group can be respectively substituted into each classifier to calculate the feature evaluation metrics of the candidate feature group under each classifier. As an example, the feature evaluation metric can be the AUC value. It should be understood that those skilled in the art can select appropriate evaluation metrics according to actual needs, and this application does not make any limitations in this regard.

[0084] Taking the example in Step 203 above, assume that the candidate feature groups corresponding to the second time-series data are candidate feature groups 1 to 10, and candidate feature group 1 contains candidate feature 1, candidate feature group 2 contains candidate feature 2, and so on. For each candidate feature group, substitute the candidate feature group into 4 classifiers respectively, and 4 feature evaluation metric values corresponding to the candidate feature group will be obtained. These 4 feature evaluation metric values form an evaluation metric group, and a total of 10 evaluation metric groups corresponding to 10 candidate feature groups can be obtained.

[0085] After that, multiple updated features can be determined from multiple candidate feature groups based on the determined feature evaluation metrics.

[0086] Specifically, Step 206: For each candidate feature group, determine the statistical value of the feature evaluation metrics of the candidate feature group under each classifier.

[0087] As an example, for each candidate feature group, the statistical value of the feature evaluation index corresponding to the candidate feature group can be the average value, median value, maximum value, or minimum value of each feature evaluation index in the evaluation index group corresponding to the candidate feature group in step 205 above. Those skilled in the art can make a choice according to actual needs, and the present application does not make any limitations here.

[0088] In this example, assume that for each candidate feature group, the average value is selected as the statistical value of the feature evaluation index corresponding to the candidate feature group. For example, the 10 statistical values corresponding to candidate feature groups 1 to 10 are 0.65, 0.3, 0.7, 0.56, 0.68, 0.28, 0.15, 0.25, 0.43, and 0.33 respectively.

[0089] Step 207, based on the statistical values corresponding to all candidate feature groups, select the target candidate feature group from all candidate feature groups.

[0090] In one example, the candidate feature group with the largest statistical value can be determined as the target candidate feature group.

[0091] For example, for each candidate feature group, the selection priority of the candidate feature group can also be determined according to the statistical value of the feature evaluation index of the candidate feature group. According to the selection priorities of multiple candidate feature groups, select the candidate feature group with the highest priority from all candidate feature groups, and determine the candidate feature group with the highest priority as the target candidate feature group. In one example, the selection priority of the candidate feature group can be positively correlated with the statistical value of the feature evaluation index of the candidate feature group. That is to say, the larger the statistical value of the feature evaluation index of the candidate feature group, the higher the selection priority of the candidate feature group; the smaller the statistical value of the feature evaluation index of the candidate feature group, the lower the selection priority of the candidate feature.

[0092] Taking the example in step 206 above, among the above 10 statistical values, the statistical value corresponding to candidate feature group 3 is 0.7, which is the highest. Therefore, candidate feature group 3 can be determined as the target candidate feature group.

[0093] Step 208, determine whether the number of candidate features in the target candidate feature group reaches a predetermined value.

[0094] Specifically, determine whether the number of candidate features in the target candidate feature group reaches a predetermined value. This predetermined value is a threshold set in advance. Here, the size of the predetermined value can be determined according to the experience of those skilled in the art, or the number of initial features can also be determined as the size of the predetermined value. The present invention does not make any limitations here.

[0095] If it is determined that the number of candidate features in the target candidate feature group does not reach a predetermined value (e.g., the number of candidate features in the target candidate feature group is less than the predetermined value), then step 209 is executed. If it is determined that the number of candidate features in the target candidate feature group reaches the predetermined value (e.g., the number of candidate features in the target candidate feature group is greater than or equal to the predetermined value), then step 210 is executed.

[0096] Taking the above example, assume that the predetermined value is 5. If the number of candidate features in the target candidate feature group is 1 and does not reach the predetermined value, then step 209 is executed. If the number of candidate features in the target candidate feature group is 5 and reaches the predetermined value, then step 210 is executed.

[0097] Step 209: Construct a new candidate feature group based on the candidate features in the target candidate feature group, update the candidate feature group with the constructed new candidate feature group, and return to execute step 204.

[0098] At this time, in step 204, for each updated candidate feature group, multiple classifiers for the candidate feature are constructed using the determined sample set, and the subsequent steps are continued.

[0099] Here, the new candidate feature group may include multiple new candidate feature groups. Each new candidate feature group includes the candidate features in the target candidate feature group and one other candidate feature, where the other candidate feature is a feature other than the candidate features in the target candidate feature group among the multiple candidate features of the second time series data.

[0100] For example, the new candidate feature group can be constructed in the following way: Each other candidate feature among the other candidate features other than the candidate features in the target feature group in the multiple candidate features of the second time series data is combined with the candidate features in the target feature group respectively to obtain multiple new candidate feature groups. That is to say, the number of candidate features in each new candidate feature group is one more than the number of candidate features in the target candidate feature group. The candidate feature group is updated with the new candidate feature group, and return to execute step 204 to process the updated candidate feature group.

[0101] Taking the above example, assume that there are a total of 10 candidate feature groups. Among them, candidate feature group 1 contains candidate feature 1, candidate feature group 2 contains candidate feature 2, and so on; the target candidate feature group is candidate feature group 3, which contains candidate feature 3. The constructed new candidate feature groups include new candidate feature groups 1 to 9. Among them, candidate feature 3 and candidate feature 1 form new candidate feature group 1, candidate feature 3 and candidate feature 2 form new candidate feature group 2, candidate feature 3 and candidate feature 4 form new candidate feature group 3, and so on, thus determining 9 new candidate feature groups; use these 9 new candidate feature groups to update the original candidate feature groups, and obtain the updated candidate feature groups 1 to 9.

[0102] Step 210, determine the candidate features in the target candidate feature group as multiple updated features.

[0103] Specifically, if the number of candidate features in the target candidate feature group reaches a predetermined value, then determine the candidate features in the target candidate feature group as multiple updated features. Since whenever the number of candidate features in the target candidate feature group does not reach the predetermined value, new candidate features will be added to the target candidate feature group to obtain a new candidate feature group, so the number of candidate features in the target candidate feature group is constantly increasing. When the number of candidate features in the target candidate feature group reaches the predetermined value, do not return to execute step 204, and determine all the candidate features in the target candidate feature group as updated features.

[0104] Taking the above example, assume that the predetermined value is 5. Then when the number of candidate features in the target candidate feature group reaches 5, determine the 5 candidate features in the target candidate feature group as updated features.

[0105] Return Figure 1 , step 105, update the initial model based on multiple updated features.

[0106] Specifically, perform data preprocessing on the second time-series data, add the multiple updated features selected above to the initial model, and retrain the model according to the processes of model training, model selection, and model deployment of the initial model.

[0107] That is to say, the multiple updated features used when updating the initial model are different from the multiple initial features used when training the initial model. That is, when updating the initial model, new data is obtained and feature selection is performed again, so that the updated features used by the updated initial model are more suitable for the new data to improve the prediction accuracy of the model.

[0108] An embodiment of the present application provides a method for automatically updating a model. An initial model is constructed based on multiple initial features of first time-series data within a first predetermined time period; it is determined whether the model evaluation index of the constructed initial model meets a predetermined requirement; if the model evaluation index of the initial model meets the predetermined requirement, then second time-series data within a second predetermined time period is obtained; multiple updated features of the obtained second time-series data are determined; and the initial model is updated based on the multiple updated features. By adopting the above method for automatically updating a model, the problem that the model prediction accuracy is reduced due to changes in feature distribution in the case of a long time span is solved.

[0109] Based on the same inventive concept, an embodiment of the present application also provides a Figure 1 model automatic update device corresponding to the model automatic update method shown. Since the principle of solving problems by the device in the embodiment of the present application is similar to the above model automatic update method in the embodiment of the present application, the implementation of the device can refer to the implementation of the method, and repeated parts will not be elaborated.

[0110] Figure 3 The structural schematic diagram of the model automatic update device provided by the embodiment of the present application is shown as Figure 3 follows. The device includes the following modules:

[0111] A construction module 301 constructs an initial model based on multiple initial features of first time-series data within a first predetermined time period;

[0112] A judgment module 302 determines whether the model evaluation index of the constructed initial model meets a predetermined requirement;

[0113] An acquisition module 303, if the model evaluation index of the initial model meets the predetermined requirement, then acquires second time-series data within a second predetermined time period;

[0114] A determination module 304 determines multiple updated features of the obtained second time-series data;

[0115] An update module 305 updates the initial model based on the multiple updated features.

[0116] As an example, the model evaluation metrics may include a model stability metric and a model prediction effect metric. In this case, the determination module 302 may determine whether the model stability metric of the initial model is greater than a first set threshold, and determine whether the model prediction effect metric of the initial model is less than a second set threshold; if the model stability metric of the initial model is greater than the first set threshold, and / or, the model prediction effect metric of the initial model is less than the second set threshold, then the determination module determines that the model evaluation metrics of the initial model meet the predetermined requirements; if the model stability metric of the initial model is not greater than the first set threshold, and the model prediction effect metric of the initial model is not less than the second set threshold, then the determination module determines that the model evaluation metrics of the initial model do not meet the predetermined requirements.

[0117] In a possible implementation manner, the determination module 304 may determine a plurality of updated features in the following manner: determine a plurality of preset time points within a second predetermined time period; for each preset time point, determine a sample set corresponding to the preset time point, where the sample set includes training data before the preset time point and test data after the preset time point obtained by dividing the second time series data according to the preset time point; determine a plurality of candidate feature groups of the second time series data, each candidate feature group including at least one candidate feature; for each candidate feature group, construct a plurality of classifiers for the candidate feature by using the determined sample set; for each candidate feature group, determine a feature evaluation metric of the candidate feature group under each classifier; based on the determined feature evaluation metrics, determine a plurality of updated features from the plurality of candidate feature groups.

[0118] For the processing flows of the respective modules in the device and the interaction flows between the modules, reference may be made to the relevant descriptions in the foregoing method embodiments, which will not be elaborated herein.

[0119] The embodiments of the present application further provide a model automatic update system. Figure 4 For the structural schematic diagram of the model automatic update system provided by the embodiments of the present application, as Figure 4 shown, the system includes:

[0120] The background management server 402 is configured to obtain first time series data within a first predetermined time period from the remote server 401, and construct an initial model based on a plurality of initial features of the obtained first time series data.

[0121] The background management server 402 determines whether the model evaluation metrics of the constructed initial model meet the predetermined requirements. If the model evaluation metrics of the initial model meet the predetermined requirements, it obtains second time series data within a second predetermined time period from the remote server 401.

[0122] The background management server 402 determines multiple update features of the obtained second time-series data, and updates the initial model based on the multiple update features.

[0123] The remote server 401 is used to store model update files, and the model update files include basic configuration files and model algorithm files.

[0124] The client management platform 403 is used to adjust various preset parameter values and send service request information to the background management server to view the returned calculation results.

[0125] Corresponding to Figure 1 In the model automatic update method, an embodiment of the present application further provides a schematic structural diagram of an electronic device 500, as shown in Figure 5 As shown, the electronic device 500 includes a processor 510, a memory 520, and a bus 530. The memory 520 stores machine-readable instructions executable by the processor 510. When the electronic device 500 runs, the processor 510 communicates with the memory 520 through the bus 530. When the machine-readable instructions are executed by the processor 510, the above-mentioned model automatic update method can be executed. When updating the model, feature variables applicable to new sample data are selected, avoiding the problem that the prediction result of the updated model is inaccurate due to the feature variables not being applicable to new sample data.

[0126] Corresponding to Figure 1 In the model automatic update method, an embodiment of the present application further provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium, and when the computer program is run by a processor, the steps of the above-mentioned model automatic update method are executed.

[0127] Specifically, the storage medium can be a general storage medium, such as a mobile disk, a hard disk, etc. When the computer program on the storage medium is run, the above-mentioned model automatic update method can be executed. When updating the model, feature variables applicable to new sample data are selected, avoiding the problem that the prediction result of the updated model is inaccurate due to the feature variables not being applicable to new sample data.

[0128] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the above method embodiments, and will not be elaborated here.

[0129] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some communication interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.

[0130] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0131] In addition, each functional unit in the embodiments provided in this application can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0132] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0133] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. In addition, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0134] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present application, used to illustrate the technical solutions of the present application, rather than limiting it. The protection scope of the present application is not limited thereto. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present application can still modify the technical solutions recorded in the foregoing embodiments or can easily think of changes, or perform equivalent replacements for some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for automatically updating a model, characterized in that, Including: Based on multiple initial features representing the number of insurance policies and premiums in the first insurance policy time series data within the first predetermined time period, construct an initial renewal prediction model, where the initial renewal prediction model is a model for predicting the user renewal rate, and the first insurance policy time series data includes insurance policy data and user behavior data; Determine whether the model evaluation indicators of the constructed initial renewal prediction model meet the predetermined requirements; If the model evaluation indicators of the initial renewal prediction model meet the predetermined requirements, obtain the second insurance policy time series data within the second predetermined time period adjacent to the current time; Determine multiple updated features of the obtained second insurance policy time series data; Update the initial renewal prediction model based on the multiple updated features.

2. The method according to claim 1, wherein Determining multiple updated features of the obtained second insurance policy time series data includes: (A) Determine multiple preset time points within the second predetermined time period; (B) For each preset time point, determine the corresponding sample set, where the sample set includes training data before the preset time point and test data after the preset time point obtained by dividing the second insurance policy time series data according to the preset time point; (C) Determine multiple candidate feature groups of the second insurance policy time series data, and each candidate feature group includes at least one candidate feature; (D) For each candidate feature group, use the determined sample set to construct multiple classifiers for the candidate feature; (E) For each candidate feature group, determine the feature evaluation indicators of the candidate feature group under each classifier; (F) Based on the determined feature evaluation indicators, determine multiple updated features from multiple candidate feature groups.

3. The method according to claim 2, wherein Based on the determined feature evaluation indicators, determining multiple updated features from multiple candidate feature groups includes: (F1) For each candidate feature group, determine the statistical value of the feature evaluation indicators of the candidate feature group under each classifier; (F2) Based on the statistical values corresponding to all candidate feature groups, select the target candidate feature group from all candidate feature groups; (F3) Determine whether the number of candidate features in the target candidate feature group reaches a predetermined value; (F4) If the number of candidate features in the target candidate feature group does not reach the predetermined value, construct a new candidate feature group based on the candidate features in the target candidate feature group, update the candidate feature group with the constructed new candidate feature group, and return to execute step (D); (F5) If the number of candidate features in the target candidate feature group reaches the predetermined value, determine the candidate features in the target candidate feature group as the multiple updated features.

4. The method according to claim 3, characterized in that, The new candidate feature group includes multiple new candidate feature groups, each new candidate feature group includes the candidate features in the target candidate feature group and one other candidate feature, and the other candidate feature is a feature other than the candidate features in the target candidate feature group among the multiple candidate features of the second insurance policy time series data.

5. The method according to claim 1, wherein The model evaluation indicators include a model stability indicator and a model prediction effect indicator, wherein, determining whether the model evaluation indicators of the constructed initial renewal prediction model meet the predetermined requirements includes: Determine whether the model stability index of the initial renewal prediction model is greater than a first set threshold, and determine whether the model prediction effect index of the initial renewal prediction model is less than a second set threshold; If the model stability index of the initial renewal prediction model is greater than the first set threshold, and / or, the model prediction effect index of the initial renewal prediction model is less than the second set threshold, then determine that the model evaluation index of the initial renewal prediction model meets the predetermined requirements; If the model stability index of the initial renewal prediction model is not greater than the first set threshold, and the model prediction effect index of the initial renewal prediction model is not less than the second set threshold, then determine that the model evaluation index of the initial renewal prediction model does not meet the predetermined requirements.

6. A model automatic update device, characterized in that, It includes: A construction module, which constructs an initial renewal prediction model based on multiple initial features representing the number of policies and premiums in the first policy time-series data within a first predetermined time period. The initial renewal prediction model is a model for predicting the user renewal rate, and the first policy time-series data includes policy data and user behavior data; A judgment module, which judges whether the model evaluation index of the constructed initial renewal prediction model meets the predetermined requirements; An acquisition module, if the model evaluation index of the initial renewal prediction model meets the predetermined requirements, then acquire the second policy time-series data within a second predetermined time period adjacent to the current time; A determination module, which determines multiple updated features of the acquired second policy time-series data; An update module, which updates the initial renewal prediction model based on the multiple updated features.

7. The device according to claim 6, wherein, The determination module determines the multiple updated features in the following manner: Determine multiple preset time points within the second predetermined time period; For each preset time point, determine a sample set corresponding to the preset time point. Among them, the sample set includes training data before the preset time point and test data after the preset time point obtained by dividing the second policy time-series data according to the preset time point; Determine multiple candidate feature groups of the second policy time-series data, and each candidate feature group includes at least one candidate feature; For each candidate feature group, construct multiple classifiers for the candidate feature by using the determined sample set; For each candidate feature group, determine the feature evaluation index of the candidate feature group under each classifier; Based on the determined feature evaluation indexes, determine multiple updated features from multiple candidate feature groups.

8. A model automatic update system, characterized in that, It includes: A background management server and a remote server; The background management server obtains the first policy time-series data within a first predetermined time period from the remote server, and constructs an initial renewal prediction model based on multiple initial features representing the number of policies and premiums in the obtained first policy time-series data. The initial renewal prediction model is a model for predicting the user renewal rate, and the first policy time-series data includes policy data and user behavior data; The background management server judges whether the model evaluation index of the constructed initial renewal prediction model meets the predetermined requirements. If the model evaluation index of the initial renewal prediction model meets the predetermined requirements, then obtain the second policy time-series data within a second predetermined time period adjacent to the current time from the remote server; The background management server determines a plurality of update features of the obtained second insurance policy time series data, and updates the initial renewal prediction model based on the plurality of update features.

9. An electronic device, characterized in that, It includes a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are run by the processor, the steps of the model automatic update method as described in any one of claims 1-5 are executed.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is run by a processor, the steps of the model automatic update method as described in any one of claims 1-5 are executed.

Citation Information

Patent Citations

  • Method of insurance retainment of insurance policy, storage medium and server

    CN108550082A

  • Scheme pushing method and device, computer device and readable storage medium

    CN110706054A