Model optimization method, model optimization device, storage medium and electronic equipment

By processing the initial model and sample data, important features are selected and new sample data are constructed, and the initial model is trained, the problem of complex and inefficient model optimization in the existing technology is solved, and efficient and convenient model optimization is achieved. The obtained target model has high accuracy and efficiency.

CN120105058APending Publication Date: 2025-06-06HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510173581.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In the prior art, model optimization operations are complex and low efficiency, making it difficult to achieve efficient and convenient model optimization.

Method used

By obtaining the initial model and multiple first sample data, using the initial model to process these data, determine the importance of the initial feature, filter out the target features, build the second sample data, and use the data to train the initial model to obtain the optimized target model.

Benefits of technology

This method reduces human participation, reduces labor and time costs, improves the accuracy and efficiency of model optimization, and the obtained target model is trained based on the characteristics importance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120105058A_ABST
    Figure CN120105058A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a model optimization method, a model optimization device, a storage medium and electronic equipment, and relates to the technical field of computers. The method comprises the steps of obtaining an initial model and multiple pieces of first sample data; wherein each piece of first sample data comprises a plurality of initial features; the initial model is adopted to process the multiple pieces of first sample data, and a prediction result is obtained; determining the importance degree value of each initial feature according to the prediction result; according to the importance degree value, determining a target feature from a plurality of initial features of the first sample data to obtain second sample data including the target feature; and training the initial model by adopting the second sample data to obtain a target model after model optimization. The importance degree of the features can be accurately and conveniently determined, and the model is accurately optimized in combination with the importance of the features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the field of computer technology. More specifically, the embodiments of the present disclosure relate to a model optimization method, a model optimization device, a computer-readable storage medium, and an electronic device. Background Art

[0002] This section is intended to provide a background or context to the embodiments of the disclosure that are recited in the claims, and no description herein is admitted to be prior art by inclusion in this section.

[0003] With the rapid development of artificial intelligence technology, in order to improve the application effect of models in multiple scenarios, it is often necessary to evaluate and optimize the models. Based on this, the prior art proposes that each time a new model is trained, the importance of features can be calculated manually by technicians, or the test results of online experimental models can be used to determine whether to optimize the model. The model optimization operation is complex and costly, and the efficiency of model optimization is low. How to optimize the model in an efficient and convenient way is a problem that needs to be solved urgently in the prior art. Summary of the invention

[0004] However, the operation complexity of model optimization in the current existing technology is high and the efficiency of model optimization is low.

[0005] To this end, there is a great need for a model optimization method that can optimize the model in an efficient and convenient way.

[0006] In this context, embodiments of the present disclosure are intended to provide a model optimization method, a model optimization device, a computer-readable storage medium, and an electronic device.

[0007] According to a first aspect of the present disclosure, a model optimization method is provided, comprising: obtaining an initial model and multiple first sample data; wherein each of the first sample data comprises multiple initial features; using the initial model to process the multiple first sample data to obtain a prediction result; determining an importance value of each of the initial features based on the prediction result; determining a target feature from the multiple initial features of the first sample data based on the importance value to obtain second sample data including the target feature; and using the second sample data to train the initial model to obtain a target model after model optimization.

[0008] In an exemplary embodiment, the using the initial model to process the multiple first sample data to obtain a prediction result includes: using the initial model to process the multiple first sample data to obtain a first prediction result; performing feature enhancement processing on the multiple first sample data to obtain multiple third sample data, and using the initial model to process the multiple third sample data to obtain a second prediction result; determining the importance value of each of the initial features based on the prediction result includes: calculating the importance value of each of the initial features based on the first prediction result and the second prediction result.

[0009] In an exemplary embodiment, the performing feature enhancement processing on the multiple first sample data to obtain multiple third sample data includes: extracting multiple sampling data from the multiple first sample data, and rearranging or changing the features in the sampling data to obtain the multiple third sample data.

[0010] In an exemplary embodiment, after extracting multiple sampling data from the multiple first sample data, the method further includes: determining whether each sampling data is sparse data based on the characteristic value of the initial feature in each sampling data; if the sampling data is determined to be sparse data, performing dense data conversion processing on the sampling data.

[0011] In an exemplary embodiment, determining the target feature from multiple initial features of the first sample data based on the importance value includes: among the multiple initial features of the first sample data, taking the initial feature whose importance value meets a first preset condition as the target feature.

[0012] In an exemplary embodiment, the use of the second sample data to train the initial model to obtain a target model after model optimization includes: using the second sample data to train the initial model to determine an intermediate model; using the target features in the second sample data as initial features, and using the intermediate model to again determine the importance values ​​of each initial feature in the second sample data; when the importance values ​​of each initial feature meet a second preset condition, using the intermediate model as the target model.

[0013] In an exemplary embodiment, determining the importance value of each of the initial features based on the prediction result includes: calculating the importance value of each of the initial features based on the prediction result, and visually displaying the importance value of each of the initial features on a display interface.

[0014] According to a second aspect of the present disclosure, a model optimization device is provided, comprising: a sample data acquisition module, configured to acquire an initial model and a plurality of first sample data; wherein each of the first sample data comprises a plurality of initial features;

[0015] A prediction result acquisition module is used to use the initial model to process the multiple first sample data to obtain a prediction result; an importance determination module is used to determine the importance value of each of the initial features according to the prediction result; a target feature determination module is used to determine the target feature from the multiple initial features of the first sample data according to the importance value to obtain second sample data including the target feature; a target model acquisition module is used to use the second sample data to train the initial model to obtain a target model after model optimization.

[0016] According to a third aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the model optimization method of the first aspect and possible implementation methods thereof are implemented.

[0017] According to a fourth aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the model optimization method of the above-mentioned first aspect and its possible implementation method by executing the executable instructions.

[0018] In the scheme disclosed in the present invention, on the one hand, compared with the method of manually calculating the importance of features, the present exemplary embodiment proposes a new optimization method, which can use the initial model to process multiple first sample data to obtain prediction results, and determine the importance value of each initial feature based on the prediction results. This process requires less human participation, reducing labor costs and time costs; on the other hand, the present exemplary embodiment provides a model optimization method, which can determine the target feature and construct second sample data including the target feature based on the importance value of the feature. Furthermore, the initial model can be trained based on the second sample data after feature optimization to obtain the target model. Since the target model is trained based on the sample data corresponding to the target feature determined by the importance value, the obtained target model is also a model trained based on the feature importance, and the accuracy and effectiveness of the model optimization are relatively high. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 A schematic diagram showing a system architecture in this exemplary embodiment.

[0020] Figure 2 A flow chart of a model optimization method in this exemplary embodiment is shown.

[0021] Figure 3 A flow chart showing another model optimization method in this exemplary embodiment.

[0022] Figure 4 A schematic diagram of a visualization display interface for configuring model evaluation indicators in this exemplary embodiment is shown.

[0023] Figure 5 A schematic diagram showing a visual display of execution results in this exemplary embodiment is shown.

[0024] Figure 6 A partial interface schematic diagram showing importance values ​​of features displayed on a visualization interface in this exemplary embodiment is shown.

[0025] Figure 7 A schematic diagram showing a framework of a model optimization method in this exemplary embodiment.

[0026] Figure 8 A schematic diagram of the overall framework of a model optimization method in this exemplary embodiment is shown.

[0027] Fig. 9 A schematic diagram showing a method of determining an importance value of a feature in this exemplary embodiment.

[0028] Fig.10 A schematic diagram showing the inferred verification results of feature deletion of a model in a music platform scenario.

[0029] Fig.11 A comparative schematic diagram showing a reliability diagram in this exemplary embodiment.

[0030] Fig.12 A schematic diagram showing a scoring distribution diagram in this exemplary embodiment.

[0031] Fig.13 A schematic structural diagram of a model optimization device in this exemplary embodiment is shown.

[0032] Fig.14 A schematic structural diagram of an electronic device in this exemplary embodiment is shown.

[0033] In the drawings, the same or corresponding reference numerals represent the same or corresponding parts. DETAILED DESCRIPTION

[0034] The principles and spirit of the present disclosure will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and implement the present disclosure, and are not intended to limit the scope of the present disclosure in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.

[0035] The embodiments of the present disclosure may be implemented as a system, device, apparatus, method or computer program product. Therefore, the present disclosure may be specifically implemented in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0036] The principle and spirit of the present disclosure are explained in detail below with reference to several representative embodiments of the present disclosure. SUMMARY OF THE INVENTION

[0038] The inventors have found that the current model optimization efficiency and accuracy need to be improved.

[0039] In view of the above, the present disclosure provides a model optimization method, a model optimization device, a computer-readable storage medium, and an electronic device. On the one hand, compared with the method of manually calculating the importance of features, this exemplary embodiment proposes a new optimization method, which can use an initial model to process multiple first sample data to obtain prediction results, and determine the importance value of each initial feature based on the prediction results. This process requires less human participation, reducing labor costs and time costs; on the other hand, this exemplary embodiment provides a model optimization method, which can determine the target feature and construct a second sample data including the target feature based on the importance value of the feature. Further, the initial model can be trained based on the second sample data after feature optimization to obtain a target model. Since the target model is trained based on the sample data corresponding to the target feature determined by the importance value, the obtained target model is also a model trained based on the feature importance, and the accuracy and effectiveness of the model optimization are high.

[0040] After introducing the basic principles of the present disclosure, various non-limiting embodiments of the present disclosure are described in detail below.

[0041] Application Scenario Overview

[0042] It should be noted that the following application scenarios are only shown to facilitate understanding of the spirit and principle of the present disclosure, and the embodiments of the present disclosure are not limited in this respect. On the contrary, the embodiments of the present disclosure can be applied to any applicable scenario.

[0043] The embodiments of the present disclosure can be applied to relevant scenarios of model optimization. The application scenarios are described in detail below in conjunction with the system architecture.

[0044] Figure 1A schematic diagram of a system architecture for model optimization is shown. The system architecture includes a user terminal 110 and a server 120. The user can configure information in the user terminal 110. The server 120 can obtain the initial model and sample data to be optimized, and calculate the feature importance value. After obtaining the calculation result, it can be returned to the user terminal 110 for display. The user can view the calculation result and determine whether to determine the target model or continue to iterate the calculation of the feature importance value or the model optimization process.

[0045] Exemplary Methods

[0046] The exemplary embodiment of the present disclosure provides a model optimization method. Figure 2 As shown, the method may include steps S210 to S250. Figure 2 Provide detailed instructions for each step.

[0047] refer to Figure 2 In step S210, an initial model and a plurality of first sample data are obtained; wherein each first sample data includes a plurality of initial features.

[0048] The initial model can be any model that needs to be optimized, such as a behavior prediction model, a behavior evaluation model or a classification model, etc. The model can be applied to a variety of application scenarios, such as predicting user behavior based on user feature data in a music platform. Different initial models can include different network structures according to actual conditions. The first sample data refers to sample data used to train or optimize the initial model, and the initial model can process the first sample data to obtain an output result.

[0049] In this exemplary embodiment, the first sample data may include multiple initial features, which can be regarded as initial dimensional indicators of feature data in the sample data. For example, the first sample data may be collected data related to the user, which may include user attribute information such as user age, user gender, user occupation, as well as user behavior information, user preference information, etc. The initial features may be five features: user age, user gender, user occupation, user behavior information, and user preference information. The first sample data is data composed of feature data corresponding to the five features: user age, user gender, user occupation, user behavior information, and user preference information. The first sample data may be a vector or a matrix. For example, when the feature data of user age, user gender, user occupation, user behavior information, and user preference information are all vectors, the first sample data may be a vector spliced ​​from these multiple vectors.

[0050] Step S220: using the initial model to process the plurality of first sample data to obtain prediction results.

[0051] In this exemplary embodiment, the initial model can process multiple first sample data separately to obtain prediction results corresponding to each first sample data. In different application scenarios, the prediction results may be different. For example, the first sample data is collected data related to the user. After the initial model for behavior prediction processes the first sample data, it can obtain the user's predicted behavior, or predict the object recommended for the user, etc.

[0052] In this exemplary embodiment, the initial model is used to process multiple first sample data to obtain prediction results, which may include different stages. Specifically, in the first stage, the initial model is used to process multiple first sample data to obtain a first prediction result, and then the multiple first sample data are updated. In the second stage, the initial model is used to process the multiple updated first sample data to obtain a second prediction result.

[0053] Step S230: Determine the importance value of each initial feature according to the prediction result.

[0054] In this exemplary embodiment, after the prediction result is determined, the importance of each initial feature can be evaluated based on the prediction result to obtain the importance value of each initial feature, wherein the importance value can be used to characterize the importance of the feature in the data, for example, the importance value can be represented by an importance weight value. This exemplary embodiment can calculate the performance indicators of the initial model based on the prediction result, such as accuracy, loss, etc., and then determine the importance value of the feature. Specifically, when the initial model is used to process multiple first sample data to obtain the first preset result and the second preset result, the importance value of each initial feature can be determined by comparing the difference in the performance indicators of the model corresponding to the first preset result and the second preset result. The larger the difference, the more important the corresponding initial feature is.

[0055] Step S240: determining a target feature from a plurality of initial features of the first sample data according to the importance value, so as to obtain second sample data including the target feature.

[0056] Among them, the target feature refers to a more important feature screened out from the initial features. This exemplary embodiment can determine the target feature from multiple initial features of the first sample data according to the importance value. For example, the initial features whose importance values ​​do not reach a preset threshold can be deleted from multiple initial features, and the remaining initial features can be used as the target feature, or a preset number of initial features before the importance values ​​are sorted can be used as the target feature among multiple initial features, and then the second sample data including the target feature can be obtained. The second sample data refers to sample data composed of feature data corresponding to the target feature. For example, the initial features included in the first sample data are user age, user gender, user occupation, user behavior information, and user preference information. When the target features are determined to be user age, user gender, and user behavior information, the second sample data can be sample data of feature data including user age, user gender, and user behavior information features.

[0057] Step S250, using the second sample data to train the initial model to obtain a target model after model optimization.

[0058] Finally, after determining the second sample data, the second sample data can be used to train the initial model again. At this time, since the feature data in the second sample data is already sample data composed of feature data corresponding to more important target features after feature screening, the optimized target model can be obtained after training the initial model with the second sample data.

[0059] In this exemplary embodiment, the target model can be used as the initial model, and the second sample data can be used as the first sample data. The process returns to step S210 to redetermine the importance values ​​of the features in the second sample data, and further evaluate whether feature screening is required. When it is determined that the model performance or the importance values ​​of the features meet the preset conditions, the loop can be stopped to determine the final target model. For example, when the performance difference between the model after feature deletion and the model before deletion is less than a preset threshold, or the importance of the features is greater than the preset threshold, the target model can be determined.

[0060] In an exemplary embodiment, if Figure 3 As shown, the above-mentioned use of the initial model to process multiple first sample data to obtain prediction results may include the following steps:

[0061] Step S310, using the initial model to process a plurality of first sample data to obtain a first prediction result;

[0062] Step S320, performing feature enhancement processing on the plurality of first sample data to obtain a plurality of third sample data, and using the initial model to process the plurality of third sample data to obtain a second prediction result;

[0063] The above-mentioned determination of the importance value of each initial feature based on the prediction result may include:

[0064] Step S330: Calculate the importance value of each initial feature according to the first prediction result and the second prediction result.

[0065] Among them, feature enhancement processing for multiple first sample data may be processing one or more features in the first sample data, which may specifically include randomly aligning and permuting each feature, or disrupting the feature data under the feature, or mixing the feature data with other feature data, etc.

[0066] The first sample data may be sample data before the feature is processed, and the third sample data may be sample data after the feature is processed. This exemplary embodiment may use the initial model to process multiple first sample data respectively to obtain a first prediction result, and process multiple third sample data to obtain a second prediction result. Then, based on the first prediction result and the second prediction result, the importance value of each initial feature is calculated. For example, the performance of the initial model may be evaluated based on the first prediction result and the second prediction result to determine the performance difference of the initial model before and after the sample data processing. The greater the difference, the more important the processed initial feature. If an initial feature is highly important to the initial model, replacing the initial feature will cause the performance of the initial model to decrease because the initial model loses its dependence on the initial feature. This exemplary embodiment may perform the above steps for each initial feature to calculate the importance value of each initial feature.

[0067] In an exemplary embodiment, performing feature enhancement processing on a plurality of first sample data to obtain a plurality of third sample data may include:

[0068] A plurality of sampling data are extracted from the plurality of first sample data, and features in the sampling data are rearranged or changed to obtain a plurality of third sample data.

[0069] In order to speed up the calculation, especially when processing large data sets, the exemplary embodiment can extract multiple sampling data from multiple first sample data, rearrange or change the features in the sampling data to obtain multiple third sample data, that is, extract part of the data from the multiple first sample data for processing to generate the third sample data. In particular, when calculating the importance of features, the exemplary embodiment can rearrange or change the features of the sampling data to solve the influence of confusion or correlation on the features, and the rearrangement or change may include shuffling, replacing, deleting, etc. the features or feature data.

[0070] In this exemplary embodiment, extracting multiple sampling data from multiple first sample data can be randomly extracted according to a preset ratio, for example, randomly extracting 20% ​​of the data from multiple first sample data as sampling data, or first pre-processing the multiple first sample data, for example, filtering abnormal data such as duplicate, invalid or empty data, and then according to the timestamp of each first sample data, taking the first sample data within a preset time period from the current time as sampling data, etc. The present disclosure does not make specific limitations on this.

[0071] In an exemplary embodiment, after extracting a plurality of sampling data from a plurality of first sample data, the above-mentioned model optimization method may further include:

[0072] Determine whether each sampled data is sparse data according to the feature value of the initial feature in each sampled data;

[0073] If it is determined that the sampled data is sparse data, dense data conversion is performed on the sampled data.

[0074] In order to improve the effectiveness and accuracy of feature importance calculation, this exemplary embodiment can determine whether each sampled data is sparse data based on the eigenvalue of the initial feature in each sampled data. For example, the sampled data is actually extracted from the first sample data and can be in the form of a vector or a matrix like the first sample data. Each column of data can correspond to the feature data of an initial feature. If the eigenvalues ​​of most of the initial features in the sampled data are 0, it means that the sampled data is relatively sparse. Therefore, this exemplary embodiment can determine whether the sampled data is sparse data by determining whether the number of eigenvalues ​​of the initial features in the sampled data that are 0 exceeds a preset number. When the number of eigenvalues ​​of the initial features that are 0 exceeds a preset number, the sampled data is determined to be sparse data.

[0075] When the sampled data is determined to be sparse data, the sampled data may be subjected to dense data conversion processing to convert it into a denser representation so as to more efficiently calculate feature importance. The dense data conversion processing may be implemented in a variety of ways, such as processing the sparse sampled data through machine learning or deep learning models to obtain dense data, or converting a sparse high-dimensional feature vector into a dense low-dimensional feature vector through linear mapping or nonlinear mapping, etc., which is not specifically limited in the present disclosure.

[0076] In an exemplary embodiment, determining the target feature from a plurality of initial features of the first sample data according to the importance value may include:

[0077] Among the multiple initial features of the first sample data, the initial features whose importance values ​​meet the first preset condition are taken as target features.

[0078] This exemplary embodiment can determine the target feature from multiple initial features according to whether the importance value of the initial feature meets the first preset condition. The first preset condition may include that the importance value is greater than a preset threshold, for example, the initial feature whose importance value is greater than the preset threshold is used as the target feature, and the first preset condition may also include a preset number before the importance value is sorted, for example, the initial features of a preset number before the importance value is sorted are used as the target feature.

[0079] In an exemplary embodiment, the above-mentioned using the second sample data to train the initial model to obtain the target model after model optimization may include:

[0080] Using the second sample data to train the initial model and determine the intermediate model;

[0081] The target feature in the second sample data is used as the initial feature, and the importance value of each initial feature in the second sample data is determined again using the intermediate model;

[0082] When the importance values ​​of the initial features meet the second preset condition, the intermediate model is used as the target model.

[0083] Among them, the intermediate model refers to the model obtained by training the initial model with the second sample data. The intermediate model can be used as the target model, or it can continue to be optimized to obtain an optimized intermediate model until it meets the requirements, and the intermediate model that meets the requirements will be used as the target model.

[0084] In this exemplary embodiment, after determining the intermediate model, the target features in the second sample data can be used as initial features, and the importance values ​​of each initial feature in the second sample data can be determined again using the intermediate model, and whether to use the intermediate model as the target model is determined based on the calculation results. For example, the first sample data is composed of feature data of 5 features. After the first round of feature importance value calculation, the importance values ​​of the 5 features are determined, and 4 relatively important features are screened out; the second sample data is composed of feature data of 4 features. The intermediate model can be obtained by training the initial model with the second sample data. The 4 features in the second sample data are used as initial features, and the importance values ​​of the second round of feature importance value calculation is performed to determine the importance values ​​of the 4 features. When the importance values ​​of the 4 features meet the second preset condition, the iteration can be stopped, and the intermediate model trained with the sample data including the 4 features is used as the target model. If the importance values ​​of the 4 features do not meet the second preset condition, the importance values ​​of the features can continue to be calculated and screened until the importance values ​​of the features meet the second preset condition, and the intermediate model trained with the sample data consisting of the feature data corresponding to the features meeting the second preset condition is determined as the target model. Among them, the second preset condition can be a judgment condition for determining whether the feature screening or deletion is stable, and the second preset condition can include that the importance values ​​of the features are all higher than the preset threshold, or the number of features with importance values ​​higher than the preset threshold is greater than the preset number, or the importance value of the feature calculated in this round is the same as the importance value of the feature calculated in the previous round, or the difference between the importance value of the feature calculated in this round and the importance value of the feature calculated in the previous round is less than the preset difference, etc. In this exemplary embodiment, if the importance value of the feature calculated in this round is the same or similar to the importance value of the feature calculated in the previous round, it means that the importance of the deleted feature is low, and there is basically no effect on the model after deletion, and the model trained with the sample data corresponding to the deleted feature can be used as the target model.

[0085] In an exemplary embodiment, determining the importance value of each initial feature according to the prediction result may include:

[0086] According to the prediction results, the importance value of each initial feature is calculated, and the importance value of each initial feature is visualized on the display interface.

[0087] This exemplary embodiment may provide a visual display interface, and visualize the calculated importance values ​​of the initial features on the display interface to display the importance values ​​of each initial feature. Specifically, it may be displayed in the form of a numerical value or in the form of a bar graph, and the present disclosure does not make any specific limitation on this.

[0088] In this exemplary embodiment, the importance values ​​of features determined in different rounds can be displayed in different interfaces. For example, the first sample data includes 5 features and the second sample data includes 4 features. The calculation results of the importance values ​​of the features in the first sample data can be displayed on the first page, and the importance values ​​of the features in the second sample data can be displayed on the second page. In addition, in order to provide a clear and effective comparison example, this exemplary embodiment can also compare and display the calculation results of different rounds, which is not specifically limited in the present disclosure.

[0089] Based on this exemplary embodiment, the sedimentation of related platforms can be achieved. Through the modular design of the platform, each module can run independently or cooperate with each other to support the needs of various businesses. It will be expanded to other business scenarios such as search, promotion and marketing in the future. This exemplary embodiment can apply the above-mentioned model evaluation and optimization method based on feature importance to the product. The product provides a simple process that can help algorithm users diagnose and evaluate models with one click, and visualize the diagnosis results for algorithm users to make subsequent decisions. The product is mainly divided into three parts: information configuration, diagnosis result aggregation, and visual display. Among them, the visual display can include a visualization page for routine analysis results of the model, a visualization page for model interpretability and inference verification results, a visualization interface for model evaluation results, and so on. The whole process is smooth and efficient, which solves the problems of opacity and difficulty in optimization encountered when the algorithm iterates the model, and efficiently helps the algorithm decide how to select model features.

[0090] Figure 4 A schematic diagram of a visual display interface for configuring model evaluation indicators in this exemplary embodiment is shown, which truly illustrates the content of information configuration, in which users can configure information such as model performance indicators, test sets, training sets, and PFI calculation parameters. Figure 5 A schematic diagram of a local page for visually displaying the execution results in this exemplary embodiment is shown, showing the aggregated content of the diagnosis results, which may specifically include model information of the model to be processed, such as the name of the model, the processing sequence number of the model, the path of the model, etc., and may also include calculation information specifically related to the model, such as instance identification, time, executor, status, running time, execution ID (Identify), execution method, training report, explainability and other attributes, wherein the user can also operate on the attribute value to browse the specific content, for example, click on the training report to view detailed information. Figure 6 A partial page schematic diagram of a visualization interface in this exemplary embodiment showing the calculation results of feature importance values ​​is shown. In this display interface, the importance values ​​of different features can be displayed in the form of a bar chart, so that the user can clearly and intuitively see the important differences between different features. In addition, configuration information or model information can also be displayed in the interface for displaying feature importance values, and the present disclosure does not make specific limitations on this.

[0091] Figure 7 A schematic diagram of a system framework for model optimization in this exemplary embodiment is shown, which may specifically include the following steps:

[0092] First, the user can select the initial model and the corresponding data set to be analyzed in the product function layer 710. The data set may include the first sample data 711. The product function layer 710 may include a data collection unit 712 for collecting the first sample data, and a sample data processing unit 713 for packaging the first sample data. The system may automatically persist the user's selection and pass the first sample data to the task processing layer 720.

[0093] After receiving the first sample data, the task processing layer 720 can convert the first sample data into the data structure required by the kernel component layer in advance through the preprocessing unit 721, and automatically create a task flow through the task flow creation unit 722 according to the indicators that the user needs to analyze, such as the importance index of the feature, the model performance index, etc.;

[0094] After the task flow is created, it will trigger the kernel component layer 730 to execute, first use the first sample data 711 to perform model training 731, and then pass the obtained initial model 732 into the model prediction module 740;

[0095] The model prediction module 740 includes a model loading unit 741 for loading an initial model 731; a sample reading unit 742 for reading first sample data 711; a first prediction unit 743 for processing the first sample data 711 using the initial model 731 to obtain a first prediction result; wherein the first sample data 711 for performing model training 730 and the first sample data 711 for obtaining the first prediction result may come from the same training set or from different training sets, which is not specifically limited in the present disclosure; the first prediction result will be passed to the model interpretable module 750;

[0096] The model interpretable module 750 may include a model loading unit 751 for loading an initial model 731; a sampling reading unit 752 for extracting sampling data from the first sample data 711; a dense processing unit 753 for converting the sampling data into dense data when the sampling data is determined to be sparse data; a confusion and scattering unit 754 for performing one or more operations such as permutation, confusion or scattering on the features in the sampling data; a second prediction unit 755 for processing the data processed by the confusion and scattering unit using the initial model 731 to obtain a second prediction result; a PFI (Permutation Feature Importance) deviation calculation unit 756 for calculating the importance value of the initial feature according to the first prediction result and the second prediction result; and a result writing unit 757 for recording the calculation result of the importance value of the initial feature;

[0097] The task processing layer 720 also includes a callback receiving unit 723 for receiving the calculation results of the kernel component layer 730; and a callback result processing unit 724 for performing processing according to the callback results, such as visually displaying the calculation results to the user, or the user determining the target model or whether to continue model optimization based on the displayed results.

[0098] Figure 8 The overall framework schematic diagram of a model optimization method of this exemplary embodiment is shown. The framework includes three layers, which respectively include:

[0099] Kernel component layer 810: This layer can mainly implement various model diagnosis kernel functions based on Python (a programming language), including model prediction 811, offline evaluation 812, and model interpretability 813. By encapsulating the above into components, combined calls are implemented, which is conducive to function expansion and maintenance and improves efficiency. The various functions are explained as follows:

[0100] Model prediction 811: the result or output obtained by a machine learning or statistical model based on input data. This exemplary embodiment can obtain a prediction result for an existing model on a specified data set for subsequent offline evaluation and calculation of model feature importance;

[0101] Offline evaluation 812: The model performance can be evaluated without interacting with real-time data. This exemplary embodiment can calculate model evaluation indicators such as AUC (area under curve, area under the ROC curve) and GAUC (Group AUC) of the model under a specified data set to evaluate the performance of the model.

[0102] Model interpretability 813: This exemplary embodiment can calculate the importance of each feature in the model through PFI based on the model prediction results, so as to judge the necessity of the feature;

[0103] Task processing layer 820: This layer can manage the creation of various model diagnosis task flows based on Java (a programming language). By hiding the complex internal implementation details, it provides a simple task flow creation capability and modularizes the system. It enables the core analysis capabilities of the kernel component layer to be automatically executed and created, giving the system automation capabilities, which mainly include:

[0104] Data processing 821: used to pre-process metadata such as data set information and model information to be analyzed, and process these metadata into a structure that can be recognized by the kernel analysis layer to standardize data links and form the premise of automated analysis. Specifically, it may include a data collection unit 8211, a data pre-processing unit 8212, and a metadata management unit 8213;

[0105] Task flow creation 822: can flexibly combine the functions of the kernel component layer and support the routine execution of the model analysis task flow, which can greatly reduce the loss of human resources and improve the efficiency of model evaluation. It can include a routine scheduling unit 8221, an active execution unit 8222, and a re-run unit 8223;

[0106] Callback processing 823: used to receive the results calculated by the kernel component layer and store the results in a standardized manner for visual display on the platform, improve system efficiency, and provide optimization direction for the algorithm; it may include a model performance result callback processing unit 8231, a model interpretability callback processing unit 8232, and an inference verification callback processing unit 8233;

[0107] Product function layer 830: This layer can be responsible for user management, authority management, model offline evaluation configuration management, model routine analysis monitoring, model interpretability configuration management, and model inference verification configuration management. Users can operate the obtained data through the visual interface, persist it in the product function, and link the task processing layer to create a new model evaluation task flow, so as to realize a one-stop model evaluation optimization process; specifically, it can include offline evaluation 831, where users can interact with terminal devices, choose to determine the calculation task indicators, or browse visualization content, such as AUC, GAUC, COPC (an indicator), scoring layout, reliability diagram; model routine monitoring 832, used to monitor the model; interpretability and inference verification 833, used to browse calculation results, historical data and determine whether to optimize the model.

[0108] In this exemplary embodiment, the evaluation results of the importance values ​​of the features of each round can be recorded and stored in a feature importance evaluation link. In response to a preset operation input by the user, it is determined that the user needs to perform model optimization. The evaluation results of the importance values ​​of the features of the corresponding round can be obtained from the feature importance evaluation link, and the model optimization can be determined based on the evaluation results. For example, the user calculated the feature importance values ​​on the first day, added the results to the feature importance evaluation link, obtained the feature importance values ​​calculated on the first day on the second day, and executed the model optimization process. Fig. 9 A schematic diagram of determining the importance value of a feature in this exemplary embodiment is shown, for example, Fig. 9 The left side shows a schematic diagram of determining the importance value of each initial feature in the first sample data. After calculations of multiple parts such as model training, model prediction, model quality analysis, and model interpretability analysis, the importance values ​​of the seven initial features included in the first sample data can be obtained. The features and their importance values ​​are "Feature 1, 0.91; Feature 2, 0.81; Feature 3, 0.46; Feature 4, 0.22; Feature 5, 0.15; Feature 6, 0.03; Feature 7, 0.01", and then, according to the importance value, the last two features with lower importance values ​​are deleted to obtain the second sample data. Then, the retraining task flow is executed again using the second sample data, and the importance values ​​of the model features after the deleted features are recalculated according to the steps of model training, model prediction, and model interpretability, as shown in FIG. Fig. 9 As shown on the right, it can be determined that the features and importance values ​​are "Feature 1, 0.91; Feature 2, 0.81; Feature 3, 0.46; Feature 4, 0.22; Feature 5, 0.15" respectively. When it is determined that the calculation result of the importance value of the current feature meets the requirements, for example, after deleting the two features at the end of the importance contribution, the index results of each feature have almost no fluctuation, indicating that the deletion of the two features with lower importance values ​​has little effect on the model. The model can be trained based on the sample data composed of the feature data after feature deletion to determine the optimized target model. It is expected that 10% of the storage space can be saved and the calculation efficiency can be improved by 15%. If the calculation result of the current feature importance value does not meet the requirements, the user can decide whether to continue the model optimization. In addition, if the calculation index of the retained feature fluctuates greatly before and after the feature is deleted or exceeds the preset threshold, it means that the deletion of the feature has a greater impact on the model. The model trained in the previous round can be returned as the target model, or a prompt message can be sent to the user to confirm whether to continue the optimization, or the model with better calculation index in the latest round can be used as the target model, etc. This exemplary embodiment is based on constructing a verification link and inferring the verification link, so that users can quickly conduct one-click experiments to evaluate the model according to the feature model interpretability index. The automated execution link can greatly improve efficiency and save human resources.

[0109] Fig.10 The figure shows a schematic diagram of the inference verification results of feature deletion of the model in a music platform scenario. It can be seen from the figure that after deleting 15 features, the storage space 1010 decreased by 24%, but the AUC indicator remained almost unchanged, achieving the optimization of the model's storage cost by 24%, the number of features by 33%, and the computing cost by 30% without affecting the actual online effect.

[0110] In this exemplary embodiment, the visualization display interface may include a model evaluation result visualization interface, which may involve model information, calculation information, reliability diagrams, and scoring distribution diagrams. Fig.11 The figure shows a comparison of two reliability graphs. The reliability graph can reflect the reliability of different prediction values ​​of the data output in the test set. According to the reliability graph, users can find out in advance whether the model output results meet expectations, infer whether feature crossing, overfitting, underfitting and other problems occur during model training, and make appropriate model optimization based on this. Fig.11 The reliability graph shown in (a) shows overestimation, and the model is overconfident, which may be due to feature crossing during training. Fig.11 The reliability graph shown in (b) shows an underestimation, indicating that the model is underfitting, which may be due to problems with the dataset. Fig.12 A schematic diagram of the score distribution is shown. The score distribution can reflect the score distribution state of the data in the test set during the calculation process. Different curves correspond to different data sets.

[0111] Exemplary Devices

[0112] The exemplary embodiment of the present disclosure also provides a model optimization device. Fig.13 As shown, the model optimization device 1300 may include the following program modules: a sample data acquisition module 1310, used to acquire an initial model and multiple first sample data; wherein each of the first sample data includes multiple initial features; a prediction result acquisition module 1320, used to process the multiple first sample data using the initial model to obtain a prediction result; an importance determination module 1330, used to determine the importance value of each of the initial features according to the prediction result; a target feature determination module 1340, used to determine the target feature from the multiple initial features of the first sample data according to the importance value, so as to obtain second sample data including the target feature; a target model acquisition module 1350, used to train the initial model using the second sample data to obtain a target model after model optimization.

[0113] In one embodiment, the prediction result acquisition module 1320 includes: a first result determination unit, which is used to use the initial model to process the multiple first sample data to obtain a first prediction result; a second result determination unit, which is used to perform feature enhancement processing on the multiple first sample data to obtain multiple third sample data, and use the initial model to process the multiple third sample data to obtain a second prediction result; the importance determination module 1330 includes: an importance value calculation unit, which is used to calculate the importance value of each of the initial features based on the first prediction result and the second prediction result.

[0114] In one embodiment, the second result determination unit includes: a feature processing subunit, configured to extract a plurality of sampling data from the plurality of first sample data, and rearrange or change features in the sampling data to obtain the plurality of third sample data.

[0115] In one embodiment, after extracting multiple sampling data from the multiple first sample data, the model optimization device also includes: a dense data conversion unit, which is used to determine whether each of the sampling data is sparse data based on the characteristic value of the initial feature in each of the sampling data; if the sampling data is determined to be sparse data, the sampling data is converted into dense data.

[0116] In one implementation, the target feature determination module 1340 includes: a target feature determination unit configured to use, among a plurality of initial features of the first sample data, an initial feature whose importance value satisfies a first preset condition as the target feature.

[0117] In one embodiment, the target model acquisition module 1350 includes: an intermediate model acquisition unit, which is used to use the second sample data to train the initial model and determine the intermediate model; a feature importance value re-determination unit, which is used to use the target feature in the second sample data as the initial feature and use the intermediate model to re-determine the importance value of each initial feature in the second sample data; a target model determination unit, which is used to use the intermediate model as the target model when the importance value of each initial feature meets the second preset condition.

[0118] In one embodiment, the importance determination module 1330 includes: a visual display unit, which is used to calculate the importance value of each of the initial features according to the prediction results, and visually display the importance value of each of the initial features on a display interface.

[0119] In addition, other specific details of the embodiments of the present disclosure have been described in detail in the embodiments of the above method and will not be repeated here.

[0120] Exemplary Storage Media

[0121] The exemplary embodiments of the present disclosure also provide a computer-readable storage medium having a computer program stored thereon, which implements the above-mentioned method of the present disclosure when the computer program is executed by a processor. The above-mentioned method can be implemented by a program product, such as a portable compact disk read-only memory (CD-ROM) and including program code, and can be run on a device, such as a personal computer. However, the program product of the present disclosure is not limited thereto, and in this document, a readable storage medium can be any tangible medium containing or storing a program, which can be used by or in combination with an instruction execution system, an apparatus or a device.

[0122] The program product may use any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0123] Computer readable signal media may include data signals propagated in baseband or as part of a carrier wave, in which readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Readable signal media may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0124] The program code contained on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RE, etc., or any suitable combination of the foregoing.

[0125] Program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, etc., and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user computing device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., using an Internet service provider to connect through the Internet).

[0126] Exemplary Electronic Devices

[0127] The exemplary embodiment of the present disclosure also provides an electronic device, which may be Figure 1 Any device in the present invention. The electronic device includes a processor and a memory, and the memory is used to store executable instructions of the processor. The processor is configured to execute the above method of the present invention by executing the executable instructions.

[0128] refer to Fig.14 An electronic device according to an exemplary embodiment of the present disclosure is described. Fig.14 The electronic device 1400 shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0129] like Fig.14 As shown, the electronic device 1400 is in the form of a general computing device. The components of the electronic device 1400 may include, but are not limited to: at least one processing unit 1410, at least one storage unit 1420, and a bus 1430 connecting different system components (including the storage unit 1420 and the processing unit 1410).

[0130] The storage unit stores program codes, which can be executed by the processing unit 1410, so that the processing unit 1410 performs the steps according to various exemplary embodiments of the present disclosure described in the above “Exemplary Method” section of this specification. For example, the processing unit 1410 can perform the following steps: Figure 2 or Figure 3 The method steps shown, etc.

[0131] The storage unit 1420 may include a volatile storage unit, such as a random access storage unit (RAM) 1421 and / or a cache storage unit 1422 , and may further include a read-only storage unit (ROM) 1423 .

[0132] The storage unit 1420 may also include a program / utility 1424 having a set (at least one) of program modules 1425, such program modules 1425 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0133] The bus 1430 may include a data bus, an address bus, and a control bus.

[0134] The electronic device 1400 may also communicate with one or more external devices 1500 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), and such communication may be performed via an input / output (I / O) interface 1440. The electronic device 1400 may also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter 1450. As shown, the network adapter 1450 communicates with other modules of the electronic device 1400 via a bus 1430. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 1400, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0135] It should be noted that, although several modules or submodules of the device are mentioned in the above detailed description, such division is merely exemplary and not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more units / modules described above can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided into multiple units / modules to be embodied.

[0136] In addition, although the operations of the disclosed method are described in a specific order in the drawings, this does not require or imply that the operations must be performed in this specific order, or that all the operations shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.

[0137] Although the spirit and principle of the present disclosure have been described with reference to several specific embodiments, it should be understood that the present disclosure is not limited to the disclosed specific embodiments, and the division of various aspects does not mean that the features in these aspects cannot be combined to benefit, and such division is only for the convenience of expression. The present disclosure is intended to cover various modifications and equivalent arrangements included in the spirit and scope of the attached claims.

Claims

1. A model optimization method, characterized in that: include: Acquire an initial model and a plurality of first sample data; wherein each of the first sample data includes a plurality of initial features; Using the initial model to process the plurality of first sample data to obtain a prediction result; Determining the importance value of each of the initial features according to the prediction result; Determining a target feature from a plurality of initial features of the first sample data according to the importance value, so as to obtain second sample data including the target feature; The initial model is trained using the second sample data to obtain a target model after model optimization.

2. The method according to claim 1, characterized in that The using the initial model to process the plurality of first sample data to obtain a prediction result includes: Using the initial model to process the plurality of first sample data to obtain a first prediction result; Performing feature enhancement processing on the plurality of first sample data to obtain a plurality of third sample data, and using the initial model to process the plurality of third sample data to obtain a second prediction result; Determining the importance value of each of the initial features according to the prediction result includes: According to the first prediction result and the second prediction result, the importance value of each of the initial features is calculated.

3. The method according to claim 2, characterized in that The performing feature enhancement processing on the plurality of first sample data to obtain a plurality of third sample data includes: A plurality of sampling data are extracted from the plurality of first sampling data, and features in the sampling data are rearranged or changed to obtain the plurality of third sampling data.

4. The method according to claim 3, characterized in that: After extracting a plurality of sampling data from the plurality of first sample data, the method further comprises: Determining whether each of the sampled data is sparse data according to a feature value of an initial feature in each of the sampled data; If it is determined that the sampled data is sparse data, dense data conversion processing is performed on the sampled data.

5. The method according to claim 1, characterized in that The step of determining a target feature from a plurality of initial features of the first sample data according to the importance value comprises: Among the multiple initial features of the first sample data, the initial features whose importance values ​​meet the first preset condition are used as the target features.

6. The method according to claim 1, characterized in that The step of using the second sample data to train the initial model to obtain a target model after model optimization includes: Using the second sample data to train the initial model and determine an intermediate model; Taking the target feature in the second sample data as the initial feature, and using the intermediate model to determine again the importance value of each initial feature in the second sample data; When the importance value of each of the initial features meets a second preset condition, the intermediate model is used as the target model.

7. The method according to claim 1, characterized in that Determining the importance value of each of the initial features according to the prediction result includes: According to the prediction results, the importance value of each of the initial features is calculated, and the importance value of each of the initial features is visually displayed on a display interface.

8. A model optimization device, characterized in that: include: A sample data acquisition module, used to acquire an initial model and a plurality of first sample data; wherein each of the first sample data includes a plurality of initial features; A prediction result obtaining module, used for processing the plurality of first sample data using the initial model to obtain a prediction result; An importance determination module, used to determine the importance value of each of the initial features according to the prediction result; a target feature determination module, configured to determine a target feature from a plurality of initial features of the first sample data according to the importance value, so as to obtain second sample data including the target feature; The target model acquisition module is used to train the initial model using the second sample data to obtain a target model after model optimization.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

10. An electronic device, characterized in that: include: processor; as well as A memory, configured to store executable instructions of the processor; The processor is configured to perform the method of any one of claims 1 to 7 by executing the executable instructions.