Hyper-parameter space analysis method and device, electronic equipment and storage medium

Through the method combined with deep learning model and CMA-ES algorithm, Sharple value and boundary eigenvalue are calculated, which solves the problem of inefficient determination of hyperparameter space boundary in precision systems, and achieves efficient and accurate hyperparameter space analysis and optimization.

CN119988930APending Publication Date: 2025-05-13NANKAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411957989.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-09-11
Filing Date
2024-12-27
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art is inefficient in determining the boundaries of the hyperparameter space of precision systems and relies on expert experience, which makes it difficult to control the accuracy and affects the system operation effect.

Method used

By obtaining the hyperparameter data of the precision system, using deep learning models for spatial modeling, calculating Shapley values ​​and boundary eigenvalues, and combining the CMA-ES algorithm to perform spatial optimization to determine the boundary eigenvalues ​​of the hyperparameter space.

Benefits of technology

Accurate and efficient analysis of the ultra-parameter space of precision systems is achieved, reducing dependence on human experience, and improving the efficiency and accuracy of system settings and optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988930A_ABST
    Figure CN119988930A_ABST
Patent Text Reader

Abstract

The invention provides a hyper-parameter space analysis method and device, electronic equipment and a storage medium. The method provided by the embodiment of the invention comprises the following steps: acquiring hyper-parameter data of the precision system; performing spatial modeling through a deep learning model by using the hyper-parameter data, so that the deep learning model represents a mapping relationship between a first dimension and a second dimension in a hyper-parameter space corresponding to the hyper-parameter data; calculating a Shapley value of the deep learning model, wherein the Shapley value is used for indicating the influence degree of the second dimension on the first dimension; using a covariance matrix adaptive evolution (CMA-ES) algorithm to perform spatial optimization in a hyper-parameter space for the deep learning model to obtain a boundary feature value, wherein the boundary feature value indicates the position of a point which changes fastest on a first dimension in the hyper-parameter space; the Shapley value and the boundary characteristic value are used for setting or optimizing the precision system. According to the embodiment of the invention, efficient, accurate and low-cost setting or optimization of the precision system can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing, and specifically to a hyperparameter space analysis method, device, electronic device and storage medium. Background Art

[0002] A precision system refers to a system that integrates advanced technology and high-precision components, and is designed to achieve specific, high-precision functions and tasks. It is increasingly used in various fields of real life. Usually for a precision system, there are many factors that affect its operating performance. These factors can be characterized by parameters and parameter values. Given the large number of types and wide range of values, they are often figuratively referred to as "super parameters" or "hyperparameters." How to select and set the range of parameters largely determines the operating effect of the precision system. Given the complexity of the operating environment, system conditions, operating conditions, etc. of the precision system, there are many types of parameters that affect the precision system. How to select parameters and determine the range of values, that is, determine the boundaries of the hyperparameter space, has become an urgent problem to be solved.

[0003] Currently, the determination of hyper-parameter space is mostly based on human experience. However, this method is inefficient on the one hand, and heavily relies on expert experience on the other hand. It is costly and difficult to control accuracy, which seriously affects the operating performance of the precision system. Summary of the invention

[0004] In view of this, the embodiments of the present application are committed to providing a hyper-parameter space analysis method, device, electronic device and storage medium to achieve accurate and efficient analysis of the hyper-parameter space, so as to facilitate the setting and optimization of the precision system.

[0005] In one aspect of the present application, a hyperparameter space analysis method is provided, comprising:

[0006] Obtaining hyperparameter data of a precision system, wherein the hyperparameter data has multiple dimensions, each dimension representing a hyperparameter, and the hyperparameter is a parameter that affects the operation effect of the precision system;

[0007] Using the hyperparameter data to perform spatial modeling through a deep learning model, so that the deep learning model represents a mapping relationship between a first dimension and a second dimension in a hyperparameter space corresponding to the hyperparameter data;

[0008] Calculating a Shapley value of the deep learning model, where the Shapley value is used to indicate the degree of influence of the second dimension on the first dimension;

[0009] Use the covariance matrix adaptive evolution CMA-ES algorithm to perform spatial optimization in the hyperparameter space for the deep learning model to obtain a boundary eigenvalue, wherein the boundary eigenvalue indicates the position of the fastest changing point in the first dimension in the hyperparameter space, and the position includes the hyperparameter value range of each dimension in the second dimension;

[0010] The Shapley value and the boundary eigenvalue are used to set or optimize the precision system;

[0011] The first dimension includes one or more different dimensions in the hyperparameter data, and the second dimension includes all dimensions in the hyperparameter data except the first dimension.

[0012] In some implementations, the method further includes: performing data preprocessing on the hyperparameter data, wherein the data preprocessing includes one or more of the following: data cleaning and normalization.

[0013] In some embodiments of the first aspect of the present application, the hyperparameter data includes at least one of the following hyperparameters: environmental parameters, component parameters of the precision system, motion parameters and performance parameters of the precision system.

[0014] In some embodiments, the method further includes: training and testing the deep learning model using the hyperparameter data so that the deep learning model can characterize the mapping relationship between the first dimension and the second dimension in the hyperparameter space corresponding to the hyperparameter data.

[0015] In some embodiments, the use of hyperparameter data to train and test the deep learning model includes: dividing the hyperparameter data into a training data set and a test data set; training the deep learning model with the values ​​of the data on the first dimension of the hyperparameter data in the training data set as the true values ​​of the output features and the values ​​of the data on the second dimension of the hyperparameter data in the training data set as the input features, and the output features of the deep learning model are the predicted values ​​of the data on the first dimension; inputting the values ​​of the data on the second dimension of the hyperparameter data in the test data set into the deep learning model to obtain the predicted values ​​of the data on the first dimension, and evaluating the deep learning model based on the values ​​of the data on the first dimension of the hyperparameter data in the test data set and the predicted values ​​of the data on the first dimension.

[0016] In some implementations, the deep learning model is constructed by one of the following:

[0017] Multilayer Perceptron;

[0018] Convolutional Neural Networks;

[0019] Recurrent Neural Networks;

[0020] Long short-term memory networks;

[0021] Attention mechanism model;

[0022] An ensemble learning model that includes machine learning decision trees.

[0023] In some embodiments, calculating the Shapley value of the deep learning model includes: creating an interpreter, which is used to estimate the Shapley value of the deep learning model; using the interpreter to call the explainer.shap_values ​​method to perform Shapley calculation in the hyperparameter space described by the deep learning model to obtain a Shapley value sequence for each dimension in the second dimension, wherein the Shapley value sequence is used to indicate the degree of influence of the numerical value of data in one dimension in the second dimension on the numerical value of data in the first dimension.

[0024] In some implementations, the method further includes: visualizing a sequence of Shapley values ​​of each dimension in the second dimension.

[0025] In one aspect of the present application, a hyperparameter spatial analysis device is provided, comprising:

[0026] An acquisition unit, used to acquire hyperparameter data of a precision system, wherein the hyperparameter data has multiple dimensions, each dimension represents a hyperparameter, and the hyperparameter is a parameter that affects an operation effect of the precision system;

[0027] A modeling unit, configured to perform spatial modeling using a deep learning model using the hyperparameter data, so that the deep learning model represents a mapping relationship between a first dimension and a second dimension in a hyperparameter space corresponding to the hyperparameter data;

[0028] A Shapley calculation unit, used to calculate a Shapley value of the deep learning model, wherein the Shapley value is used to indicate the degree of influence of the second dimension on the first dimension;

[0029] An eigenvalue determination unit, used to use a covariance matrix adaptive evolution CMA-ES algorithm to perform spatial optimization in a hyperparameter space for the deep learning model to obtain a boundary eigenvalue, wherein the boundary eigenvalue indicates the position of the fastest changing point in a first dimension in the hyperparameter space, and the position includes a hyperparameter value range of each dimension in the second dimension;

[0030] The Shapley value and the boundary eigenvalue are used to set or optimize the precision system;

[0031] The first dimension includes one or more different dimensions in the hyperparameter data, and the second dimension includes all dimensions in the hyperparameter data except the first dimension.

[0032] In one aspect of the present application, an electronic device is provided, including: a processor and a memory;

[0033] Wherein, the memory is connected to the processor, and the memory is used to store a computer program;

[0034] The processor is used to implement the above method by running the computer program stored in the memory.

[0035] In one aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the above method is implemented.

[0036] According to the embodiments of the present application, by combining the deep learning model, the CMA-ES algorithm and the SHAP technology, it is possible to analyze the degree of influence of the second dimension in the hyperparameter data on the first dimension, and then determine the position of the fastest changing point in the first dimension in the hyperparameter space. Based on this, the precision system can be set up or optimized without relying on human experience, which can improve efficiency and accuracy and reduce labor costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0038] Figure 1 It is a flowchart of the hyperparameter space analysis method provided in the embodiment of the present application;

[0039] Figure 2 is an exemplary implementation flow chart of spatial modeling using a deep learning model using hyperparameter data according to an embodiment of the present application;

[0040] Figure 3 is a structural schematic diagram of a hyperparameter spatial analysis device provided in an embodiment of the present application;

[0041] Figure 4 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0042] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0043] Terminology explanation:

[0044] Hyperparameter space: hyperparameter space, but the "hyperparameter" involved in the embodiments of the present application refers to the parameters that affect the operating effect of the precision system, which are vividly called "hyperparameters" or "super parameters" because of their many types and wide range of values. Accordingly, the hyperparameter space, i.e., the hyperparameter space (Hyperparameter Space), refers to the set of all possible hyperparameter values. In the hyperparameter space, each dimension represents a hyperparameter, and different dimensions have different degrees of correlation, and each point represents a specific set of hyperparameter settings. For example, if there are four hyperparameters: wind direction, wind speed, air pressure, and temperature, then the hyperparameter space is a four-dimensional space. The size and complexity of the hyperparameter space depends on the number of hyperparameters and the range of their possible values. In view of the large number of hyperparameter types that affect the operating effect of the precision system, the hyperparameter space is usually a high-dimensional space, and a high-dimensional space or a hyperparameter with a large number of possible values ​​makes it more difficult to search for the optimal hyperparameter configuration, which is the so-called "dimensionality disaster". Therefore, choosing a suitable hyperparameter space is crucial for precision system tuning.

[0045] Hyperparameter space boundary: the boundary of the hyperparameter space, which refers to a complete set of possible value ranges of the hyperparameters. The boundary of the hyperparameter space has an important impact on the setting, testing, and optimization of precision systems. In the disclosed embodiment, the hyperparameter space boundary refers to the sampling interval where the other dimensions in the hyperparameter space have the greatest impact on a specified dimension.

[0046] Shapley Additive Explanations (SHAP): is a widely used explanation technique. SHAP quantifies and explains the contribution of features to model output based on Shapley values, providing a consistent and fair explanation framework so that the impact of each feature on model output can be understood and explained.

[0047] Exemplary Methods

[0048] Figure 1 FIG. 1 is a flow chart of a hyperparameter space analysis method provided by an embodiment of the present disclosure. Figure 1 , the method may include the following steps:

[0049] Step 101, obtaining hyperparameter data of a precision system, where the hyperparameter data has multiple dimensions, each dimension representing a hyperparameter, and the hyperparameter is a parameter that affects the operation effect of the precision system;

[0050] Step 102, using the hyperparameter data to perform spatial modeling through a deep learning model, so that the deep learning model represents a mapping relationship between a first dimension and a second dimension in a hyperparameter space corresponding to the hyperparameter data;

[0051] The first dimension includes one or more different dimensions in the hyperparameter data, and the second dimension includes all dimensions in the hyperparameter data except the first dimension.

[0052] Step 103, calculating the Shapley value of the deep learning model, where the Shapley value is used to indicate the degree of influence of the second dimension on the first dimension;

[0053] Step 104, use the covariance matrix adaptive evolution CMA-ES algorithm to perform spatial optimization in the hyperparameter space for the deep learning model to obtain boundary eigenvalues, where the boundary eigenvalues ​​indicate the position of the fastest changing point in the first dimension in the hyperparameter space, and the position includes the hyperparameter value range of each dimension in the second dimension. The above-mentioned Shapley value and boundary eigenvalue are used to set or optimize the precision system.

[0054] In the disclosed embodiment, the hyperparameter data is a data set of hyperparameters of a precision system, which may include all parameters that affect the operation of the precision system. The precision system may be any type of precision system. For example, aircraft, guided weapons, precision positioning systems, etc. in the aerospace field. Accordingly, the hyperparameters that affect the precision system may include, but are not limited to, environmental parameters, component parameters of the precision system, motion parameters of the precision system, performance parameters, etc.

[0055] Environmental parameters may include parameters of the production or manufacturing environment of precision systems, parameters of the transportation environment, parameters of the operating environment, etc. For example, environmental parameters for manufacturing aircraft include: manufacturing temperature, humidity, manufacturing equipment accuracy, etc. Aircraft operating environment parameters include wind direction, wind speed, air pressure, temperature, space altitude particles, space magnetic field, noise, vibration, etc.

[0056] The number of component parameters of precision systems is even greater. Taking aircraft as an example, they can include frame parameters, power system parameters, control system parameters, etc. Frame parameters can specifically include weight, wheelbase, materials, etc. Power system parameters can include motor parameters (such as nominal no-load KV value, maximum peak current, maximum peak power, etc.), battery parameters (such as capacity, voltage, discharge rate, etc.), propeller parameters (such as pitch, model, chord length, number of blades, safe speed, etc.). Control system parameters include receiver parameters (such as frequency, modulation mode, number of channels, remote control distance), autopilot parameters (such as GPS, IMU, barometer, etc.).

[0057] The motion parameters of a precision system may include speed, attitude, acceleration, etc. Taking an aircraft as an example, they may include take-off speed, cruising speed, maximum speed, glide speed, climb rate, turning radius, pitch angle, yaw angle, roll angle, airflow angle, angular velocity, acceleration, etc.

[0058] Taking an aircraft as an example, performance parameters may include: ceiling, range, endurance, activity radius, maneuverability parameters, take-off performance parameters, landing performance parameters, stability parameters, target hitting accuracy, etc.

[0059] As can be seen from the above description, there are usually a large number of types of hyperparameters that affect the operation of precision systems.

[0060] In some embodiments, if there are N kinds of hyperparameters of the model, N is an integer greater than 1, the hyperparameter data of the model is N-dimensional data, and the hyperparameter space corresponding to the hyperparameter data is N-dimensional space. If the model has thousands or more hyperparameters, its hyperparameter data is high-dimensional data of thousands or ten thousand dimensions, and the corresponding hyperparameter space is high-dimensional space such as thousands of dimensions and ten thousand dimensions. Exemplarily, the hyperparameter data can be represented as multidimensional feature data (X1, X2, ..., Xn), n represents the total number of dimensions of the hyperparameter data, n is an integer greater than 1, Xi (i = 1, 2, ..., n) represents the data on dimension i, that is, the possible values ​​of hyperparameter i.

[0061] In step 101, the hyper-parameter data may be input by a staff member or may come from other electronic devices. The specific method of obtaining the hyper-parameter data is not limited in the embodiments of the present disclosure.

[0062] In some implementations, before step 102, the following may also be included: performing data preprocessing on the hyperparameter data, where the data preprocessing includes one or more of the following: data cleaning and normalization. Data preprocessing can effectively improve the data quality of the hyperparameter data, reduce the risk of errors, and provide a reliable basis for subsequent processing.

[0063] In some examples, data cleaning can include checking, correcting, and screening hyperparameter data to ensure the accuracy and completeness of the hyperparameter data.

[0064] In some examples, the normalization process may include: unifying the units and scales of data of different dimensions in the hyperparameter data. By normalizing the hyperparameter data, the dimensional differences between different features may be eliminated.

[0065] In other implementations, data preprocessing may also include: inconsistency processing, missing value processing, etc. The specific process and implementation of data preprocessing are not limited in the embodiments of the present disclosure.

[0066] In step 102, using hyperparameter data to perform spatial modeling through a deep learning model involves steps such as parameter initialization, hyperparameter optimization, training, and testing of the deep learning model. Parameter initialization and hyperparameter optimization are optional operations. That is, step 102 may include: using hyperparameter data to train and test the deep learning model so that the deep learning model can characterize the mapping relationship between the first dimension and the second dimension in the hyperparameter space corresponding to the hyperparameter data.

[0067] For precision systems, one or more performance parameters of the precision system in the hyperparameter data can be used as the first dimension (corresponding to the output features of the deep learning model), and the remaining parameters can be used as the second dimension (corresponding to the input features of the deep learning model). In this way, the influence of other parameters on the performance parameters can be analyzed by the method provided in the embodiment of the present application, and then the position of the point with the fastest change in the performance parameter (i.e., the first dimension) can be determined, i.e., the boundary feature value. Taking an aircraft as an example, one or some specific performance parameters of the aircraft can be used as the first dimension, and other parameters such as various environmental parameters, motion parameters, component parameters, etc. of the aircraft can be used as the second dimension, so as to analyze the position of the point with the fastest change in the hyperparameter space.

[0068] Figure 2 FIG. 1 shows an exemplary implementation flow chart of using hyperparameter data to perform spatial modeling through a deep learning model in step 102. Figure 2 In some implementations, step 102 may include:

[0069] Step 201, divide the hyperparameter data into a training data set and a test data set;

[0070] Specifically, the hyperparameter data can be randomly divided into a training data set and a test data set according to a preset ratio, that is, the training data set contains part of the hyperparameter data, and the test data set contains another part of the hyperparameter data. Appropriate division of the training data set and the test data set can prevent the deep learning model from overfitting.

[0071] Step 202, using the values ​​of the data on the first dimension of the hyperparameter data in the training data set as the true values ​​of the output features, and the values ​​of the data on the second dimension of the hyperparameter data in the training data set as the input features to train the deep learning model, and the output features of the deep learning model are the predicted values ​​of the data on the first dimension;

[0072] Specifically, grid search and other methods can be used to determine the hyperparameter value, and then the hyperparameter is introduced into the deep learning model. The deep learning model is then trained using the data values ​​on the first dimension of the hyperparameter data in the training dataset as the true value of the output feature and the data values ​​on the second dimension of the hyperparameter data in the training dataset as the input feature, until the loss value reaches the requirement and converges.

[0073] Step 203: Input the values ​​of the data on the second dimension of the hyperparameter data in the test data set into the deep learning model to obtain the predicted values ​​of the data on the first dimension, and evaluate the deep learning model based on the values ​​of the data on the first dimension of the hyperparameter data in the test data set and the predicted values ​​of the data on the first dimension.

[0074] Specifically, after the training is completed, the trained deep learning model can be evaluated on the test data set to verify the performance of the deep learning model on the data it has not seen. Specifically, the trained deep learning model is used to predict the test data set, and the loss value between the prediction result (i.e., the predicted value of the data on the first dimension) and the true label (i.e., the value of the data on the first dimension in the test data set) is calculated. If the loss value is less than the current best loss value, the parameters of the deep learning model are updated until the convergence conditions are met. In some embodiments, the loss function used to calculate the loss value can be, but is not limited to, the mean square error.

[0075] Step 204: Save the trained deep learning model.

[0076] Specifically, saving the deep learning model means saving all parameters of the deep learning model, where all parameters include model parameters and hyperparameters of the deep learning model.

[0077] In some implementations, after training is completed, the parameters of the deep learning model can be saved as a file of a specified type and the file can be stored in a specified path for use in subsequent processing.

[0078] In a specific application, the deep learning model in step 102 can be constructed by one of the following: a multilayer perceptron (MLP), a convolutional neural network, a recurrent neural network, a long short-term memory network, an attention mechanism model, an integrated learning model of a machine learning decision tree such as eXtreme Gradient Boosting (XGBoost), etc. Exemplarily, the deep learning model in step 102 can be an MLP model constructed by MLP.

[0079] Taking the deep learning model as an MLP model as an example, the specific implementation process of step 102 may include: randomly dividing the preprocessed hyperparameter data into a training data set and a test data set according to an appropriate proportion, selecting a multilayer perceptron (MLP) as a modeling tool to build an MLP model using a deep learning framework, training the MLP model using the training data set, and testing and evaluating the trained MLP model using the test data set, thereby completing the modeling of the hyperparameter space.

[0080] Among them, the training process of the MLP model may include: first, initializing the MLP model, optimizing the hyperparameters of the MLP model to determine the hyperparameters of the MLP model, using the hyperparameters of the MLP model to train and test the MLP model, and when the MLP model meets the requirements, the training is completed and the parameters of the MLP model are saved.

[0081] In step 102, spatial modeling is performed through a deep learning model to obtain a continuous hyperparameter space, which includes not only the original hyperparameter data, but also the hyperparameter data estimated by the deep learning model. Thus, the discretized hyperparameter data is converted into continuously distributed hyperparameter data through spatial modeling, that is, the coarse-grained hyperparameter data is converted into fine-grained hyperparameters, so that subsequent steps 103 and 104 can perform related processing based on the continuously distributed hyperparameter data, thereby obtaining Shapley values ​​and boundary eigenvalues ​​with finer granularity and higher accuracy.

[0082] In some implementations, step 103 is implemented by SHAP. Specifically, step 103 may include the following steps a1 to a2:

[0083] Step a1, creating an interpreter, which is used to estimate the Shapley value of the deep learning model;

[0084] SHAP supports many types of interpreters, such as deep, gradient, kernel, linear, tree, sampling, etc. In specific applications, you can choose a suitable interpreter as needed. In some examples, step a1 can use Kernel Explainer as the interpreter of the deep learning model.

[0085] Specifically, a file containing deep learning model parameters is loaded to generate a function expression of the deep learning model, and a KernelExplainer interpreter is created based on the function expression of the deep learning model. Two parameters are passed in when creating the interpreter: the function expression of the deep learning model and its input data. The prediction function of the model is used to predict the input data, which is usually preprocessed and standardized input feature data.

[0086] Specifically, you can construct a single-sample explanation object and a full-data explanation object. By creating these two explanation objects, you can easily store the explanation information of the model prediction for subsequent analysis and visualization.

[0087] Step a2, use the explainer to call the explainer.shap_values ​​method to perform Shapley calculation in the hyperparameter space described by the deep learning model to obtain a Shapley value sequence for each dimension in the second dimension. The Shapley value sequence is used to indicate the degree of influence of the numerical value of the data in one dimension in the second dimension on the numerical value of the data in the first dimension.

[0088] Specifically, using the created KernelExplainer interpreter, call the explainer.shap_values ​​method to calculate the Shapley value. This method requires input data to be passed in and is used to calculate the Shapley value on all input data. After the calculation is completed, the Shapley value sequence of each dimension in the second dimension is obtained. The Shapley value sequence is used to indicate the degree of influence of the value of the data in one dimension in the second dimension on the value of the data in the first dimension. These Shapley values ​​can help understand the formation process of the model prediction results and the contribution of different features to the model prediction.

[0089] In some implementations, after step 103, the method may further include: visualizing the Shapley value sequence of each dimension in the second dimension, so that the user can intuitively view the influence of other dimensions on the specified dimension.

[0090] For example, visualizing the Shapley value sequence of each dimension in the second dimension may include: generating a bar chart and an explanation report indicating the Shapley value sequence of each dimension in the second dimension, and displaying the bar chart and the explanation report so that the user can intuitively and clearly view the SHAP calculation results.

[0091] In the disclosed embodiment, the Shapley value obtained in step 103 not only provides an importance ranking of the data values ​​on the second dimension in the hyperparameter data, but also dynamically reveals the changes in the importance of the numerical values ​​of the data on the first dimension in the hyperparameter data in different scenarios, thereby providing a more in-depth and refined hyperparameter space analysis. This dynamic evaluation method can truly reflect the complex interactions between data of different dimensions (i.e., different hyperparameter values) in the hyperparameter space, thereby improving the comprehensiveness and accuracy of the hyperparameter space importance analysis.

[0092] In some implementations, step 104 may include the following steps b1 to b2:

[0093] Step b1, initializing the parameters of the CMA-ES algorithm according to the mapping relationship indicated by the deep learning model;

[0094] The parameters of the CMA-ES algorithm include: the number of parent individuals, the number of search space dimensions, the upper and lower boundaries of the search space, the standard deviation of the population, and adaptive parameters. The search space belongs to the hyperparameter space. The adaptive parameters are used to control the update of the covariance matrix, the update of the standard deviation, and the damping coefficient.

[0095] Specifically, step b1 may include the following items:

[0096] 1) According to the incoming population size, calculate the number of parent individuals mu to be selected, which is usually half of the population size.

[0097] 2) Initialize the search space dimension n_dim.

[0098] 3) Initialize the upper and lower bounds lb and ub of the search space.

[0099] 4) Initialize the standard deviation sigma of the population.

[0100] 5) Initialize the initial mean xmean of the population. If not specified, it is randomly sampled uniformly from the upper and lower bounds of the search space.

[0101] 6) Initialize and select the weight vector weights of the parent individual, which is calculated using the weight formula specific to the CMA-ES algorithm.

[0102] 7) The weighted average of the fitness of individuals in the population, mueff, is calculated.

[0103] 8) Initialize the adaptive parameters cc, cs, c1, cmu and damps. These parameters respectively control the update of the covariance matrix, the update of the standard deviation and the damping coefficient, etc., and play an important role in the convergence and stability of the algorithm.

[0104] Step b2, perform the following processing until the step length reaches the standard deviation of the population to determine the boundary eigenvalue:

[0105] First, calculate the transformation matrix of the population: Calculate the transformation matrix of the population. The transformation matrix is ​​the product of matrix B and matrix D. Matrix B is the eigenvector matrix of the covariance matrix, and matrix D is the eigenvalue diagonal matrix.

[0106] Secondly, generate samples of lambda_pop individuals through a loop, use the samples of each individual to get a random vector by sampling from the standard normal distribution, and use the covariance matrix and the random vector to calculate the true sample vector and its fitness value. The sample vector is within the boundary of the search space.

[0107] Third, the samples are sorted according to the fitness value, and the mu individuals with the highest fitness are selected as the parent individuals. At the same time, the adaptive parameters in the CMA-ES algorithm are updated, such as the adaptive learning rates cc, cs, c1, cmu and damps, as well as the eigenvectors and eigenvalues ​​of the covariance matrix.

[0108] Fourth, calculate the new mean vector xmean and the new transformation matrix based on the parent individuals, and calculate the new step length sigma based on the cumulative sum of the evolutionary path;

[0109] Finally, the covariance matrix is ​​decomposed into eigenvalues ​​to update the transformation matrix and adjust the step size sigma. If a flat area is found in the fitness function, the step size sigma is also adjusted to try to jump out of the flat area.

[0110] Assuming that the hyperparameter data is represented as multidimensional feature data (X1, X2, ..., Xn), the boundary feature value can be expressed as [x1, x2, ..., xn]. Where Xi represents the feature of a certain dimension, and xi represents the position of the point with the fastest change in a certain dimension i.

[0111] Compared with the traditional evolutionary strategy algorithm, the CMA-ES algorithm has stronger adaptability and global search capabilities, and is particularly suitable for high-dimensional, non-convex and complex optimization problems. CMA-ES is a gradient-free optimization and does not use gradient information. It performs well on complex optimization problems such as non-separable, ill-conditioned, orrugged / multi-modal. In the disclosed embodiment, step 104 combines the CMA-ES algorithm to find the point in the output space of the deep learning model that causes the fastest change in the output of the deep learning model, and can more accurately find the value of the data in the second dimension that has the greatest impact on the value of the data in the first dimension. Since the CMA-ES algorithm can efficiently explore high-dimensional space and concentrate on searching the areas that are most sensitive to the output of the deep learning model, it significantly improves the efficiency and accuracy of the spatial optimization process of the hyperparameter space, while also reducing the waste of computing resources, and is particularly suitable for spatial optimization of hyperparameter data with a high number of dimensions.

[0112] In the disclosed embodiment, the hyperparameter data is first used to perform spatial modeling through a deep learning model, and the feature contribution of the model is analyzed through the SHAP technology to better explain the mapping relationship between other dimensions (i.e., the second dimension) and the specified dimension (i.e., the first dimension) in the hyperparameter data, and the CMA-ES algorithm is used to perform spatial optimization in the hyperparameter space for the trained deep learning model to find the spatial optimal point, which is the point where the value of the first dimension in the hyperparameter space changes the fastest. It can be seen that the disclosed embodiment systematically solves the shortcomings of the prior art in the interpretability of hyperparameter data, optimization efficiency and feature importance evaluation by combining deep learning models, CMA-ES algorithms and SHAP technology, and provides a more comprehensive, transparent and efficient hyperparameter space analysis solution.

[0113] After determining the boundary eigenvalue (i.e., the point where the value of the first dimension in the hyperparameter space obtained by the above spatial optimization changes the fastest), it means that the position where the first dimension data changes the fastest is determined. The position refers to the value range of each hyperparameter of the second dimension, and the precision system can be set or optimized accordingly. For example, through the above process, it can be determined based on the Shapley value that the wind speed has a greater impact on the accuracy of the aircraft in capturing the target, and based on the boundary eigenvalue, it can be determined that when the wind speed is between 13 and 20 (m / s), the accuracy of the aircraft in capturing the target changes the fastest, which means that the accuracy of the aircraft in capturing the target is most sensitive to the wind speed in this interval, and the aircraft can be further set or optimized accordingly. For example, the weight of the wind speed parameter within this interval is increased.

[0114] Exemplary Devices

[0115] Figure 3 FIG. 1 shows a schematic diagram of the structure of a hyperparameter data analysis device provided in an embodiment of the present application. Figure 3 , the hyper-parameter data analysis device 300 may include:

[0116] An acquisition unit 301 is used to acquire hyperparameter data of a precision system, where the hyperparameter data has multiple dimensions, each dimension represents a hyperparameter, and the hyperparameter is a parameter that affects the operation effect of the precision system;

[0117] A modeling unit 302 is used to perform spatial modeling using a deep learning model using the hyperparameter data, so that the deep learning model represents a mapping relationship between a first dimension and a second dimension in a hyperparameter space corresponding to the hyperparameter data;

[0118] A Shapley calculation unit 303, used to calculate a Shapley value of the deep learning model, where the Shapley value is used to indicate the degree of influence of the second dimension on the first dimension;

[0119] An eigenvalue determination unit 304 is used to use a covariance matrix adaptive evolution CMA-ES algorithm to perform spatial optimization in a hyperparameter space for a deep learning model to obtain a boundary eigenvalue, where the boundary eigenvalue indicates the position of the point with the fastest change in the first dimension in the hyperparameter space;

[0120] Shapley values ​​and boundary eigenvalues ​​are used to set up or optimize precise systems.

[0121] The first dimension includes one or more different dimensions in the hyperparameter data, and the second dimension includes all dimensions in the hyperparameter data except the first dimension.

[0122] In some implementations, the hyper-parameter data analysis device 300 may further include: a preprocessing unit 305, for performing data preprocessing on the hyper-parameter data, where the data preprocessing includes one or more of the following: data cleaning and normalization.

[0123] In some embodiments, the modeling unit 302 can be specifically used to: divide the hyperparameter data into a training data set and a test data set; use the values ​​of the data on the first dimension of the hyperparameter data in the training data set as the true values ​​of the output features, and the values ​​of the data on the second dimension of the hyperparameter data in the training data set as the input features to train the deep learning model, and the output features of the deep learning model are the predicted values ​​of the data on the first dimension; input the values ​​of the data on the second dimension of the hyperparameter data in the test data set into the deep learning model to obtain the predicted values ​​of the data on the first dimension, and evaluate the deep learning model according to the values ​​of the data on the first dimension of the hyperparameter data in the test data set and the predicted values ​​of the data on the first dimension.

[0124] In some implementations, the Shapley calculation unit 303 can be specifically used to: create an interpreter, which is used to estimate the Shapley value of the deep learning model; use the interpreter to call the explainer.shap_values ​​method to perform Shapley calculations in the hyperparameter space described by the deep learning model to obtain a Shapley value sequence for each dimension in the second dimension, and the Shapley value sequence is used to indicate the degree of influence of the numerical value of data in one dimension in the second dimension on the numerical value of data in the first dimension.

[0125] In some implementations, the hyperparameter data analysis device 300 may further include: a visualization unit 306, configured to visualize the Shapley value sequence of each dimension in the second dimension.

[0126] Other technical details of the hyper-parameter data analysis device 300 in the embodiment of the present application can be found in the previous method section and will not be repeated here.

[0127] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can refer to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment. The system and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without creative work.

[0128] Electronic devices

[0129] Figure 4 FIG. 1 shows an example diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 4 As shown, the electronic device may include: one or more processors 401, and a memory 402 storing one or more programs, which are executed by the one or more processors 401 to implement the method flow shown in the above embodiments of the present disclosure and / or program units corresponding to each unit in the device.

[0130] The various components are interconnected using different buses and can be mounted on a common motherboard or otherwise as needed. The processor 401 can process instructions executed within the electronic device, including instructions stored in or on the memory to display graphical information of a user interface on an external input / output device (such as a display device coupled to the interface). In other embodiments, multiple processors and / or multiple buses can be used with multiple memories and multiple memories if desired.

[0131] The processor 401 may include one or more single-core processors or multi-core processors. The processor 401 may include any combination of general-purpose processors or dedicated processors (such as image processors, application processors, baseband processors, etc.).

[0132] The memory 402 is a computer-readable storage medium provided by the present disclosure, which can be used to store non-transient software programs, non-transient computer executable programs and units, such as the following in the embodiments of the present disclosure: Figure 1 The processor 401 executes the non-transient software programs, instructions and units stored in the memory 402, thereby executing the above method embodiments. Figure 1 The methods shown correspond to programs, instructions, and units.

[0133] The electronic device may further include: an input device 403 and an output device 404. The processor 401, the memory 402, the input device 403 and the output device 404 may be connected via a bus or other means. Figure 4 The example of connecting through bus is taken in the following.

[0134] The input device 403 can receive input digital or character information, and generate signal input related to user settings and function control of the hyperparameter space analysis device, such as a touch screen, a keypad, a mouse, a track pad, a touch pad, an indicator rod, one or more mouse buttons, a trackball, a joystick and other input devices. The output device 404 may include a display device, an auxiliary lighting device (e.g., an LED) and a tactile feedback device (e.g., a vibration motor), etc. The display device may include, but is not limited to, a liquid crystal display (LCD), a light emitting diode (LED) display and a plasma display. In some embodiments, the display device may be a touch screen.

[0135] The above-mentioned programs (also referred to as software, software applications, or codes) include machine instructions for programmable processors, and these computer programs can be implemented using object-oriented programming languages, assembly or machine languages.

[0136] With the development of time and technology, the meaning of medium is becoming more and more extensive, and the propagation path of computer programs is no longer limited to tangible media, and can also be downloaded directly from the network, etc. Any combination of one or more computer-readable storage media can be used. Computer-readable storage media can be used but not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or devices, or any combination of the above. More specific examples (non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this document, computer-readable storage media can be any tangible medium containing or storing programs, which can be used by or in combination with instruction execution systems, devices or devices.

[0137] An embodiment of the present application also provides a computer-readable storage medium on which a computer program is stored. The program includes instructions. When the instructions are executed by one or more processors of a computing device, the steps of the method described in any one of the aforementioned method embodiments are executed.

[0138] The embodiment of the present application also provides a computer program product. When the computer program product is run on a computer, it enables the computer to execute the steps of any one of the methods described in the aforementioned method embodiments.

[0139] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A hyperparameter space analysis method, comprising: Obtaining hyperparameter data of a precision system, wherein the hyperparameter data has multiple dimensions, each dimension representing a hyperparameter, and the hyperparameter is a parameter that affects the operation effect of the precision system; Using the hyperparameter data to perform spatial modeling through a deep learning model, so that the deep learning model represents a mapping relationship between a first dimension and a second dimension in a hyperparameter space corresponding to the hyperparameter data; Calculating a Shapley value of the deep learning model, where the Shapley value is used to indicate the degree of influence of the second dimension on the first dimension; Use the covariance matrix adaptive evolution CMA-ES algorithm to perform spatial optimization in the hyperparameter space for the deep learning model to obtain a boundary eigenvalue, wherein the boundary eigenvalue indicates the position of the fastest changing point in the first dimension in the hyperparameter space, and the position includes the hyperparameter value range of each dimension in the second dimension; The Shapley value and the boundary eigenvalue are used to set or optimize the precision system; The first dimension includes one or more different dimensions in the hyperparameter data, and the second dimension includes all dimensions in the hyperparameter data except the first dimension.

2. The method according to claim 1, characterized in that The hyperparameter data includes at least one of the following hyperparameters: environmental parameters, component parameters of the precision system, motion parameters and performance parameters of the precision system.

3. The method according to claim 1, characterized in that The method also includes: using the hyperparameter data to train and test the deep learning model so that the deep learning model can characterize the mapping relationship between the first dimension and the second dimension in the hyperparameter space corresponding to the hyperparameter data.

4. The method according to claim 3, characterized in that The use of hyperparameter data to train and test the deep learning model includes: Dividing the hyperparameter data into a training data set and a test data set; The deep learning model is trained by taking the values ​​of the data on the first dimension of the hyperparameter data in the training data set as the true values ​​of the output features and the values ​​of the data on the second dimension of the hyperparameter data in the training data set as the input features, and the output features of the deep learning model are the predicted values ​​of the data on the first dimension; The values ​​of the data on the second dimension of the hyperparameter data in the test data set are input into the deep learning model to obtain the predicted values ​​of the data on the first dimension, and the deep learning model is evaluated based on the values ​​of the data on the first dimension of the hyperparameter data in the test data set and the predicted values ​​of the data on the first dimension.

5. The method according to claim 1, characterized in that The deep learning model is constructed by one of the following: Multilayer Perceptron; Convolutional Neural Networks; Recurrent Neural Networks; Long short-term memory networks; Attention mechanism model; An ensemble learning model that includes machine learning decision trees.

6. The method according to claim 1, characterized in that Calculating the Shapley value of the deep learning model includes: Creating an interpreter for estimating the Shapley value of the deep learning model; The interpreter is used to perform Shapley calculations in the hyperparameter space described by the deep learning model to obtain a sequence of Shapley values ​​for each dimension in the second dimension, wherein the sequence of Shapley values ​​is used to indicate the degree of influence of a numerical value of data in one dimension in the second dimension on a numerical value of data in the first dimension.

7. The method according to claim 5, characterized in that Also includes: Visualize the sequence of Shapley values ​​for each of the second dimensions.

8. A hyperparameter spatial analysis device, characterized in that: include: An acquisition unit, used to acquire hyperparameter data of a precision system, wherein the hyperparameter data has multiple dimensions, each dimension represents a hyperparameter, and the hyperparameter is a parameter that affects the operation effect of the precision system; A modeling unit, configured to perform spatial modeling using a deep learning model using the hyperparameter data, so that the deep learning model represents a mapping relationship between a first dimension and a second dimension in a hyperparameter space corresponding to the hyperparameter data; A Shapley calculation unit, used to calculate a Shapley value of the deep learning model, wherein the Shapley value is used to indicate the degree of influence of the second dimension on the first dimension; An eigenvalue determination unit, used to use a covariance matrix adaptive evolution CMA-ES algorithm to perform spatial optimization in a hyperparameter space for the deep learning model to obtain a boundary eigenvalue, wherein the boundary eigenvalue indicates the position of the fastest changing point in a first dimension in the hyperparameter space, and the position includes a hyperparameter value range of each dimension in the second dimension; The Shapley value and the boundary eigenvalue are used to set or optimize the precision system; The first dimension includes one or more different dimensions in the hyperparameter data, and the second dimension includes all dimensions in the hyperparameter data except the first dimension.

9. An electronic device, characterized in that: include: Processor and memory; Wherein, the memory is connected to the processor, and the memory is used to store a computer program; The processor is configured to implement the method according to any one of claims 1 to 7 by running the computer program stored in the memory.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Software defect prediction model based on deep neural network and probabilistic decision forest

    CN109446090A

  • Hyper-parameter space analysis method and device, electronic equipment and storage medium

    CN119089175A