Training method for classification model, search method for hyperparameters, and device
By introducing scaling invariant linear layer and target training method into the neural network, the problem of superparameter adjustment consumes a lot of computing resources during the training process is solved, and efficient computing resource utilization and precision maintenance are achieved.
Patent Information
- Application Number
- CN202010787442.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-08-07
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2040-08-07
AI Technical Summary
In the prior art, a large amount of manpower and computing resources are required during the training of neural networks to adjust hyperparameters, especially learning rate and weight decay, resulting in inefficient training.
By introducing a scaling invariant linear layer and adopting the target training method, the module length of the weight parameters is fixed, so that only the target hyperparameters need to be configured during the training process, achieving the effect of reducing computing resource consumption.
While ensuring the accuracy of the classification model, it significantly reduces the consumption of computing resources during the training process, improves the efficiency of hyperparameters, and reduces the search cost.
Smart Images

Figure CN114078195B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and more particularly, to a method for training a classification model, a method for searching hyperparameters, and an apparatus. Background Art
[0002] Artificial intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a branch of computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making. Research in the field of artificial intelligence includes robots, natural language processing, computer vision, decision-making and reasoning, human-computer interaction, recommendation and search, AI basic theory, etc.
[0003] A neural network usually contains millions or more trainable parameter weights, and these parameters can be trained through a series of optimization algorithms; among these parameters, there are also some parameters that need to be determined in advance, called hyperparameters. Hyperparameters have a significant impact on the training effect of the neural network. Inappropriate hyperparameters often lead to non-convergence or poor training effect of the neural network; however, currently, the setting of hyperparameters usually relies on manual experience and requires multiple adjustments, thus consuming a large amount of human and computing resources.
[0004] Therefore, how to effectively reduce the computing resources consumed in training a classification model has become a technical problem that urgently needs to be solved. Summary of the Invention
[0005] This application provides a method for training a classification model, a method for searching hyperparameters, and an apparatus. Through the method for training a classification model provided by the embodiments of this application, the computing resources consumed in training the classification model can be reduced while ensuring the accuracy of the classification model.
[0006] Furthermore, the method for searching hyperparameters provided by the embodiments of this application can reduce the search space of hyperparameters, improve the search efficiency of hyperparameters; at the same time, reduce the search cost of hyperparameters.
[0007] In a first aspect, there is provided a method for training a classification model, including:
[0008] Obtain the target hyperparameters of the classification model to be trained, where the target hyperparameters are used to control the gradient update step size of the classification model to be trained. The classification model to be trained includes a scale-invariant linear layer, and the scale-invariant linear layer enables the predicted classification results output when the weight parameters of the classification model to be trained are multiplied by any scaling factor to remain unchanged; update the weight parameters of the classification model to be trained according to the target hyperparameters and the target training method to obtain the trained classification model, and the target training method enables the norms of the weight parameters of the classification model to be trained before and after the update to be the same.
[0009] It should be understood that when a neural network has scale invariance, the step size of its weight parameters is inversely proportional to the square of the norm of the weight parameters; among them, scale invariance means that when the weight w is multiplied by any scaling factor, the output y of this layer can remain unchanged. The main function of the weight decay hyperparameter is to constrain the norm of the parameters, so as to avoid the step size from decreasing sharply as the norm of the weight parameters increases. However, since weight decay is an implicit constraint, the norm of the weight parameters will still change during the training process, resulting in the need for manual adjustment of the optimal weight decay hyperparameters for different models. In the embodiments of the present application, the norm of the weight parameters can be explicitly fixed through the target training method, that is, the norms of the weight parameters of the classification model to be trained before and after the update are the same, so that the influence of the norm of the weight parameters on the step size remains constant, achieving the effect of weight decay, and at the same time, manual adjustment is no longer required. Since the premise of adopting the target training method is that the neural network needs to have scale invariance, generally for most neural networks, most layers have scale invariance, but the classification layers in the classification model are mostly scale-sensitive; therefore, in the embodiments of the present application, the original classification layer is replaced by a scale-invariant linear layer, so that each layer of the classification model to be trained has scale invariance.
[0010] In the embodiments of the present application, by replacing the classification layer of the classification model to be trained with a scale-invariant linear layer and adopting the target training method, that is, fixing the norms of the weight parameters of the classification model to be trained before and after the update, two important hyperparameters, namely the learning rate and weight decay, can be equivalently replaced by the target hyperparameters when training the classification model to be trained; therefore, when training the classification model to be trained, only the target hyperparameters of the model to be trained need to be configured, so as to reduce the computing resources consumed for training the classification model to be trained.
[0011] In a possible implementation manner, the above classification model may include any neural network model with a classification function; for example, the classification model may include, but is not limited to: detection models, segmentation models, and recognition models, etc.
[0012] In a possible implementation, the target hyperparameter may refer to an equivalent learning rate. The role of the equivalent learning rate in the classification model to be trained can be regarded as the same as or similar to the role of the learning rate. The equivalent learning rate is used to control the gradient update step size of the classification model to be trained.
[0013] Combined with the first aspect, in some implementations of the first aspect, the weight parameters of the trained classification model are obtained by iteratively updating through the backpropagation algorithm according to the target hyperparameter and the target training method.
[0014] Combined with the first aspect, in some implementations of the first aspect, the scaling-invariant linear layer obtains the predicted classification result according to the following formula:
[0015]
[0016] where Y i represents the predicted classification result corresponding to the weight parameters updated in the i-th iteration; W i represents the weight parameters updated in the i-th iteration; X represents the feature to be classified; S represents the scale constant.
[0017] It should be noted that the above formula is an example for the scaling-invariant processing. The purpose of the scaling-invariant linear layer is to make the weight w multiplied by any scaling factor, and the output y of this layer can remain unchanged; in other words, other specific processing methods can also be adopted for the scaling-invariant processing, and the present application does not make any limitation thereto.
[0018] Combined with the first aspect, in some implementations of the first aspect, the target training method includes:
[0019] Processing the updated weight parameters through the following formula to make the norm lengths of the weight parameters before and after the update of the classification model to be trained the same:
[0020]
[0021] where W i+1 represents the weight parameters updated in the (i + 1)-th iteration; W i represents the weight parameters updated in the i-th iteration; Norm0 represents the initial weight norm length of the classification model to be trained.
[0022] The second aspect provides a method for searching hyperparameters, including:
[0023] Obtain candidate values of a target hyperparameter, where the target hyperparameter is used to control the gradient update step size of a classification model to be trained, and the classification model to be trained includes a scale-invariant linear layer, and the scale-invariant linear layer is used to make the predicted classification result output when the weight parameters of the classification model to be trained are multiplied by any scaling factor remain unchanged;
[0024] According to the candidate values and a target training method, obtain performance parameters of the classification model to be trained, where the target training method makes the norms of the weight parameters of the classification model to be trained before and after updating the same, and the performance parameters include the accuracy of the classification model to be trained;
[0025] Determine the target value of the target hyperparameter from the candidate values according to the performance parameters.
[0026] It should be noted that in the embodiments of the present application, by adopting a scale-invariant linear layer, the classification layer in the classification model to be trained can be replaced, so that each layer in the classification model to be trained has scale invariance, that is, scale insensitivity.
[0027] It should be understood that in the embodiments of the present application, by adopting a target training method, two important hyperparameters included in the classification model to be trained, namely the learning rate and weight decay, can be equivalently replaced by a target hyperparameter; since the target training method can make the norms of the weight parameters of the classification model to be trained before and after updating the same; therefore, when updating the weight parameters of the classification model to be trained, only the target hyperparameter needs to be searched to determine the target value of the target hyperparameter, where the target hyperparameter can refer to an equivalent learning rate, and the role of the equivalent learning rate in the classification model to be trained can be regarded as the same or similar to the role of the learning rate, and the equivalent learning rate is used to control the gradient update step size of the classification model to be trained.
[0028] In the embodiments of the present application, by adopting a scale-invariant linear layer in the classification model to be trained, the classification layer in the classification model to be trained can be replaced, so that each layer of the classification model to be trained has scale invariance; further, by adopting a target training method, only adjusting one target hyperparameter can achieve the optimal accuracy that can be achieved by adjusting two hyperparameters (learning rate and weight decay) under normal training; that is, the search process of the two hyperparameters is equivalently replaced by the search process of one target hyperparameter, thereby reducing the search dimension of the hyperparameters and improving the search efficiency of the target hyperparameter; at the same time, the search cost of the hyperparameters is reduced.
[0029] Combined with the second aspect, in some implementation manners of the second aspect, the accuracy of the classification model to be trained corresponding to the target value is greater than the accuracy of the classification model to be trained corresponding to other candidate values among the candidate values.
[0030] In an embodiment of the present application, the target value of the target hyperparameter may refer to the candidate value corresponding to the optimal accuracy of the to-be-trained classification model when assigning each candidate value of the target hyperparameter to the to-be-trained classification model.
[0031] In combination with the second aspect, in some implementation manners of the second aspect, the obtaining of the candidate values of the target hyperparameter includes:
[0032] Performing uniform partitioning according to the initial search range of the target hyperparameter to obtain the candidate values of the target hyperparameter.
[0033] In an embodiment of the present application, the candidate values of the target hyperparameter may be multiple candidate values of the target hyperparameter obtained by performing uniform partitioning according to the initial search range set by the user.
[0034] In a possible implementation manner, in addition to the above-mentioned uniform partitioning of the initial search range to obtain multiple candidate values, other ways may also be used to partition the initial search range.
[0035] In combination with the second aspect, in some implementation manners of the second aspect, it further includes:
[0036] Updating the initial search range of the target hyperparameter according to the current training step, the pre-configured training step, and the change trend of the accuracy of the to-be-trained classification model.
[0037] In a possible implementation manner, update the initial search range of the target hyperparameter according to the current training step, the pre-configured training step, and the monotonicity of the accuracy of the to-be-trained classification model.
[0038] In an embodiment of the present application, in order to quickly search for the target value of the target hyperparameter, that is, the optimal target hyperparameter, during the search for the target hyperparameter, the initial search range of the target hyperparameter may be updated according to the monotonicity of the performance parameters of the to-be-trained classification model, reducing the search range of the target hyperparameter and improving the search efficiency of the target hyperparameter.
[0039] In combination with the second aspect, in some implementation manners of the second aspect, the updating of the initial search range of the target hyperparameter according to the current training step, the pre-configured training step, and the change trend of the accuracy of the to-be-trained classification model includes:
[0040] If the current training step is less than the pre-configured training step, update the upper boundary of the initial search range of the target hyperparameter to the candidate value of the target hyperparameter corresponding to the optimal accuracy of the to-be-trained classification model in the current training step.
[0041] In combination with the second aspect, in some implementations of the second aspect, updating the initial search range of the target hyperparameter according to the current training step, the pre-configured training step, and the change trend of the accuracy of the to-be-trained classification model includes:
[0042] If the current training step is equal to the pre-configured training step, update the upper boundary of the initial search range of the target hyperparameter to a first candidate value, and update the lower boundary of the search range of the target hyperparameter to a second candidate value. The first candidate value and the second candidate value refer to candidate values adjacent to the candidate values of the target hyperparameter corresponding to the optimal accuracy of the to-be-trained classification model.
[0043] In combination with the second aspect, in some implementations of the second aspect, the scale-invariant linear layer obtains the predicted classification result according to the following formula:
[0044]
[0045] where Y i represents the predicted classification result corresponding to the weight parameter updated in the i-th iteration; W i represents the weight parameter updated in the i-th iteration; X represents the feature to be classified; S represents the scale constant.
[0046] In combination with the second aspect, in some implementations of the second aspect, the target training method includes:
[0047] Process the updated weight parameter through the following formula so that the norm lengths of the weight parameters before and after the update of the to-be-trained classification model are the same:
[0048]
[0049] where W i+1 represents the weight parameter updated in the (i + 1)-th iteration; W i represents the weight parameter updated in the i-th iteration; Norm0 represents the initial weight norm length of the to-be-trained classification model.
[0050] In the embodiments of the present application, when each layer of the to-be-trained classification model has scale invariance, iteratively updating the weights of the to-be-trained classification model through the target training method can make the norm lengths of the weight parameters before and after the update the same; that is, the effect of weight decay can be achieved through the target training method, and manual adjustment is no longer required.
[0051] In a third aspect, a training device for a classification model is provided, including:
[0052] An acquisition unit for acquiring target hyperparameters of a classification model to be trained, where the target hyperparameters are used to control the gradient update step size of the classification model to be trained, and the classification model to be trained includes a scale-invariant linear layer, and the scale-invariant linear layer enables the prediction classification result output when the weight parameters of the classification model to be trained are multiplied by any scaling factor to remain unchanged; a processing unit for updating the weight parameters of the classification model to be trained according to the target hyperparameters and a target training method to obtain a trained classification model, and the target training method enables the magnitudes of the weight parameters of the classification model to be trained before and after the update to be the same.
[0053] Combined with the third aspect, in some implementation manners of the third aspect, the weight parameters of the trained classification model are obtained by iteratively updating through the backpropagation algorithm according to the target hyperparameters and the target training method.
[0054] Combined with the third aspect, in some implementation manners of the third aspect, the scale-invariant linear layer obtains the prediction classification result according to the following formula:
[0055]
[0056] where Y i represents the prediction classification result corresponding to the weight parameters updated in the i-th iteration; W i represents the weight parameters updated in the i-th iteration; X represents the feature to be classified; S represents the scale constant.
[0057] Combined with the third aspect, in some implementation manners of the third aspect, the target training method includes:
[0058] Processing the updated weight parameters through the following formula to make the magnitudes of the weight parameters of the classification model to be trained before and after the update the same:
[0059]
[0060] where W i+1 represents the weight parameters updated in the (i + 1)-th iteration; W i represents the weight parameters updated in the i-th iteration; Norm0 represents the initial weight magnitude of the classification model to be trained.
[0061] In a fourth aspect, a hyperparameter search device is provided, including:
[0062] An acquisition unit for acquiring candidate values of a target hyperparameter, where the target hyperparameter is used to control the gradient update step of a classification model to be trained, and the classification model to be trained includes a scale-invariant linear layer, and the scale-invariant linear layer enables the predicted classification result output when the weight parameters of the classification model to be trained are multiplied by any scaling factor to remain unchanged; a processing unit for obtaining a performance parameter of the classification model to be trained according to the candidate value and a target training method, where the target training method enables the norm lengths of the weight parameters of the classification model to be trained before and after update to be the same, and the performance parameter includes the accuracy of the classification model to be trained; determining a target value of the target hyperparameter from the candidate values according to the performance parameter.
[0063] It should be noted that in the embodiments of the present application, by adopting a scale-invariant linear layer, the classification layer in the classification model to be trained can be replaced, so that each layer in the classification model to be trained has scale invariance, that is, scale insensitivity.
[0064] It should be understood that in the embodiments of the present application, by adopting a target training method, two important hyperparameters included in the classification model to be trained, namely the learning rate and the weight decay, can be equivalently replaced with a target hyperparameter; since the target training method can make the norm lengths of the weight parameters of the classification model to be trained before and after update the same; therefore, when updating the weight parameters of the classification model to be trained, only the target hyperparameter needs to be searched to determine the target value of the target hyperparameter, where the target hyperparameter can refer to an equivalent learning rate, and the role of the equivalent learning rate in the classification model to be trained can be regarded as the same or similar to the role of the learning rate, and the equivalent learning rate is used to control the gradient update step of the classification model to be trained.
[0065] In the embodiments of the present application, by adopting a scale-invariant linear layer in the classification model to be trained, the classification layer in the classification model to be trained can be replaced, so that each layer of the classification model to be trained has scale invariance; further, by adopting a target training method, it is possible to achieve the optimal accuracy that can be achieved by adjusting two hyperparameters (learning rate and weight decay) under normal training by only adjusting one target hyperparameter; that is, the search process of two hyperparameters is equivalently replaced with the search process of one target hyperparameter, thereby reducing the search dimension of the hyperparameter and improving the search efficiency of the target hyperparameter; at the same time, the search cost of the hyperparameter is reduced.
[0066] Combined with the fourth aspect, in some implementation manners of the fourth aspect, the accuracy of the classification model to be trained corresponding to the target value is greater than the accuracy of the classification model to be trained corresponding to other candidate values among the candidate values.
[0067] In an embodiment of the present application, the target value of the target hyperparameter may refer to the candidate value corresponding to the optimal accuracy of the classification model to be trained when assigning each candidate value of the target hyperparameter to the classification model to be trained.
[0068] In combination with the fourth aspect, in some implementation manners of the fourth aspect, the processing unit is specifically configured to:
[0069] Uniformly divide according to the initial search range of the target hyperparameter to obtain candidate values of the target hyperparameter.
[0070] In an embodiment of the present application, the candidate values of the target hyperparameter may be multiple candidate values of the target hyperparameter obtained by uniformly dividing according to the initial search range set by the user.
[0071] In combination with the fourth aspect, in some implementation manners of the fourth aspect, the processing unit is further configured to:
[0072] Update the initial search range of the target hyperparameter according to the current training step, the pre-configured training step, and the change trend of the accuracy of the classification model to be trained.
[0073] In an embodiment of the present application, in order to quickly search for the target value of the target hyperparameter, that is, the optimal target hyperparameter, during the search for the target hyperparameter, the initial search range of the target hyperparameter may be updated according to the monotonicity of the performance parameters of the classification model to be trained, the search range of the target hyperparameter may be reduced, and the search efficiency of the target hyperparameter may be improved.
[0074] In combination with the fourth aspect, in some implementation manners of the fourth aspect, the processing unit is specifically configured to:
[0075] If the current training step is less than the pre-configured training step, update the upper boundary of the initial search range of the target hyperparameter to the candidate value of the target hyperparameter corresponding to the optimal accuracy of the classification model to be trained in the current training step.
[0076] In combination with the fourth aspect, in some implementation manners of the fourth aspect, the processing unit is specifically configured to:
[0077] If the current training step is equal to the pre-configured training step, update the upper boundary of the initial search range of the target hyperparameter to a first candidate value, and update the lower boundary of the search range of the target hyperparameter to a second candidate value. The first candidate value and the second candidate value refer to candidate values adjacent to the candidate value of the target hyperparameter corresponding to the optimal accuracy of the classification model to be trained.
[0078] In connection with the fourth aspect, in some implementations of the fourth aspect, the scale-invariant linear layer obtains the predicted classification result according to the following formula:
[0079]
[0080] where Y i represents the predicted classification result corresponding to the weight parameter updated in the i-th iteration; W i represents the weight parameter updated in the i-th iteration; X represents the feature to be classified; and S represents the scale constant.
[0081] In connection with the fourth aspect, in some implementations of the fourth aspect, the target training method includes:
[0082] Processing the updated weight parameter through the following formula to make the norm lengths of the weight parameters of the classification model to be trained before and after the update the same:
[0083]
[0084] where W i+1 represents the weight parameter updated in the (i + 1)-th iteration; W i represents the weight parameter updated in the i-th iteration; and Norm0 represents the initial weight norm of the classification model to be trained.
[0085] In the embodiments of the present application, when each layer of the classification model to be trained has scale invariance, iteratively updating the weights of the classification model to be trained through the target training method can make the norm lengths of the weight parameters before and after the update the same; that is, the effect of weight decay can be achieved through the target training method, and manual adjustment is no longer required.
[0086] Fifth aspect, a training method of a classification model is provided, including a memory for storing a program; a processor for executing the program stored in the memory. When the program stored in the memory is executed by the processor, the processor is configured to: obtain target hyperparameters of the classification model to be trained, the target hyperparameters being used to control the gradient update step of the classification model to be trained, the classification model to be trained including a scale-invariant linear layer, and the scale-invariant linear layer making the predicted classification result remain unchanged when the weight parameter of the classification model to be trained is multiplied by any scaling coefficient; update the weight parameter of the classification model to be trained according to the target hyperparameters and the target training method to obtain the trained classification model, and the target training method makes the norm lengths of the weight parameters of the classification model to be trained before and after the update the same.
[0087] In a possible implementation, the processor included in the above search device is further configured to execute the training method in any one of the implementations in the first aspect.
[0088] It should be understood that the extensions, limitations, explanations, and descriptions of the relevant content in the above first aspect also apply to the same content in the third aspect.
[0089] In a sixth aspect, a hyperparameter search device is provided, including a memory for storing a program; a processor for executing the program stored in the memory. When the program stored in the memory is executed by the processor, the processor is configured to: obtain candidate values of a target hyperparameter, where the target hyperparameter is used to control the gradient update step of a classification model to be trained, and the classification model to be trained includes a scale-invariant linear layer, and the scale-invariant linear layer enables the predicted classification result output when the weight parameters of the classification model to be trained are multiplied by any scaling factor to remain unchanged; obtain performance parameters of the classification model to be trained according to the candidate values and a target training method, where the target training method enables the magnitudes of the weight parameters of the classification model to be trained before and after update to be the same, and the performance parameters include the accuracy of the classification model to be trained; and determine a target value of the target hyperparameter from the candidate values according to the performance parameters.
[0090] In a possible implementation, the processor included in the above search device is further configured to execute the search method in any one of the implementations in the second aspect.
[0091] It should be understood that the extensions, limitations, explanations, and descriptions of the relevant content in the above second aspect also apply to the same content in the third aspect.
[0092] In a seventh aspect, a computer-readable medium is provided, which stores program code for a device to execute, and the program code includes code for executing the training method in the first aspect and any one of the implementations in the first aspect.
[0093] In an eighth aspect, a computer-readable medium is provided, which stores program code for a device to execute, and the program code includes code for executing the search method in the second aspect and any one of the implementations in the second aspect.
[0094] In a ninth aspect, a computer-readable medium is provided, which stores program code for a device to execute, and the program code includes code for executing the training method in the first aspect and any one of the implementations in the first aspect.
[0095] In a tenth aspect, there is provided a computer program product including instructions, which, when the computer program product runs on a computer, cause the computer to execute the search method in the second aspect and any one implementation manner of the second aspect described above.
[0096] In an eleventh aspect, there is provided a chip, which includes a processor and a data interface. The processor reads instructions stored on a memory through the data interface and executes the training method in the first aspect and any one implementation manner of the first aspect described above.
[0097] Optionally, as an implementation manner, the chip may further include a memory. Instructions are stored in the memory, and the processor is configured to execute the instructions stored on the memory. When the instructions are executed, the processor is configured to execute the training method in the first aspect and any one implementation manner of the first aspect described above.
[0098] In a twelfth aspect, there is provided a chip, which includes a processor and a data interface. The processor reads instructions stored on a memory through the data interface and executes the search method in the second aspect and any one implementation manner of the second aspect described above.
[0099] Optionally, as an implementation manner, the chip may further include a memory. Instructions are stored in the memory, and the processor is configured to execute the instructions stored on the memory. When the instructions are executed, the processor is configured to execute the search method in the second aspect and any one implementation manner of the second aspect described above. Description of the Drawings
[0100] Figure 1 is a schematic diagram of the system architecture of the hyperparameter search method provided by an embodiment of the present application;
[0101] Figure 2 is a schematic diagram of a system architecture provided by an embodiment of the present application;
[0102] Figure 3 is a schematic diagram of the hardware structure of a chip provided by an embodiment of the present application;
[0103] Figure 4 is a schematic diagram of a system architecture provided by an embodiment of the present application;
[0104] Figure 5 is a schematic flowchart of the training method of the classification model provided by an embodiment of the present application;
[0105] Figure 6 is a schematic diagram of iteratively updating the weights of the classification model to be trained based on the scale-invariant linear layer and the fixed norm method provided by an embodiment of the present application;
[0106] Figure 7 It is a schematic flowchart of a method for searching hyperparameters provided by an embodiment of the present application;
[0107] Figure 8 It is a schematic diagram of a method for searching hyperparameters provided by an embodiment of the present application;
[0108] Figure 9 It is a schematic block diagram of a training device for a classification model provided by an embodiment of the present application;
[0109] Figure 10 It is a schematic block diagram of a device for searching hyperparameters provided by an embodiment of the present application;
[0110] Figure 11 It is a schematic diagram of the hardware structure of a training device for a classification model provided by an embodiment of the present application;
[0111] Figure 12 It is a schematic diagram of the hardware structure of a device for searching hyperparameters provided by an embodiment of the present application. Detailed implementation manners
[0112] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application; obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments.
[0113] First, a brief description of the concepts involved in the embodiments of the present application will be given.
[0114] 1. Hyperparameters
[0115] In machine learning, hyperparameters refer to the parameters whose values are set before the start of the learning process, rather than the parameter data obtained through training. Generally, hyperparameters can be optimized to select a set of optimal hyperparameters for the learning machine to improve the performance and effect of learning.
[0116] 2. Learning rate
[0117] The learning rate is an important hyperparameter when training a neural network; the learning rate can be used to control the step size of gradient update during the training process of the neural network.
[0118] 3. Weight decay
[0119] Weight decay is an important hyperparameter when training a neural network; weight decay can be used to control the intensity of the constraint on the weight size.
[0120] Next, in combination with Figure 1 the system architecture of the method for searching hyperparameters in the embodiments of the present application will be described.
[0121] As Figure 1As shown, the system architecture 100 may include an equivalent learning rate search module 110 and a scale-invariant training module 111. Among them, the equivalent learning rate search module 110 is used to optimize the hyperparameters of the neural network model, so as to obtain the optimal accuracy corresponding to the neural network model and the optimal model weights corresponding to the optimal accuracy. The scale-invariant training module 111 is used to iteratively update the weights of the neural network model according to different input equivalent learning rate hyperparameters to obtain the parameters of the neural network model.
[0122] It should be noted that the equivalent learning rate hyperparameter can be used to control the step size of gradient update during the neural network training process. By only adjusting the equivalent learning rate hyperparameter in the neural network, the neural network model can achieve the optimal accuracy of the model that can be achieved by adjusting the two hyperparameters of the learning rate and weight decay under normal training.
[0123] Exemplarily, as Figure 1 shown, the equivalent learning rate search module 110 and the scale-invariant training module 111 can be deployed in the cloud or on the server.
[0124] Figure 2 Fig. 200 shows a system architecture provided by an embodiment of the present application.
[0125] In Figure 2 , the data acquisition device 260 is used to acquire training data. For the classification model to be trained in the embodiment of the present application, the classification model to be trained can be trained with the training data acquired by the data acquisition device 260.
[0126] Exemplarily, in the embodiment of the present application, for the classification model to be trained, the training data may include the data to be classified and the sample labels corresponding to the data to be classified.
[0127] After the training data is acquired, the data acquisition device 260 stores the training data in the database 230, and the training device 220 trains to obtain the target model / rule 201 based on the training data maintained in the database 230.
[0128] The following describes how the training device 220 obtains the target model / rule 201 based on the training data.
[0129] For example, the training device 220 processes the training data of the classification model to be trained, compares the predicted classification result corresponding to the data to be classified output by the classification model to be trained with the true value corresponding to the data to be classified, until the difference between the predicted classification result output by the training device 220 and the true value is less than a certain threshold, thereby completing the training of the classification model.
[0130] It should be noted that in actual applications, the training data maintained in the database 230 may not necessarily all come from the collection of the data collection device 260, and it is also possible to be received from other devices.
[0131] In addition, it should be noted that the training device 220 does not necessarily train the target model / rule 201 completely based on the training data maintained in the database 230. It is also possible to obtain training data from the cloud or other places for model training. The above description should not be regarded as a limitation to the embodiments of the present application.
[0132] The target model / rule 201 trained according to the training device 220 can be applied to different systems or devices, such as being applied to Figure 2 the execution device 210 shown. The execution device 210 can be a terminal, such as a mobile phone terminal, a tablet computer, a laptop computer, augmented reality (AR) / virtual reality (VR), a vehicle-mounted terminal, etc., or it can also be a server, or a cloud server, etc. In Figure 2 , the execution device 210 is configured with an input / output (I / O) interface 212 for data interaction with external devices. The user can input data to the I / O interface 212 through the client device 240. The input data in the embodiments of the present application may include: training data input by the client device.
[0133] The preprocessing modules 213 and 214 are used to preprocess the input data received through the I / O interface 212. In the embodiments of the present application, it is also possible not to have the preprocessing modules 213 and 214 (or only one of the preprocessing modules), and directly use the computing module 211 to process the input data.
[0134] During the preprocessing of the input data by the execution device 210, or during the relevant processing such as the computing module 211 of the execution device 210 performing calculations, the execution device 210 can call data, code, etc. in the data storage system 250 for corresponding processing, or can also store the data, instructions, etc. obtained from the corresponding processing into the data storage system 250.
[0135] Finally, the I / O interface 212 returns the processing result, for example, the target value of the target hyperparameter, to the client device 240, so as to be provided to the user.
[0136] It should be noted that the training device 220 can generate corresponding target models / rules 201 based on different training data for different targets or tasks. The corresponding target models / rules 201 can be used to achieve the above-mentioned targets or complete the above-mentioned tasks, so as to provide the required results for users.
[0137] In Figure 2 the situation shown, in one case, the user can manually give input data, and this manual giving can be operated through the interface provided by the I / O interface 212.
[0138] In another case, the client device 240 can automatically send input data to the I / O interface 212. If the client device 240 is required to automatically send input data and needs to obtain the user's authorization, the user can set the corresponding permissions in the client device 240. The user can view the results output by the execution device 210 in the client device 240, and the specific presentation forms can be display, sound, action and other specific ways. The client device 240 can also be used as a data acquisition end to collect the input data input to the I / O interface 212 and the output results output from the I / O interface 212 as new sample data and store them in the database 230. Of course, it can also be collected without passing through the client device 240, but the input data input to the I / O interface 212 and the output results output from the I / O interface 212 as shown in the figure are directly stored in the database 130 as new sample data.
[0139] It should be noted that Figure 2 is only a schematic diagram of a system architecture provided by an embodiment of the present application. The positional relationship between the devices, components, modules, etc. shown in the figure does not constitute any limitation. For example, in Figure 2 the data storage system 250 is an external memory relative to the execution device 210; in other cases, the data storage system 250 can also be placed in the execution device 210.
[0140] Figure 3 is a schematic diagram of the hardware structure of a chip provided by an embodiment of the present application.
[0141] Figure 3 The chip shown may include a neural network processor 300 (neural-network processing unit, NPU); the chip can be set in the execution device 210 as shown in Figure 2 to complete the calculation work of the calculation module 211. The chip can also be set in the training device 220 as shown in Figure 2 to complete the training work of the training device 220 and output the target model / rule 201.
[0142] The NPU 300 is mounted on the main central processing unit (CPU) as a coprocessor and tasks are assigned by the main CPU. The core part of the NPU 300 is the arithmetic circuit 303, and the controller 304 controls the arithmetic circuit 303 to extract data from the memory (weight memory or input memory) and perform operations.
[0143] In some implementations, the arithmetic circuit 303 includes multiple processing engines (PEs) inside. In some implementations, the arithmetic circuit 303 is a two-dimensional systolic array; the arithmetic circuit 303 can also be a one-dimensional systolic array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 303 is a general matrix processor.
[0144] For example, assume there is an input matrix A, a weight matrix B, and an output matrix C; the arithmetic circuit 303 fetches the corresponding data of matrix B from the weight memory 302 and caches it on each PE in the arithmetic circuit 303; the arithmetic circuit 303 fetches the data of matrix A from the input memory 301 and performs matrix operations with matrix B, and the partial results or final results of the obtained matrix are stored in the accumulator 308.
[0145] The vector calculation unit 307 can further process the output of the arithmetic circuit 303, such as vector multiplication, vector addition, exponential operation, logarithmic operation, magnitude comparison, etc. For example, the vector calculation unit 307 can be used for network calculations in non-convolution / non-FC layers of a neural network, such as pooling, batch normalization, local response normalization, etc.
[0146] In some implementations, the vector calculation unit 307 can store the processed output vector into the unified memory 306. For example, the vector calculation unit 307 can apply a non-linear function to the output of the arithmetic circuit 303, such as a vector of accumulated values, to generate activation values. In some implementations, the vector calculation unit 307 generates normalized values, combined values, or both.
[0147] In some implementations, the processed output vector can be used as the activation input to the arithmetic circuit 303, for example, for use in subsequent layers in a neural network.
[0148] The unified memory 306 is used to store input data and output data. The weight data directly stores the input data in the external memory into the input memory 301 and / or the unified memory 306, stores the weight data in the external memory into the weight memory 302, and stores the data in the unified memory 306 into the external memory through the memory access controller 305 (direct memory access controller, DMAC).
[0149] The bus interface unit 310 (bus interface unit, BIU) is used to realize the interaction between the main CPU, the DMAC, and the instruction fetch buffer 309 through the bus.
[0150] The instruction fetch buffer 309 connected to the controller 304 is used to store the instructions used by the controller 304; the controller 304 is used to call the instructions cached in the instruction fetch buffer 309 to control the working process of the arithmetic accelerator.
[0151] Generally, the unified memory 306, the input memory 301, the weight memory 302, and the instruction fetch buffer 309 are all on-chip memories, and the external memory is the memory outside the NPU. The external memory can be a double data rate synchronous dynamic random access memory (DDR SDRAM), a high bandwidth memory (HBM), or other readable and writable memories.
[0152] Exemplarily, the related operations for the iterative update of the weight parameters of the classification model to be trained in the embodiments of the present application can be executed by the arithmetic circuit 303 or the vector calculation unit 307.
[0153] As introduced above Figure 2 The execution device 210 in can execute each step of the hyperparameter search method in the embodiments of the present application, and the chip shown in FIG. 3 can also be used to execute each step of the hyperparameter search method in the embodiments of the present application.
[0154] Figure 4 FIG. shows a system architecture provided by an embodiment of the present application. The system architecture 400 may include a local device 420, a local device 430, an execution device 410, and a data storage system 450. Among them, the local device 420 and the local device 430 are connected to the execution device 410 through a communication network.
[0155] Exemplarily, the execution device 410 can be implemented by one or more servers.
[0156] Optionally, the execution device 410 can be used in cooperation with other computing devices. For example: devices such as data memories, routers, load balancers, etc. The execution device 410 can be arranged on one physical site, or distributed on multiple physical sites. The execution device 410 can use the data in the data storage system 450, or call the program code in the data storage system 450 to implement the processing method of the combination optimization task in the embodiments of the present application.
[0157] It should be noted that the above execution device 410 can also be referred to as a cloud device. In this case, the execution device 410 can be deployed in the cloud.
[0158] In one example, the execution device 410 can perform the following process: obtaining target hyperparameters of the classification model to be trained, where the target hyperparameters are used to control the gradient update step of the classification model to be trained, and the classification model to be trained includes a scale-invariant linear layer, and the scale-invariant linear layer makes the prediction classification result remain unchanged when the weight parameters of the classification model to be trained are multiplied by any scaling factor; updating the weight parameters of the classification model to be trained according to the target hyperparameters and the target training method to obtain the trained classification model, and the target training method makes the magnitudes of the weight parameters of the classification model to be trained before and after the update the same.
[0159] In one possible implementation manner, the training method of the classification model in the embodiments of the present application can be an offline method executed in the cloud. For example, the training method in the embodiments of the present application can be executed by the above execution device 410.
[0160] In one possible implementation manner, the training method of the classification model in the embodiments of the present application can be executed by the local device 420 or the local device 430.
[0161] In one example, the execution device 410 can perform the following process: obtaining candidate values of target hyperparameters, where the target hyperparameters are used to control the gradient update step of the classification model to be trained, and the classification model to be trained includes a scale-invariant linear layer, and the scale-invariant linear layer makes the prediction classification result remain unchanged when the weight parameters of the classification model to be trained are multiplied by any scaling factor; obtaining performance parameters of the classification model to be trained according to the candidate values and the target training method, where the target training method makes the magnitudes of the weight parameters of the classification model to be trained before and after the update the same, and the performance parameters include the accuracy of the classification model to be trained; determining the target value of the target hyperparameters from the candidate values according to the performance parameters.
[0162] In a possible implementation, the hyperparameter search method of the embodiments of the present application may be an offline method executed in the cloud. For example, the search method of the embodiments of the present application may be executed by the execution device 410 above.
[0163] In a possible implementation, the hyperparameter search method of the embodiments of the present application may be executed by the local device 420 or the local device 430.
[0164] For example, users can operate their respective user devices (e.g., the local device 420 and the local device 430) to interact with the execution device 410. Each local device may represent any computing device, such as a personal computer, a computer workstation, a smart phone, a tablet computer, a smart camera, a smart car, or other types of cellular phones, media consumption devices, wearable devices, set-top boxes, game consoles, etc. Each user's local device can interact with the execution device 410 through a communication network of any communication mechanism / communication standard. The communication network can be a wide area network, a local area network, a point-to-point connection, etc., or any combination thereof.
[0165] The following will Figures 5 to 8 elaborate on the technical solutions of the embodiments of the present application in detail.
[0166] Figure 5 is a schematic flowchart of the training method of the classification model provided by the embodiments of the present application. In some examples, the training method 500 may be executed by Figure 2 the execution device 210 in Figure 3 the chip shown in Figure 4 the execution device 410 or a local device in Figure 5 The method 500 in
[0167] S510. Obtain the target hyperparameters of the classification model to be trained.
[0168] Among them, the target hyperparameters are used to control the gradient update step size of the classification model to be trained. The classification model to be trained includes a scale-invariant linear layer, and the scale-invariant linear layer enables the predicted classification results to remain unchanged when the weight parameters of the classification model to be trained are multiplied by any scaling factor.
[0169] It should be noted that when the parameters of a neural network have scale invariance, the effective step size of its weight parameters is inversely proportional to the square of the norm of the weight parameters; where scale invariance means that when the weight w is multiplied by any scaling factor, the output y of this layer remains unchanged. The main role of the weight decay hyperparameter is to constrain the norm of the parameters, thereby preventing the step size from decreasing sharply as the norm of the weight parameters increases. However, since weight decay is an implicit constraint, the norm of the weight parameters will still change during the training process, resulting in the need for manual adjustment of the optimal weight decay hyperparameter for different models.
[0170] In the embodiments of the present application, the target training method can be used to explicitly fix the norm of the parameters, that is, to make the norm of the weight parameters before and after the update of the classification model to be trained the same, so that the influence of the norm of the weight parameters on the step size remains constant, achieving the effect of weight decay and eliminating the need for manual adjustment. Since the premise of using the target training method is that the neural network needs to have scale invariance, generally for most neural networks, most layers have scale invariance, but most classification layers in the classification model are scale-sensitive; therefore, in the embodiments of the present application, a scale-invariant linear layer is used to replace the original classification layer, so that each layer of the classification model to be trained has scale invariance. For example, the scale-invariant linear layer can be seen in the following Figure 6 as shown.
[0171] Exemplarily, the above classification model can include any neural network model with classification functions; for example, the classification model can include but is not limited to: detection models, segmentation models, and recognition models, etc.
[0172] Exemplarily, the target hyperparameter can refer to the equivalent learning rate. The role of the equivalent learning rate in the classification model to be trained can be regarded as the same or similar to the role of the learning rate, and the equivalent learning rate is used to control the gradient update step size of the classification model to be trained.
[0173] S520. Update the weight parameters of the classification model to be trained according to the target hyperparameter and the target training method to obtain the trained classification model.
[0174] Among them, the target training method makes the norm of the weight parameters before and after the update of the classification model to be trained the same.
[0175] It should be understood that when training a classification model, it is usually necessary to configure target hyperparameters to iteratively update the weight parameters in the classification model, so that the classification model converges; that is, the difference between the predicted classification result output by the classification model for the sample classification features and the true value corresponding to the sample features is less than or equal to a preset range; among them, the target hyperparameter can refer to the equivalent learning rate, and the role of the equivalent learning rate in the classification model to be trained can be regarded as the same or similar to the role of the learning rate, and the equivalent learning rate is used to control the gradient update step of the classification model to be trained.
[0176] Optionally, in a possible implementation manner, the weight parameters of the trained classification model are obtained by iteratively updating multiple times through the backpropagation algorithm according to the target hyperparameters and the target training method.
[0177] Optionally, in a possible implementation manner, the scaling invariance linear layer obtains the predicted classification result according to the following formula:
[0178]
[0179] where Y i represents the predicted classification result corresponding to the weight parameter updated in the i-th iteration; W i represents the weight parameter updated in the i-th iteration; X represents the feature to be classified; S represents the scale constant.
[0180] It should be noted that the above formula is an example for scaling invariance processing. The purpose of the scaling invariance linear layer is to make the weight w multiplied by any scaling factor, and the output y of this layer can remain unchanged; in other words, other specific processing methods can also be used for scaling invariance processing, and the present application does not make any limitations on this.
[0181] Optionally, in a possible implementation manner, the target training method includes:
[0182] Processing the updated weight parameters through the following formula to make the norm lengths of the weight parameters before and after the update of the classification model to be trained the same:
[0183]
[0184] where W i+1 represents the weight parameter updated in the (i + 1)-th iteration; W i represents the weight parameter updated in the i-th iteration; Norm0 represents the initial weight norm length of the classification model to be trained.
[0185] In an embodiment of the present application, by replacing the classification layer of the classification model to be trained with a scale-invariant linear layer and adopting a target training method, that is, fixing the norms of the weight parameters before and after updating the classification model to be trained, two important hyperparameters, namely the learning rate and the weight decay, can be equivalently replaced with target hyperparameters when training the classification model to be trained; therefore, when training the classification model to be trained, only the target hyperparameters of the model to be trained need to be configured, thereby reducing the computing resources consumed for training the classification model to be trained.
[0186] Figure 7 is a schematic flowchart of a method for searching hyperparameters provided by an embodiment of the present application. In some examples, the search method 600 may be executed by Figure 2 the execution device 210 in Figure 3 the chip shown in Figure 4 and the execution device 410 or a local device in Figure 7 The method 600 in
[0187] S610. Obtain candidate values of the target hyperparameters.
[0188] Among them, the target hyperparameters are used to control the gradient update step of the classification model to be trained, and the classification model to be trained includes a scale-invariant linear layer, and the scale-invariant linear layer enables the predicted classification result output when the weight parameters of the classification model to be trained are multiplied by any scaling factor to remain unchanged.
[0189] It should be noted that in the embodiment of the present application, by adopting a scale-invariant linear layer, the classification layer in the classification model to be trained can be replaced, so that each layer in the classification training model has scale invariance, that is, scale insensitivity. For example, the scale-invariant linear layer can be referred to Figure 6 shown in
[0190] Optionally, in a possible implementation manner, the scale-invariant linear layer obtains the predicted classification result according to the following formula:
[0191]
[0192] where Y i represents the predicted classification result corresponding to the weight parameter updated at the i-th iteration; W i represents the weight parameter updated at the i-th iteration; X represents the feature to be classified; and S represents the scale constant.
[0193] S620. Obtain the performance parameters of the classification model to be trained according to the candidate values and the target training method.
[0194] Among them, the target training method makes the norm lengths of the weight parameters of the classification model to be trained the same before and after updating, and the performance parameter includes the accuracy of the classification model to be trained.
[0195] It should be understood that the performance parameter of the classification model to be trained can be used to evaluate the performance parameter of the classification model to be trained after configuring a certain target hyperparameter for the classification model to be trained; for example, the performance parameter can include the model accuracy of the classification model to be trained, that is, it can refer to the accuracy of the predicted classification result of the classification model to be trained for classifying the features to be classified.
[0196] It should be noted that in the embodiments of the present application, by adopting the target training method, two important hyperparameters included in the classification model to be trained, namely the learning rate and the weight decay, can be equivalently replaced with the target hyperparameter; since the target training method can make the norm lengths of the weight parameters of the classification model to be trained the same before and after updating; therefore, when training the classification model to be trained, only the target hyperparameter needs to be searched to determine the target value of the target hyperparameter, where the target hyperparameter can refer to the equivalent learning rate, and the role of the equivalent learning rate in the classification model to be trained can be regarded as the same or similar to the role of the learning rate, and the equivalent learning rate is used to control the gradient update step of the classification model to be trained.
[0197] Optionally, in a possible implementation manner, the target training method includes:
[0198] Process the updated weight parameter through the following formula so that the norm lengths of the weight parameters of the classification model to be trained are the same before and after updating:
[0199]
[0200] Among them, W i+1 represents the weight parameter updated at the (i + 1)-th iteration; W i represents the weight parameter updated at the i-th iteration; Norm0 represents the initial weight norm of the classification model to be trained.
[0201] In the embodiments of the present application, when each layer of the classification model to be trained has scale invariance, iteratively updating the weights of the classification model to be trained by the target training method can make the norm lengths of the weight parameters the same before and after updating; that is, the effect of weight decay can be achieved through the target training method, and manual adjustment is no longer required.
[0202] S630. Determine the target value of the target hyperparameter from the candidate values according to the performance parameter of the classification model to be trained.
[0203] Among them, the target value of the target hyperparameter can refer to the value of the target hyperparameter when the classification model to be trained can meet the preset conditions.
[0204] In the embodiments of the present application, by adopting a scale-invariant linear layer in the classification model to be trained, the classification layer in the classification model to be trained can be replaced, so that each layer of the classification model to be trained has scale invariance; further, by adopting a target training method, it is possible to achieve the optimal accuracy that can be achieved by adjusting two hyperparameters (learning rate and weight decay) under normal training by only adjusting one target hyperparameter; that is, the search process of two hyperparameters is equivalently replaced by the search process of one target hyperparameter, thereby reducing the search dimension of the hyperparameter and accelerating the search speed.
[0205] Optionally, in a possible implementation manner, the accuracy of the classification model to be trained corresponding to the target value is greater than the accuracy of the classification model to be trained corresponding to other candidate values among the candidate values.
[0206] In the embodiments of the present application, the target value of the target hyperparameter can refer to the candidate value corresponding to the optimal accuracy of the classification model to be trained when assigning each candidate value of the target hyperparameter to the classification model to be trained.
[0207] Optionally, in a possible implementation manner, the obtaining of the candidate values of the target hyperparameter includes:
[0208] According to the initial search range of the target hyperparameter, it is evenly divided to obtain the candidate values of the target hyperparameter.
[0209] In the embodiments of the present application, the candidate values of the target hyperparameter can be obtained by evenly dividing the initial search range set by the user to obtain multiple candidate values of the target hyperparameter.
[0210] In one example, in addition to the above-mentioned evenly dividing the initial search range to obtain multiple candidate values, other methods can also be used to divide the initial search range.
[0211] Furthermore, in the embodiments of the present application, in order to quickly search for the target value of the target hyperparameter, that is, the optimal target hyperparameter, during the search for the target hyperparameter, the initial search range of the target hyperparameter can be updated according to the monotonicity of the performance parameters of the classification model to be trained, reducing the search range of the target hyperparameter and improving the search efficiency of the target hyperparameter. The specific process can be seen in the following Figure 8 as shown.
[0212] Optionally, in a possible implementation manner, it further includes: updating the initial search range of the target hyperparameter according to the current training step number, the pre-configured training step number, and the change trend of the accuracy of the classification model to be trained.
[0213] Exemplarily, the initial search range of the target hyperparameter can be updated according to the current training step number, the pre-configured training step number, and the monotonicity of the accuracy of the classification model to be trained.
[0214] Optionally, in a possible implementation manner, the updating the initial search range of the target hyperparameter according to the current training step number, the pre-configured training step number, and the change trend of the accuracy of the classification model to be trained includes:
[0215] If the current training step number is less than the pre-configured training step number, update the upper boundary of the initial search range of the target hyperparameter to the candidate value of the target hyperparameter corresponding to the optimal accuracy of the classification model to be trained in the current training step number.
[0216] Optionally, in a possible implementation manner, the updating the initial search range of the target hyperparameter according to the current training step number, the pre-configured training step number, and the change trend of the accuracy of the classification model to be trained includes:
[0217] If the current training step number is equal to the pre-configured training step number, update the upper boundary of the initial search range of the target hyperparameter to the first candidate value, and update the lower boundary of the search range of the target hyperparameter to the second candidate value. The first candidate value and the second candidate value refer to the candidate values adjacent to the candidate value of the target hyperparameter corresponding to the optimal accuracy of the classification model to be trained.
[0218] Schematically, Figure 6 FIG. is a schematic diagram of iteratively updating the weights of the classification model to be trained based on the scale-invariant linear layer provided by the embodiments of the present application.
[0219] It should be noted that when the parameters of the neural network have scale invariance, the equivalent step size of its weight parameters is inversely proportional to the square of the parameter norm; wherein, scale invariance means that when the weight w is multiplied by any scaling coefficient, the output y of this layer can remain unchanged. The main function of the weight decay hyperparameter is to constrain the norm of the parameters, thereby avoiding the sharp decrease of the equivalent step size as the parameter norm increases. However, since weight decay is an implicit constraint, the parameter norm will still change during the training process, resulting in the need for manual adjustment of the optimal weight decay for different models.
[0220] Figure 6 The weight training method of the classification model to be trained shown can make the influence of the parameter norm on the equivalent step size remain constant by the target training method, that is, by fixing the parameter norm explicitly, achieving the effect of weight decay and no longer requiring manual adjustment. Since Figure 6The premise of the method for training the weights of the neural network shown is that the neural network needs to have scale invariance. Generally, for most neural networks, most layers are scale invariant, but the last classification layer is generally scale sensitive. Therefore, as Figure 6 shown, in the embodiments of the present application, the original classification layer is replaced by a scale-invariant linear layer, so that each layer of the neural network has scale invariance.
[0221] As Figure 6 shown, when training the parameters of the classification model to be trained, the initial weight parameters of the classification model to be trained can be obtained, such as the first weight parameter W1 and the feature X obtained from the training data; the first weight parameter W1 and the feature X are subjected to scale-invariant processing through the scale-invariant linear layer to obtain the prediction result Y1; for example, for the classification model, the predicted output result Y1 can refer to the predicted classification label corresponding to the feature X.
[0222] Schematically, the scale-invariant linear layer can process the first weight parameter and the feature by using the following formula:
[0223]
[0224] where Y i represents the predicted classification result corresponding to the weight parameter updated in the i-th iteration; W i represents the weight parameter updated in the i-th iteration; X represents the feature to be classified; S represents the scale constant.
[0225] It should be understood that the above formula is an example for the scale-invariant processing. The purpose of the scale-invariant linear layer is to make the weight w multiplied by any scaling factor, and the output y of this layer can remain unchanged; in other words, the scale-invariant processing can also adopt other specific processing methods, and the present application does not make any limitation on this.
[0226] Further, the initial weight norm of the classification model to be trained can be obtained according to the initial weight parameters of the classification model to be trained, denoted as Norm0.
[0227] Then, the first weight parameter is updated by using forward propagation and backward propagation according to the input equivalent learning rate parameter and the prediction result Y1 output by the scale-invariant linear layer.
[0228] Exemplarily, the first weight parameter can be updated according to the difference between the prediction result Y1 and the true value corresponding to the feature X, so that the accuracy of the output prediction result of the classification model to be trained meets the preset range.
[0229] Further, in the embodiments of the present application, when updating the first weight parameter, the modulus length of the updated second weight parameter can be fixed. The modulus length fixing process refers to making the weight parameter the same before and after the update to achieve the effect of the weight decay hyperparameter.
[0230] Schematically, the following formula can be used to fix the modulus length of the weight parameter updated in the (i + 1)-th iteration update:
[0231]
[0232] where, W i+1 represents the weight parameter updated in the (i + 1)-th iteration; W i represents the weight parameter updated in the i-th iteration; Norm0 represents the initial weight modulus length of the to-be-trained classification model.
[0233] By iteratively updating the weight parameter and repeating the execution of the training steps E times, the performance parameter corresponding to the to-be-trained classification model and the weight of the to-be-trained classification model are returned. Among them, the performance parameter can refer to the accuracy when the to-be-trained classification model performs classification processing.
[0234] In the embodiments of the present application, by adopting a scale-invariant linear layer in the to-be-trained classification model, the classification layer in the to-be-trained classification model can be replaced, so that each layer of the to-be-trained classification model has scale invariance; further, by adopting a target training method, it is possible to achieve the optimal accuracy that can be achieved by adjusting two hyperparameters (learning rate and weight decay) under normal training by only adjusting one target hyperparameter; that is, the search process of the two hyperparameters is equivalently replaced by the search process of one target hyperparameter, thereby reducing the search dimension of the hyperparameter and improving the search efficiency of the target hyperparameter; at the same time, reducing the search cost of the hyperparameter.
[0235] Figure 8 is a schematic diagram of the hyperparameter search method provided by the embodiments of the present application. It should be understood that Figure 8 the process of any one of the trainings from training 1 to training K can refer to the Figure 6 schematic diagram shown.
[0236] Exemplarily, assume that the fast equivalent learning rate search is: search(LR min , LR max , K, T{E t}); where, the range of the input equivalent learning rate is [LR min , LR max ; the number of divisions is K; the number of search rounds is T; the number of training steps per round of search is E t ; and E T is the maximum number of training rounds. The output is the global optimal accuracy acc best ; the global optimal model weight W best .
[0237] The search method includes the following steps:
[0238] Step 1: Determine that the search process is looped for T rounds; for t in {1, 2, …, T - 1}.
[0239] Step 2: Divide the range of the equivalent learning rate to obtain K candidate values of the equivalent learning rate.
[0240] For example, the range of the equivalent learning rate can be evenly divided to obtain K candidate values of the equivalent learning rate; that is
[0241] Step 3: Perform the training process as shown below on the K candidate values LR i of the equivalent learning rate respectively, the number of training steps is E Figure 6 , and obtain the corresponding accuracy Acc of the classification model to be trained and the weight W of the classification model to be trained. t
[0242] Step 4: Determine the identifier corresponding to the highest accuracy of the classification model to be trained; that is, idx = argmax(Acc).
[0243] Furthermore, in the embodiments of the present application, the search range can be updated according to the determined optimal accuracy.
[0244] Step 5: If the current number of training steps is less than the maximum number of training steps E t , then update the upper bound of the range of the equivalent learning rate to the candidate value of the equivalent learning rate corresponding to the highest accuracy; that is, LR max = LR idx .
[0245] Step 6: If the current number of training steps is equal to the maximum number of training steps E t , then update the upper and lower bounds of the range of the equivalent learning rate to LR idx+1 and LR idx-1 respectively; that is, LR max = LR idx+1 ; LR min = LR idx-1 .
[0246] Step 7: If the current highest accuracy is greater than the global optimal accuracy, then update the global optimal accuracy and the global optimal model weight; that is, acc best = Acc idx ; W best = W idx .
[0247] In an embodiment of the present application, by utilizing the unimodal trend of the accuracy of the classification model to be trained with respect to the equivalent learning rate, the range of the equivalent learning rate can be evenly divided, so as to quickly determine the optimal range of the equivalent learning rate at the current training round; in other words, by the monotonicity of the optimal equivalent learning rate at different training rounds, the range of the optimal equivalent learning rate can be quickly reduced in stages, so as to quickly search for the optimal equivalent learning rate.
[0248] In one example, the hyperparameter search method provided in the embodiment of the present application is applied to hyperparameter optimization in the ImageNet image classification task; taking the deep neural network as ResNet50 as an example for illustration.
[0249] Step 1: First, set the initial parameters of the target hyperparameter. For example, the range of the learning rate can be [0.1, 3.2], the number of divisions K = 5, the number of search rounds T = 2, and the number of search steps E per round = {25000, 125000}. Next, start the search for the target hyperparameter.
[0250] Step 2: The first round of search.
[0251] Exemplarily, 5 points are evenly taken within the range of the learning rate [0.1, 3.2], and the Figure 6 shown training method is used to train the weight parameters of the deep neural network. The number of training rounds can be 25000, and the test accuracy corresponding to each learning rate candidate value is obtained.
[0252] For example, the optimal accuracy is 74.12, and the corresponding learning rate is 1.55; then the range of the learning rate can be updated to [0.1, 1.55], and the optimal accuracy is updated to 74.12.
[0253] Step 3: The second round of search.
[0254] Exemplarily, 5 points are evenly taken within the range of the learning rate [0.1, 1.55], and the Figure 6 shown training method is used to train the weight parameters of the deep neural network. The number of training rounds can be 125000, and the test accuracy corresponding to each learning rate candidate value is obtained.
[0255] For example, the optimal accuracy is 77.64, and the corresponding learning rate is 1.0875, and the optimal accuracy is updated to 77.64.
[0256] Step 4: The search is completed.
[0257] For example, the optimal learning rate is 1.0875, and the optimal accuracy of the corresponding deep neural network is 77.64, and the model weights of the corresponding deep neural network are returned.
[0258] Table 1
[0259] Model Baseline Bayesian Bayesian Cost Our Method Our Method Cost Model 1 76.3 77.664 35 77.64 6 Model 2 72.0 72.346 35 72.634 6 Model 3 74.7 75.934 30 76.048 6
[0260] Table 1 shows the test results of model accuracy and resource overhead corresponding to different models without hyperparameter optimization, with Bayesian hyperparameter optimization, and with the hyperparameter search method proposed in this application.
[0261] Among them, Model 1 represents a residual model (e.g., ResNet50); Model 2 represents the MobileNetV2 model; Model 3 represents MobileNetV2×1.4. Baseline represents the accuracy without hyperparameter optimization (unit: %), Bayesian represents the accuracy after hyperparameter optimization using the Bayesian optimization method (unit: %), Bayesian Cost represents the resource overhead of using the Bayesian optimization method (unit: multiple of the single training overhead), Our Method represents the accuracy after hyperparameter optimization using the method of this application (unit: %), and Our Method Cost represents the resource overhead of using the hyperparameter search method proposed in this application for hyperparameter optimization (unit: multiple of the single training overhead).
[0262] It can be seen from the test results shown in Table 1 that in the ImageNet image classification task, for different deep neural network models, the models obtained by the hyperparameter search method proposed in this application are significantly better than the models without hyperparameter optimization, reaching or exceeding the model accuracy obtained by the Bayesian hyperparameter optimization method; and the resource overhead is much smaller than that of the model obtained by the Bayesian hyperparameter optimization method.
[0263] It should be understood that the above examples are for helping those skilled in the art to understand the embodiments of this application, rather than limiting the embodiments of this application to the specific numerical values or specific scenarios illustrated. Those skilled in the art can obviously make various equivalent modifications or changes according to the above examples, and such modifications or changes also fall within the scope of the embodiments of this application.
[0264] Above, in combination with Figures 1 to 8 , the training method of the classification model and the hyperparameter search method provided by the embodiments of this application are described in detail; below, in combination with Figures 9 to 12 , the device embodiments of this application will be described in detail. It should be understood that the training device in the embodiments of this application can execute the training method of the foregoing embodiments of this application; the search device in the embodiments of this application can execute various search methods of the foregoing embodiments of this application, that is, the specific working processes of the following various products can refer to the corresponding processes in the foregoing method embodiments.
[0265] Figure 9It is a schematic block diagram of a training device for a classification model provided by the present application.
[0266] It should be understood that the training device 700 can execute Figure 5 or Figure 6 the training method shown. The training device 700 includes: an acquisition unit 710 and a processing unit 720.
[0267] Among them, the acquisition unit 710 is used to acquire the target hyperparameters of the classification model to be trained, and the target hyperparameters are used to control the gradient update step size of the classification model to be trained. The classification model to be trained includes a scale-invariant linear layer, and the scale-invariant linear layer makes the prediction classification result remain unchanged when the weight parameters of the classification model to be trained are multiplied by any scaling factor; the processing unit 720 is used to update the weight parameters of the classification model to be trained according to the target hyperparameters and the target training method to obtain the trained classification model, and the target training method makes the norm lengths of the weight parameters of the classification model to be trained before and after the update the same.
[0268] Optionally, as an embodiment, the weight parameters of the trained classification model are obtained by iteratively updating through the backpropagation algorithm multiple times according to the target hyperparameters and the target training method.
[0269] Optionally, as an embodiment, the scale-invariant linear layer obtains the prediction classification result according to the following formula:
[0270]
[0271] where Y i represents the prediction classification result corresponding to the weight parameters updated in the i-th iteration; W i represents the weight parameters updated in the i-th iteration; X represents the feature to be classified; S represents the scale constant.
[0272] Optionally, as an embodiment, the target training method includes:
[0273] Processing the updated weight parameters through the following formula to make the norm lengths of the weight parameters of the classification model to be trained before and after the update the same:
[0274]
[0275] where W i+1 represents the weight parameters updated in the (i + 1)-th iteration; W i represents the weight parameters updated in the i-th iteration; Norm0 represents the initial weight norm length of the classification model to be trained.
[0276] In one example, the training device 700 can be used to execute Figure 5 or Figure 6 all or part of the operations in the training methods shown by any one of them. For example, the obtaining unit 710 can be used to execute S510, or all or part of the operations of obtaining the model to be trained, the target hyperparameter (e.g., equivalent learning rate), and the training data; the processing unit 720 can be used to execute S520 or all or part of the operations of iteratively updating the weight parameters of the model to be trained. Among them, the obtaining unit 710 can refer to the communication interface or transceiver in the training device 700; the processing unit 720 can be a processor or chip with computing power in the training device 700.
[0277] Figure 10 is a schematic block diagram of the hyperparameter search device provided by this application.
[0278] It should be understood that the search device 800 can execute Figure 7 or Figure 8 the search methods shown. The search device 800 includes: an obtaining unit 810 and a processing unit 820.
[0279] Among them, the obtaining unit 810 is used to obtain candidate values of the target hyperparameter, where the target hyperparameter is used to control the gradient update step of the model to be trained for classification, and the model to be trained for classification includes a scale-invariant linear layer, and the scale-invariant linear layer makes the prediction classification result remain unchanged when the weight parameters of the model to be trained for classification are multiplied by any scaling factor; the processing unit 820 is used to obtain the performance parameter of the model to be trained for classification according to the candidate value and the target training method, where the target training method makes the norm lengths of the weight parameters of the model to be trained for classification before and after update the same, and the performance parameter includes the accuracy of the model to be trained for classification; determine the target value of the target hyperparameter from the candidate values according to the performance parameter.
[0280] Optionally, as an embodiment, the accuracy of the model to be trained for classification corresponding to the target value is greater than the accuracy of the model to be trained for classification corresponding to other candidate values among the candidate values.
[0281] Optionally, as an embodiment, the processing unit 810 is specifically used for:
[0282] Uniformly divide according to the initial search range of the target hyperparameter to obtain candidate values of the target hyperparameter.
[0283] Optionally, as an embodiment, the processing unit 820 is further used for:
[0284] Update the initial search range of the target hyperparameter according to the current training step number, the pre-configured training step number, and the change trend of the accuracy of the classification model to be trained.
[0285] Optionally, as an embodiment, the processing unit 820 is specifically configured to:
[0286] If the current training step number is less than the pre-configured training step number, update the upper boundary of the initial search range of the target hyperparameter to the candidate value of the target hyperparameter corresponding to the optimal accuracy of the classification model to be trained in the current training step number.
[0287] Optionally, as an embodiment, the processing unit 820 is specifically configured to:
[0288] If the current training step number is equal to the pre-configured training step number, update the upper boundary of the initial search range of the target hyperparameter to a first candidate value, and update the lower boundary of the search range of the target hyperparameter to a second candidate value. The first candidate value and the second candidate value refer to the candidate values adjacent to the candidate value of the target hyperparameter corresponding to the optimal accuracy of the classification model to be trained.
[0289] Optionally, as an embodiment, the scale-invariant linear layer obtains the predicted classification result according to the following formula:
[0290]
[0291] where Y i represents the predicted classification result corresponding to the weight parameter updated in the i-th iteration; W i represents the weight parameter updated in the i-th iteration; X represents the feature to be classified; S represents the scale constant.
[0292] Optionally, as an embodiment, the target training method includes:
[0293] Process the updated weight parameter through the following formula so that the norm lengths of the weight parameters before and after the update of the classification model to be trained are the same:
[0294]
[0295] where W i+1 represents the weight parameter updated in the (i + 1)-th iteration; W i represents the weight parameter updated in the i-th iteration; Norm0 represents the initial weight norm length of the classification model to be trained.
[0296] In one example, the search device 800 can be used to execute Figure 7 Or Figure 8All or part of the operations in the search method shown in any of them. For example, the obtaining unit 810 can be used to execute S610, obtaining the initial search range or obtaining all or part of the operations in the data and the model; the processing unit 820 can be used to execute S620, S630 or all or part of the operations in the search for the target hyperparameters. Among them, the obtaining unit 810 can refer to the communication interface or transceiver in the search device 800; the processing unit 820 can be a processor or chip with computing capabilities in the search device 800.
[0297] It should be noted that the above training device 700 and search device 800 are embodied in the form of functional units. The term "unit" here can be implemented in software and / or hardware forms, and no specific limitation is made thereto.
[0298] For example, the "unit" can be a software program, a hardware circuit, or a combination of both that implements the above functions. The hardware circuit may include an application specific integrated circuit (ASIC), an electronic circuit, a processor (such as a shared processor, a dedicated processor, or a group of processors, etc.) for executing one or more software or firmware programs, a memory, a merged logic circuit, and / or other suitable components that support the described functions.
[0299] Therefore, the units in each example described in the embodiments of the present application can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0300] Figure 11 It is a schematic diagram of the hardware structure of the training device of the classification model provided by the embodiments of the present application.
[0301] Figure 11 The training device 900 shown includes a memory 910, a processor 920, a communication interface 930, and a bus 940. Among them, the memory 910, the processor 920, and the communication interface 930 are communicatively connected to each other through the bus 940.
[0302] The memory 910 may be a read only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 910 may store a program, and when the program stored in the memory 910 is executed by the processor 920, the processor 920 is configured to execute each step of the training method of the classification model according to the embodiments of the present application; for example, to execute Figure 5 or Figure 6 each step shown.
[0303] It should be understood that the training device shown in the embodiments of the present application may be a server or a chip configured in the server.
[0304] Among them, the training device may be a device with the function of training a classification model. For example, it may include any device known in the current technology; alternatively, the training device may also refer to a chip with the function of training a classification model. The training device may include a memory and a processor; the memory may be used to store program codes, and the processor may be used to call the program codes stored in the memory to implement the corresponding functions of the computing device. The processor and the memory included in the computing device may be implemented by a chip, and no specific limitation is made here.
[0305] For example, the memory may be used to store the relevant program instructions of the training method of the classification model provided in the embodiments of the present application, and the processor may be used to call the relevant program instructions of the training method of the classification model stored in the memory.
[0306] The processor 920 may be a general-purpose central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute relevant programs to implement the training method of the classification model according to the method embodiments of the present application.
[0307] The processor 920 may also be an integrated circuit chip with the ability to process signals. During the implementation process, each step of the training method of the classification model of the present application may be completed by the integrated logic circuit in the hardware of the processor 920 or the instructions in the form of software.
[0308] The above-mentioned processor 920 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 910, and the processor 920 reads the information in the memory 910 and combines its hardware to complete the implementation of the present application Figure 9 the functions that the units included in the training device shown need to execute, or, execute the Figure 5 or Figure 6 training method of the classification model shown.
[0309] The communication interface 930 uses a transceiver device such as, but not limited to, a transceiver to implement the communication between the training device 900 and other devices or communication networks.
[0310] The bus 940 may include a path for transmitting information between various components of the training device 900 (for example, the memory 910, the processor 920, the communication interface 930).
[0311] Figure 12 It is a schematic diagram of the hardware structure of the hyperparameter search device provided by the embodiments of the present application.
[0312] Figure 12 The search device 1000 shown includes a memory 1010, a processor 1020, a communication interface 1030, and a bus 1040. Among them, the memory 1010, the processor 1020, and the communication interface 1030 are communicatively connected to each other through the bus 1040.
[0313] The memory 1010 may be a read only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1010 may store a program. When the program stored in the memory 1010 is executed by the processor 1020, the processor 1020 is configured to perform each step of the hyperparameter search method according to the embodiments of the present application. For example, it is configured to perform Figures 7 to 8 each step shown.
[0314] It should be understood that the search device shown in the embodiments of the present application may be a server or a chip configured in the server.
[0315] Among them, the search device may be a device with the hyperparameter search function. For example, it may include any device known in the current technology. Alternatively, the search device may also refer to a chip with the hyperparameter search function. The search device may include a memory and a processor. The memory may be configured to store program codes, and the processor may be configured to call the program codes stored in the memory to implement the corresponding functions of the computing device. The processor and the memory included in the computing device may be implemented by a chip, which is not specifically limited herein.
[0316] For example, the memory may be configured to store the relevant program instructions of the hyperparameter search method provided in the embodiments of the present application, and the processor may be configured to call the relevant program instructions of the hyperparameter search method stored in the memory.
[0317] The processor 1020 may be a general-purpose central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), or one or more integrated circuits, configured to execute relevant programs to implement the hyperparameter search method according to the method embodiments of the present application.
[0318] The processor 1020 may also be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the hyperparameter search method of the present application may be completed by the integrated logic circuit in the hardware of the processor 1020 or the instructions in software form.
[0319] The above-mentioned processor 1020 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, registers, etc. This storage medium is located in the memory 1010, and the processor 1020 reads the information in the memory 1010 and combines its hardware to complete the implementation of the present application Figure 10 the functions required to be executed by the units included in the search device shown in Figures 7 to 8 the search method for hyperparameters shown in
[0320] The communication interface 1030 uses a transceiver device such as, but not limited to, a transceiver to implement the communication between the search device 1000 and other devices or communication networks.
[0321] The bus 1040 may include a path for transmitting information between the various components of the search device 1000 (for example, the memory 1010, the processor 1020, the communication interface 1030).
[0322] It should be noted that although the above-mentioned training device 900 and search device 1000 only show the memory, the processor, and the communication interface, in the specific implementation process, those skilled in the art should understand that the training device 900 and the search device 1000 may also include other devices necessary for normal operation. At the same time, according to specific needs, those skilled in the art should understand that the above-mentioned training device 900 and search device 1000 may also include hardware devices for implementing other additional functions. In addition, those skilled in the art should understand that the above-mentioned search device 700 may also only include the devices necessary for implementing the embodiments of the present application, and do not have to include Figure 11 or Figure 12 all the devices shown in
[0323] Exemplarily, an embodiment of the present application further provides a chip, which includes a transceiver unit and a processing unit. Among them, the transceiver unit may be an input / output circuit or a communication interface; the processing unit is a processor, a microprocessor or an integrated circuit integrated on the chip; the chip can execute the training method in the above method embodiment.
[0324] Exemplarily, an embodiment of the present application further provides a chip, which includes a transceiver unit and a processing unit. Among them, the transceiver unit may be an input / output circuit or a communication interface; the processing unit is a processor, a microprocessor or an integrated circuit integrated on the chip; the chip can execute the search method in the above method embodiment.
[0325] Exemplarily, an embodiment of the present application further provides a computer-readable storage medium, on which instructions are stored, and when the instructions are executed, the training method in the above method embodiment is executed.
[0326] Exemplarily, an embodiment of the present application further provides a computer-readable storage medium, on which instructions are stored, and when the instructions are executed, the search method in the above method embodiment is executed.
[0327] Exemplarily, an embodiment of the present application further provides a computer program product containing instructions, and when the instructions are executed, the training method in the above method embodiment is executed.
[0328] Exemplarily, an embodiment of the present application further provides a computer program product containing instructions, and when the instructions are executed, the search method in the above method embodiment is executed.
[0329] It should be understood that the processor in the embodiment of the present application may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0330] It should also be understood that the memory in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0331] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0332] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. Additionally, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context.
[0333] In this application, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0334] It should be understood that in various embodiments of the present application, the magnitudes of the sequence numbers of the above processes do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0335] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0336] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0337] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling, direct coupling, or communication connection can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.
[0338] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0339] In addition, the functional units in each embodiment of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0340] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0341] As described above, the above are only specific implementation manners of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A training method for a classification model, characterized in that, Including: Obtain the target hyperparameter of the classification model to be trained, where the target hyperparameter is the equivalent learning rate, and the target hyperparameter is used to control the gradient update step of the classification model to be trained. The classification model to be trained includes a scale-invariant linear layer, and the scale-invariant linear layer makes the prediction classification result remain unchanged when the weight parameter of the classification model to be trained is multiplied by any scaling factor. Update the weight parameter of the classification model to be trained according to the target hyperparameter and the target training method to obtain the trained classification model. The target training method makes the norm lengths of the weight parameters of the classification model to be trained before and after the update the same.
2. The training method according to claim 1, characterized in that, The weight parameter of the trained classification model is obtained by iteratively updating through the backpropagation algorithm multiple times according to the target hyperparameter and the target training method.
3. The training method according to claim 1 or 2, characterized in that The scale-invariant linear layer obtains the prediction classification result according to the following formula: Among them, Y i represents the predicted classification result corresponding to the weight parameter updated in the i-th iteration; W i represents the weight parameter updated in the i-th iteration; X represents the feature to be classified; S represents the scale constant.
4. The training method according to claim 1 or 2, characterized in that The target training method includes: Process the updated weight parameter through the following formula to make the norm lengths of the weight parameters of the classification model to be trained before and after the update the same: Among them, W i+1 represents the weight parameter updated in the (i + 1)-th iteration; W i represents the weight parameter updated in the i-th iteration; Norm0 represents the initial weight norm of the classification model to be trained.
5. A search method for hyperparameters, characterized in that, Including: Obtain the candidate value of the target hyperparameter, where the target hyperparameter is the equivalent learning rate, and the target hyperparameter is used to control the gradient update step of the classification model to be trained. The classification model to be trained includes a scale-invariant linear layer, and the scale-invariant linear layer makes the prediction classification result remain unchanged when the weight parameter of the classification model to be trained is multiplied by any scaling factor. Obtain the performance parameter of the classification model to be trained according to the candidate value and the target training method. The target training method makes the norm lengths of the weight parameters of the classification model to be trained before and after the update the same. The performance parameter includes the accuracy of the classification model to be trained. Determine the target value of the target hyperparameter from the candidate values according to the performance parameter.
6. The search method according to claim 5, wherein The accuracy of the classification model to be trained corresponding to the target value is greater than the accuracy of the classification model to be trained corresponding to other candidate values among the candidate values.
7. The search method according to claim 5 or 6, characterized in that, The obtaining of the candidate value of the target hyperparameter includes: Perform uniform partitioning according to the initial search range of the target hyperparameter to obtain the candidate value of the target hyperparameter.
8. The search method according to claim 7, wherein Also including: Update the initial search range of the target hyperparameter according to the current training step number, the pre-configured training step number, and the change trend of the accuracy of the classification model to be trained.
9. The search method according to claim 8, wherein, The updating of the initial search range of the target hyperparameter according to the current training step number, the pre-configured training step number, and the change trend of the accuracy of the classification model to be trained includes: If the current training step number is less than the pre-configured training step number, update the upper boundary of the initial search range of the target hyperparameter to the candidate value of the target hyperparameter corresponding to the optimal accuracy of the classification model to be trained in the current training step number.
10. The search method according to claim 8, wherein The updating of the initial search range of the target hyperparameter according to the current training step number, the pre-configured training step number, and the change trend of the accuracy of the classification model to be trained includes: If the current training step is equal to the pre-configured training step, update the upper boundary of the initial search range of the target hyperparameter to a first candidate value, and update the lower boundary of the search range of the target hyperparameter to a second candidate value. The first candidate value and the second candidate value refer to the candidate values adjacent to the candidate value of the target hyperparameter corresponding to the optimal accuracy of the classification model to be trained.
11. The search method according to any one of claims 5 or 6, 8 to 10, characterized in that, The scaling-invariant linear layer obtains the predicted classification result according to the following formula: Among them, Y i represents the predicted classification result corresponding to the weight parameter updated in the i-th iteration; W i represents the weight parameter updated in the i-th iteration; X represents the feature to be classified; S represents the scale constant.
12. The search method according to any one of claims 5 or 6, 8 to 10, characterized in that The target training method includes: Process the updated weight parameters through the following formula to make the norms of the weight parameters of the classification model to be trained before and after the update the same: Among them, W i+1 represents the weight parameter updated in the (i + 1)-th iteration; W i represents the weight parameter updated in the i-th iteration; Norm0 represents the initial weight norm of the classification model to be trained.
13. A training device for a classification model, characterized in that including: An acquisition unit, configured to acquire a target hyperparameter of a classification model to be trained. The target hyperparameter is an equivalent learning rate, and the target hyperparameter is used to control the gradient update step of the classification model to be trained. The classification model to be trained includes a scaling-invariant linear layer, and the scaling-invariant linear layer enables the predicted classification result output when the weight parameter of the classification model to be trained is multiplied by any scaling factor to remain unchanged. A processing unit, configured to update the weight parameter of the classification model to be trained according to the target hyperparameter and the target training method to obtain a trained classification model. The target training method makes the norms of the weight parameters of the classification model to be trained before and after the update the same.
14. The training device according to claim 13, wherein The weight parameter of the trained classification model is obtained by iteratively updating through the backpropagation algorithm according to the target hyperparameter and the target training method.
15. The training device according to claim 13 or 14, characterized in that, The scaling-invariant linear layer obtains the predicted classification result according to the following formula: Among them, Y i represents the predicted classification result corresponding to the weight parameter updated in the i-th iteration; W i represents the weight parameter updated in the i-th iteration; X represents the feature to be classified; S represents the scale constant.
16. The training device according to claim 13 or 14, characterized in that, The target training method includes: Process the updated weight parameters through the following formula to make the norms of the weight parameters of the classification model to be trained before and after the update the same: Among them, W i+1 represents the weight parameter updated in the (i + 1)-th iteration; W i represents the weight parameter updated in the i-th iteration; Norm0 represents the initial weight norm of the classification model to be trained.
17. A search device for hyperparameters, characterized in that, including: An acquisition unit, configured to acquire candidate values of a target hyperparameter. The target hyperparameter is an equivalent learning rate, and the target hyperparameter is used to control the gradient update step of a classification model to be trained. The classification model to be trained includes a scaling-invariant linear layer, and the scaling-invariant linear layer enables the predicted classification result output when the weight parameter of the classification model to be trained is multiplied by any scaling factor to remain unchanged. A processing unit, configured to obtain a performance parameter of the classification model to be trained according to the candidate values and the target training method. The target training method makes the norms of the weight parameters of the classification model to be trained before and after the update the same. The performance parameter includes the accuracy of the classification model to be trained. Determine the target value of the target hyperparameter from the candidate values according to the performance parameter.
18. The search device according to claim 17, characterized in that, The accuracy of the classification model to be trained corresponding to the target value is greater than the accuracy of the classification model to be trained corresponding to other candidate values among the candidate values.
19. The search device according to claim 17 or 18, characterized in that, Specifically, the processing unit is configured to: Evenly divide the initial search range of the target hyperparameter to obtain candidate values of the target hyperparameter.
20. The search device according to claim 19, characterized in that, The processing unit is further configured to: Update the initial search range of the target hyperparameter according to the current training step number, the pre-configured training step number, and the changing trend of the accuracy of the classification model to be trained.
21. The search device according to claim 20, characterized in that, The processing unit is specifically configured to: If the current training step number is less than the pre-configured training step number, update the upper boundary of the initial search range of the target hyperparameter to the candidate value of the target hyperparameter corresponding to the optimal accuracy of the classification model to be trained in the current training step number.
22. The search device according to claim 20, characterized in that, The processing unit is specifically configured to: If the current training step number is equal to the pre-configured training step number, update the upper boundary of the initial search range of the target hyperparameter to a first candidate value, and update the lower boundary of the search range of the target hyperparameter to a second candidate value. The first candidate value and the second candidate value refer to the candidate values adjacent to the candidate value of the target hyperparameter corresponding to the optimal accuracy of the classification model to be trained.
23. The search device according to any one of claims 17 or 18, 20 to 22, characterized in that The scale-invariant linear layer obtains the predicted classification result according to the following formula: Among them, Y i represents the predicted classification result corresponding to the weight parameter updated in the i-th iteration; W i represents the weight parameter updated in the i-th iteration; X represents the feature to be classified; S represents the scale constant.
24. The search device according to any one of claims 17 or 18, 20 to 22, characterized in that The target training method includes: Process the updated weight parameters through the following formula so that the magnitudes of the weight parameters of the classification model to be trained before and after the update are the same: Among them, W i+1 represents the weight parameter updated in the (i + 1)-th iteration; W i represents the weight parameter updated in the i-th iteration; Norm0 represents the initial weight norm of the classification model to be trained.
25. A training device for a classification model, characterized in that, Including: A memory for storing programs; A processor for executing the programs stored in the memory. When the processor executes the programs stored in the memory, the processor is used to execute the training method according to any one of claims 1 to 4.
26. A search device for hyperparameters, characterized in that, Including: A memory for storing programs; A processor for executing the programs stored in the memory. When the processor executes the programs stored in the memory, the processor is used to execute the search method according to any one of claims 5 to 12.
27. A computer-readable storage medium, characterized in that, Program instructions are stored in the computer-readable storage medium. When the program instructions are run by a processor, the training method according to any one of claims 1 to 4 is implemented.
28. A computer-readable storage medium, characterized in that, Program instructions are stored in the computer-readable storage medium. When the program instructions are run by a processor, the search method according to any one of claims 5 to 12 is implemented.
Citation Information
Cited By
Classification model training method, hyper-parameter searching method, and device
WO2022028323A1