Model hyperparameter value taking method and device, processing core, equipment, chip and medium
By using a many-core processor in a neuromorphic simulation system for parallel simulation, the problem of difficulty in determining hyperparameter values for neural network models was solved, and efficient hyperparameter configuration was achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-21
- Publication Date
- 2026-03-27
AI Technical Summary
In existing technologies, the configuration of hyperparameters in neural network models affects model performance, but suitable values cannot be known in advance, resulting in time-consuming and inefficient methods based on experience or repeated trials.
A model hyperparameter value determination method based on a brain-like simulation system is adopted. Parallel simulation is performed through the processing cores in a many-core system. Brain-like simulation is conducted using sample data and different value batches to determine the target value.
It shortens the time to determine the target values of model hyperparameters, improves the acquisition efficiency, and achieves fast and efficient hyperparameter configuration.
Smart Images

Figure CN116542286B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of processor, and in particular, to a model hyperparameter value determination method and device, a processing core, an equipment, a chip and a medium. BACKGROUND
[0002] For a neural network model, the configuration of model hyperparameters will affect the performance of the neural network model. When using the neural network model to process a given problem, the suitable values of the hyperparameters cannot be known in advance. The suitable values of the hyperparameters are explored through experience or through repeated experiments, which is time-consuming and inefficient. SUMMARY
[0003] The present disclosure provides a model hyperparameter value determination method and device, a processing core, an equipment, a chip and a medium.
[0004] In a first aspect, the present disclosure provides a model hyperparameter value determination method based on a brain-like simulation system, the brain-like simulation system comprising at least one cluster of many-core systems, a cluster in the many-core system comprising one or more processing cores, the method comprising: obtaining, by the processing cores of each cluster, at least one value batch of a model hyperparameter of a neural network model from a value set of the model hyperparameter, each value batch comprising at least two different values of the model hyperparameter; broadcasting sample data of the neural network model to the processing cores of each cluster, so that the processing cores of each cluster perform brain-like simulation running on the neural network model according to the sample data and the at least one value batch of the model hyperparameter, to obtain a calculation result of the at least one value batch of each cluster; and determining a target value of the model hyperparameter corresponding to the sample data according to the calculation result of the at least one value batch of each cluster.
[0005] In a second aspect, the present disclosure provides a model hyperparameter value determination device based on a brain-like simulation system, the brain-like simulation system comprising at least one cluster of many-core systems, a cluster in the many-core system comprising one or more processing cores, the device comprising: a batch determination module configured to obtain, by the processing cores of each cluster, at least one value batch of a model hyperparameter of a neural network model from a value set of the model hyperparameter, each value batch comprising at least two different values of the model hyperparameter; a simulation running module configured to broadcast sample data of the neural network model to the processing cores of each cluster, so that the processing cores of each cluster perform brain-like simulation running on the neural network model according to the sample data and the at least one value batch of the model hyperparameter, to obtain a calculation result of the at least one value batch of each cluster; and a parameter value determination module configured to determine a target value of the model hyperparameter corresponding to the sample data according to the calculation result of the at least one value batch of each cluster.
[0006] In a third aspect, the disclosure provides a processing core for loading a neural network model to complete deep learning processing, wherein a value of a model hyperparameter of the neural network model is obtained according to the model hyperparameter value determination method.
[0007] In a fourth aspect, the disclosure provides an electronic device, comprising: a plurality of processing cores; and an on-chip network configured to interact data between the plurality of processing cores and external data; wherein one or more instructions are stored in one or more of the processing cores, and the one or more instructions are executed by the one or more processing cores to enable the one or more processing cores to execute the model hyperparameter value determination method.
[0008] In a fifth aspect, the disclosure provides a brain-like computing chip, comprising a many-core system, the many-core system comprising at least one cluster, each cluster comprising the processing core.
[0009] In a sixth aspect, the disclosure provides a computer readable medium having a computer program stored thereon, wherein the computer program, when executed by the processing core, implements the model hyperparameter value determination method.
[0010] The model hyperparameter value determination method and device, processing core, equipment, chip and medium provided by the disclosure can be used to perform brain-like simulation running on hyperparameters of a neural network model in parallel based on a cluster comprising one or more processing cores in a brain-like simulation system, and the sample data input and broadcast to the corresponding hyperparameter instances in each processing core of the cluster undergo different values of the model hyperparameters and the same calculation logic. Through the brain-like simulation running, the calculation results corresponding to different values of the model hyperparameters can be obtained, that is, the calculation results corresponding to different values of the model hyperparameters of each cluster in the many-core system can be obtained through one simulation, the simulation running process is time-saving and efficient, the time spent for determining the target value of the model hyperparameter can be shortened, and the efficiency of obtaining the target value of the model hyperparameter can be improved.
[0011] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the disclosure, nor is it used to limit the scope of the disclosure. Other features of the disclosure will become apparent to those skilled in the art through the following description. BRIEF DESCRIPTION OF DRAWINGS
[0012] The accompanying drawings are included to provide a further understanding of the disclosure and constitute a part of the specification, which together with the embodiments of the disclosure are used to explain the disclosure and do not constitute a limitation of the disclosure. The above and other features and advantages will become more apparent to those skilled in the art through the description of the detailed example embodiments, with reference to the accompanying drawings, in which:
[0013] Figure 1 A flowchart of a model hyperparameter value determination method provided by an embodiment of the present disclosure is shown in FIG. 1.
[0014] Figure 2 A specific flowchart of a simulation running of a neural network model provided by an embodiment of the present disclosure is shown in FIG. 2.
[0015] Figure 3 A specific flowchart of a single brain simulation running provided by an embodiment of the present disclosure is shown in FIG. 3.
[0016] Figure 4 A specific flowchart of determining a target value of a model hyperparameter provided by an embodiment of the present disclosure is shown in FIG. 4.
[0017] Figure 5 A flowchart of a model hyperparameter value determination method provided by an embodiment of the present disclosure is shown in FIG. 1.
[0018] Figure 6 A flowchart of a model hyperparameter value determination method provided by an embodiment of the present disclosure is shown in FIG. 1.
[0019] Figure 7 A block diagram of a model hyperparameter value determination apparatus provided by an embodiment of the present disclosure is shown in FIG. 6.
[0020] Figure 8 A specific block diagram of a simulation running module provided by an embodiment of the present disclosure is shown in FIG. 7.
[0021] Figure 9 A specific block diagram of a parameter value determination module provided by an embodiment of the present disclosure is shown in FIG. 8.
[0022] Figure 10 A block diagram of an electronic device provided by an embodiment of the present disclosure is shown in FIG. 9. DETAILED DESCRIPTION
[0023] In order to enable those skilled in the art to better understand the technical solutions of the present disclosure, the exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to help understanding, and should be considered as merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Also, in order to be clear and concise, the description below omits the description of well-known functions and structures.
[0024] The embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.
[0025] As used herein, the term “and / or” includes any and all combinations of one or more of the associated listed items.
[0026] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. "Coupled" or "connected" or similar terms are not restricted to physical or mechanical connections or associations, but can also include electrical connections, whether direct or indirect.
[0027] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and the present disclosure, and will not be interpreted in an overly literal or overly formal sense unless expressly so defined herein.
[0028] In the field of machine learning, a neural network (NN) is a computing system that processes data by mimicking the structure of biological neural networks, which are generated by a large number of neurons that are widely interconnected, and has strong non-linear and adaptive data processing capabilities. A neural network is mathematically modeled by using a mathematical model of a neuron to describe the neural network, i.e., to obtain a neural network model in the embodiments of the present disclosure.
[0029] The neural network model usually includes two types of parameters: model parameters and model hyperparameters.
[0030] Among them, the model parameter is a configuration variable inside the model; the value of the model parameter can be obtained by data estimation or data learning. As an example, the model parameters may, for example, include weights and biases in a neural network, support vectors in a support vector machine, etc. The model hyperparameter (which can be referred to as a hyperparameter in the following embodiments) is a configuration outside the model, and the value of the hyperparameter cannot be obtained from data estimation or data learning, but needs to be manually set by a practitioner. As an example, the model hyperparameters may, for example, include: learning rate, number of iterations, number of hidden layers, number of neurons per layer, etc. in a neural network.
[0031] In practical application scenarios, the configuration of the model hyperparameters does not affect the calculation process and shares time dimension information, such as the number of time steps and the step length. However, the configuration of the model hyperparameters affects the performance of the neural network model, such as the use cost of the computing resources (memory and running time), the learning speed, the quality of the learned model, and the ability to obtain correct results on new input data.
[0032] For example, the configuration of the hyperparameters affects the score of the model on the key metrics of the validation set / test set, such as the classification accuracy of the model on the validation set.
[0033] For a specific problem to be solved by the neural network model, the appropriate values of the model hyperparameters cannot be known in advance. Therefore, the values of the model hyperparameters are often found by practitioners using empirical rules, copied from values used for other problems, or obtained through repeated trials, but these methods are time-consuming and inefficient.
[0034] The model hyperparameter value determination method and device, processing core, equipment, chip, and medium provided by the embodiments of the present disclosure can quickly and efficiently determine the target values of the model hyperparameters. In order to better understand the present disclosure, the model hyperparameter value determination method and device, processing core, equipment, chip, and medium according to the embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. It should be noted that these embodiments are not intended to limit the scope of the present disclosure.
[0035] Figure 1 The flowchart of the model hyperparameter value determination method provided by the embodiments of the present disclosure.
[0036] With reference to Figure 1 The embodiments of the present disclosure provide a model hyperparameter value determination method 100 based on a brain-like simulation system, the brain-like simulation system includes at least one cluster of many-core systems, and the cluster in the many-core system includes one or more processing cores. In some embodiments, the model hyperparameter value determination method 100 can include the following steps.
[0037] S110, at least one value batch of the model hyperparameters is obtained from a value set of the model hyperparameters of the neural network model by the processing cores of each cluster, and each value batch includes at least two different values of the model hyperparameters.
[0038] S120, the sample data of the neural network model is broadcast to the processing cores of each cluster, so that the processing cores of each cluster perform brain-like simulation running on the neural network model according to the sample data and the at least one value batch of the model hyperparameters, and obtain the calculation results of the at least one value batch of each cluster.
[0039] S130, the target values of the model hyperparameters corresponding to the sample data are determined according to the calculation results of the at least one value batch of each cluster.
[0040] In the embodiments of the present disclosure, each cluster in the many-core system includes one or more processing cores, each cluster can perform brain-like computing on the neural network according to instances of multiple hyperparameters of the neural network in parallel through the included processing cores, and the many-core system can broadcast input sample data to the processing cores of each cluster, that is, broadcast the sample data input to correspond to each instance of the hyperparameters, and the calculation results of the brain-like simulation running of each cluster, so as to determine the target value of the model hyperparameters corresponding to the sample data according to the calculation results.
[0041] In the model hyperparameter value determination method of the embodiments of the present disclosure, when the processing core performs simulation running of the neural network model, the sample data undergoes different values of the model hyperparameters and the same calculation logic, and through the simulation running, the calculation results corresponding to different values of the model hyperparameters can be obtained. The simulation running process is time-consuming and efficient, which shortens the time spent in determining the target value of the model hyperparameters and improves the acquisition efficiency of the target value of the model hyperparameters.
[0042] In some embodiments, the neural network model can have many hyperparameters. Generally, when the value of the hyperparameter is determined by experience, copying the value used for other problems, or repeated trials, etc., the suitable value of the hyperparameter needs to be explored by trying different values of the model hyperparameter one by one and simulating multiple times, which results in a long time-consuming and low-efficiency process of determining the value of the model hyperparameter. According to the model hyperparameter value determination method of the embodiments of the present disclosure, when the neural network model has many hyperparameters, the optimal hyperparameters can be quickly and efficiently determined.
[0043] In the embodiments of the present disclosure, the neural network can include an artificial neural network (ANN) and a spiking neural network (SNN), etc.
[0044] In the embodiments of the present disclosure, the neural network can process information through deep learning (DL) or brain-like computing (Neuromorphic Computing), etc.
[0045] Deep learning is a kind of machine learning. Deep learning forms more abstract high-level representation attribute classes or features by combining low-level features to discover distributed feature representation of data. The purpose of deep learning is to establish a neural network that simulates the brain to analyze and learn. It simulates the mechanism of the human brain to process data, such as images, sounds and texts, to perform target recognition, semantic segmentation, reinforcement learning, etc.
[0046] Brain-inspired computing is to simulate the process of brain processing information by using neuromorphic computing. The brain-inspired computing can generally perform model computing based on a spiking neural network, in which data is transmitted between neurons in the form of pulses. The internal information transmission of the spiking neural network is completed by a pulse sequence, and a neuron outputs an electrical pulse signal, which is weighted into a current signal of different intensities.
[0047] Exemplarily, the neural network model in the embodiments of the present disclosure can adopt any one of the following neuron models: a Leaky Integrate-and-Fire (LIF) model, a Hodgkin-Huxley model (HH model for short), and a FitzHugh-Nagumo model (FHN model for short).
[0048] In the embodiments of the present disclosure, the model hyperparameter value setting method based on the brain-inspired simulation system can be applied to a many-core processor or a circuit adopting a many-core system architecture. For example, the many-core processor can be a neuromorphic chip. The neuromorphic chip can simulate the working mechanism of the human brain. In some embodiments, the neuromorphic chip can include a plurality of processing cores, each of which can be used to perform different computing tasks (for example, computing tasks of a neural network), and the same computing task can also be allocated to different processing cores for execution; each processing core can simulate a predetermined number of neurons with biological significance inside, so that the neuromorphic chip performs different computing tasks by simulating neurons. In some embodiments, the neuromorphic chip can perform both artificial neural network computing and spiking neural network computing.
[0049] It should be understood that, whether it is deep learning or brain-inspired computing, the underlying computing model used is a neural network model, which is encoded based on the frequency of neural pulses. The neural network model in the embodiments of the present disclosure includes but is not limited to the above-mentioned neuron models. In actual application scenarios, the neural network model involved in the model hyperparameter value setting method of the embodiments of the present disclosure can be selected according to actual conditions, and the specific type of neural network is not specifically limited in the embodiments of the present disclosure.
[0050] In steps S110 and S120, the processing cores in each cluster can obtain the value set of the model hyperparameters and the sample data of the neural network model in the following multiple ways.
[0051] In some embodiments, the value set of the model hyperparameters and the sample set of the neural network can be input to the processing cores in advance, so that the processing cores can directly obtain the value set when performing the model hyperparameter value setting method, and obtain any sample data from the sample set.
[0052] In some embodiments, the processing core can obtain the value set of the model hyperparameters and the sample data of the neural network model from the received computing instruction.
[0053] As an example, the computing instruction can include the value set of the model hyperparameters and a sample set of the neural network model, the sample set including one or more sample data of the neural network model. The processing core obtains any sample data from the sample set in the case of executing the model hyperparameter value method.
[0054] As another example, in the case that the processing core in each cluster of the many-core system executes the model hyperparameter value method for the first time, the received computing instruction includes the value set of the model hyperparameters and one sample data of the neural network model. In the case that the processing core executes the model hyperparameter value method again, the received new computing instruction can only include other sample data of the neural network model.
[0055] In the embodiments of the present disclosure, the processing core in each cluster of the many-core system can flexibly select the information obtaining manner of the value set of the model hyperparameters of the neural network model and the sample data of the neural network model, which is not specifically limited in the embodiments of the present disclosure.
[0056] In some embodiments, the model hyperparameters of the neural network model in step S110 can be a single hyperparameter or a combination of hyperparameters; and the number of values of at least one model hyperparameter is greater than 1.
[0057] As an example, the model hyperparameters of the neural network model are a single hyperparameter A, and the value set of the hyperparameter A includes the following values {A1, A2, …, An}, where n is an integer greater than 1. According to the value set of the hyperparameter A, n groups of different values of the hyperparameter A can be determined: the first group of values is A1, the second group of values is A2, …, and the n-th group of values is An.
[0058] As an example, the model hyperparameters can be a combination of hyperparameters, such as hyperparameter A, hyperparameter B, and hyperparameter C; and the value set of the combination of hyperparameters includes the following values: hyperparameter A {A1, A2}, hyperparameter B {B1, B2}, and hyperparameter C {C1, C2}. According to the value set of the combination of hyperparameters, different combinations of values of the model hyperparameters include: the first group of values is A1, B1, and C1, the second group of values is A1, B1, and C2, the third group of values is A1, B2, and C1, the fourth group of values is A1, B2, and C2, the fifth group of values is A2, B1, and C1, the sixth group of values is A2, B1, and C2, the seventh group of values is A2, B2, and C1, and the eighth group of values is A2, B2, and C2.
[0059] In some embodiments, if the computing capability of the processing cores of each cluster in the many-core system is sufficient and / or the number of value combination of the value set is small, one value batch of the model hyperparameters can include all value combinations of each model hyperparameter of the neural network model; if the computing capability of the processing cores is insufficient and / or the number of value combination of the value set is large, one value batch of the model hyperparameters can include part of the value combinations of each model hyperparameter of the neural network model, that is, all value combinations are divided into multiple value batches.
[0060] As can be known from the description of the above embodiments, the model hyperparameters of the neural network model are not limited to be a single hyperparameter or a combination of hyperparameters, and the total number of different values of the model hyperparameters is equal to the product of the number of values of each hyperparameter. Therefore, when the number of hyperparameters of the neural network model is large, the advantages of the model hyperparameter value method of the embodiments of the present disclosure can be embodied, and the target values of the model hyperparameters of the neural network model can be quickly and efficiently obtained, and the processing efficiency is improved.
[0061] In some embodiments, in the case where the model hyperparameters are combinations of hyperparameters, the combinations of hyperparameters include shared hyperparameters, and the number of elements in the value set of the shared hyperparameters is 1.
[0062] In this embodiment, considering that in actual application scenarios, the values of one or more hyperparameters in the combination of hyperparameters can be shared by different parameter value combinations, only one hyperparameter value can be defined in the value set of the shared hyperparameters, so as to reduce data redundancy and save data storage space.
[0063] Figure 2 A specific flowchart of simulating and running the neural network model in the embodiments of the present disclosure is shown. As shown in Figure 2 In some embodiments, step S120 can specifically include: S21, for any value batch of the model hyperparameters, using the neural network model, performing single brain simulation running on each group of values of the value batch and the sample data, to obtain the calculation results of each cluster corresponding to each group of values of the value batch, respectively.
[0064] In the embodiments of the present disclosure, through the once brain simulation running of the neural network model by the processing cores of each cluster in the many-core system, the calculation results of each group of different values of the model hyperparameters in the current value batch output by each cluster can be obtained. Compared with the case where the calculation result corresponding to one combination of hyperparameters can be obtained in one simulation of the neural network model, the model simulation running of the present disclosure shortens the total time consumption of the model simulation running, thereby improving the acquisition efficiency of the target values of the model hyperparameters.
[0065] Figure 3 A specific flowchart of single brain simulation running in the embodiments of the present disclosure is shown. As shown inFigure 3 As shown in the foregoing step S21, the step of performing single brain simulation running on each group of values and sample data of the value batch using the neural network model to obtain the calculation results corresponding to each group of values of the value batch, can specifically include the following sub-steps.
[0066] S31, compiling each group of values of the value batch through the graph programming model to obtain the computational graph of the value batch.
[0067] In the embodiments of the present disclosure, the graph programming model is used to determine the content, i.e., the computational graph, that needs to be executed by the parallel computing architecture through neural network programming. Through neural network programming, all behaviors of the computer are determined by the neural network structure and knowledge parameters. The knowledge obtained through training is stored in the form of data and used by specific hardware. The computational graph obtained through the graph programming model is also called a data flow graph. The computational graph can be a directed graph, in which the nodes represent operations and the edges represent data flow between different operations. Through the computational graph of the value batch, parallel computing of each group of values of the model hyperparameters in the current value batch during model simulation running can be implemented, thereby improving the computing capability of the data processing process, saving resources, and reducing power consumption.
[0068] S32, copying the sample data according to the number of value groups of the value batch to obtain the sample data of the number of value groups.
[0069] In the embodiments of the present disclosure, copying the sample data according to the number of value groups of the value batch can realize dimension expansion of the sample data, so that each group of different values of the model hyperparameters can correspond to a sample data, and the sample data corresponding to each group of different values is the same for the current batch.
[0070] S33, performing single brain simulation running on the computational graph and the sample data of the number of value groups using the neural network model to obtain the calculation results corresponding to each group of values of the value batch.
[0071] Through the above steps S31-S33, each group of values of the model hyperparameters is compiled through the compilation of the graph programming model, and the sample data is copied according to the number of value groups of the value batch to expand the sample data. Then, by performing single brain simulation running on the compiled computational graph and the same sample data corresponding to each group of different values of the model hyperparameters using the neural network model, the calculation results corresponding to each group of values of the value batch of each cluster can be obtained. The entire model simulation running process is time-saving and efficient, thereby realizing efficient determination of the target values of the hyperparameters of the neural network model.
[0072] Figure 4A specific flowchart for determining the target value of the model hyperparameters in the embodiments of the present disclosure is shown. As shown in Figure 4 In some embodiments, the step S130 can include the following sub-steps.
[0073] S41, according to the calculation results of each cluster, determine the model effect evaluation index of each group of values of the model hyperparameters in each cluster for the sample data.
[0074] In some embodiments, the model effect evaluation index includes, but is not limited to, at least one of the following index items: accuracy, precision, recall, balanced F score (F1-score), area under the curve (AUC) of the receiver operating characteristic (ROC) curve, lift, gain, model risk discrimination ability index (K-S), and model stability index (PSI).
[0075] It should be noted that different types of neural network models can use different model effect evaluation indexes. In actual application, the number and type of model effect evaluation indexes can be selected according to actual needs.
[0076] S42, take the value of the model hyperparameters with the best processing index as the target value of the model hyperparameters.
[0077] Through the above steps S41 and S42, the model effect evaluation is performed through the calculation results of a single brain simulation run under the sample data of the current neural network model, and the value of the group of model hyperparameters with the best model effect evaluation index in each cluster is taken as the target value of the model hyperparameters under the sample data of the current neural network model. Compared with the method of exploring the suitable value of the hyperparameters through experience or through repeated experiments, the embodiments of the present disclosure can quickly and efficiently obtain the target value of the model hyperparameters.
[0078] Figure 5 A flowchart of the model hyperparameter value determination method of the embodiments of the present disclosure is shown. Figure 5 The same or equivalent steps in Figure 1 are denoted by the same reference numerals. As shown in Figure 5 , the model hyperparameter value determination method 500 is basically the same as the model hyperparameter value determination method 100, except that the model hyperparameter value determination method 500 further includes the following steps.
[0079] S140, determine the target value of the model hyperparameters corresponding to each of the plurality of sample data in the sample data set for each cluster.
[0080] S150, determine the final value of the model hyperparameter from the target values of the model hyperparameter corresponding to the plurality of sample data respectively.
[0081] The model hyperparameter value determination method of the embodiments of the present disclosure can obtain, for each sample data in the sample data set, the calculation results corresponding to different values of the model hyperparameter in at least one value batch through a single brain simulation run of the processing cores of each cluster in the many-core system, and can filter the target values of the model hyperparameter corresponding to the plurality of sample data in the sample data set from the running output results of each cluster through steps S140 and S150, so as to select the final value of the model hyperparameter from the plurality of target values. In the determination process of the final value of the model hyperparameter, the data referred to includes the different values of the model hyperparameter and the plurality of sample data in the sample data set, and the final value of the model hyperparameter obtained is beneficial to obtaining better results when the neural network model processes the sample data set.
[0082] In the embodiments of the present disclosure, the model hyperparameter value determination method is applicable to various types of neural network models. For example, the model hyperparameter value determination method can be used to determine the model hyperparameter value for LIF model, HH model or FHN model. In order to simplify the description, the implementation of the model hyperparameter value determination method is described below by taking LIF model as an example. However, this description cannot be interpreted as limiting the scope or implementation possibility of the present scheme, and the processing method of other neural network models other than LIF model is consistent with the processing method of LIF model.
[0083] The specific flow of the model hyperparameter value determination method of the exemplary embodiments of the present disclosure will be described below in combination with Figure 6 and the operation logic of LIF model. Figure 6 The flowchart of the model hyperparameter value determination method of the embodiments of the present disclosure is shown.
[0084] In some embodiments, the operation logic performed by LIF model can be represented as the following expression (1).
[0085]
[0086] In the above expression (1), F represents the fired pulse signal, V represents the membrane potential, F = (V >= V th )? 1:0 represents that the neuron node fires the pulse signal F in the case that the membrane potential V is greater than or equal to the preset firing threshold V th (threshold value of membrane potential); V upd represents the membrane potential, V reset represents the reset voltage (also referred to as reset potential), V upd = (F) x Vreset +(1-F) x V represents resetting the membrane potential V while emitting the pulse signal F upd .
[0087] As can be seen from expression (1), the LIF model has hyperparameters such as a firing threshold V th , a reset voltage Vreset, and so on. The sample dimension of V is [1, D], where D is the number of neurons. When exploring the value of the hyperparameter, the value set of the hyperparameter V reset is {-65mV, -60mV}, and the value set of the hyperparameter V th is {10mV, 20mV}.
[0088] As shown in Figure 6 , in some embodiments, the model hyperparameter value determination method includes the following steps.
[0089] S61, according to the value set of each model hyperparameter of the neural network model, at least one value batch of the model hyperparameters is determined, and each value batch includes at least two groups of different values of the model hyperparameters.
[0090] In this step, the value set of each model hyperparameter is used to represent the value traversal range of each model hyperparameter. Obtaining the value of the model hyperparameter from the value traversal range can be called value exploration of the model hyperparameter.
[0091] As an example, when exploring the value of the hyperparameter, the exploration space of the hyperparameter can be determined according to the obtained value set of each model hyperparameter. Next, according to the exploration space of the hyperparameters V reset and V th of the LIF model, the exploration space is used to indicate at least two groups of different values of the model hyperparameters.
[0092] Table 1: Each group of values of the model hyperparameters
[0093] Each group of values of the index Reset voltage Vreset Dispensing threshold V th ]]> 0 -65 mV 10 mV 1 -60 mV 10 mV 2 -65 mV 20 mV 3 -60 mV 20 mV
[0094] In some embodiments, at least one value batch of the model hyperparameters can be determined according to the processing capability of the processing core and the total number of groups of values of the model hyperparameters.
[0095] As an example, according to the processing capability of the processing core, the processing core determines one value batch of the model hyperparameters according to each group of values of the model hyperparameters shown in Table 1.
[0096] S62, the batch size is determined according to the number of value groups of the value batch.
[0097] In some embodiments, if each group of values in the value batch of the model hyperparameters is stored in the form of a data table, the batch size of the value batch can be determined according to the length of the data table corresponding to each group of values in the value batch, that is, the number of value groups in the value batch.
[0098] S63, compiling each group of values in the value batch by the graph programming model to obtain a computation graph of the value batch.
[0099] In this step, the computation graph of the value batch contains the corresponding model hyperparameter value array.
[0100] S64, replicating the sample data of the input neural network model according to the batch size to obtain sample data of the batch size.
[0101] In this step, the sample data is replicated according to the number of value groups in the value batch, which expands the dimension of the sample data for subsequent processing.
[0102] Through the above steps S61-S64, the value batch of the model hyperparameters required for the simulation and running of the neural network model can be obtained, and the same sample data corresponding to each group of values in the value batch is obtained for subsequent parallel operation of the neural network model on the compiled computation graph of the value batch and the sample data.
[0103] In the embodiments of the present disclosure, after determining at least two different groups of values of the model hyperparameters according to the value set of each model hyperparameter of the neural network model in the above step S61, the data dimension of each model hyperparameter of the neural network model is expanded; in the above step S64, the data dimension of the sample data is expanded based on the number of value groups in the value batch. When the processing core uses the neural network model for parallel operation, the data dimension of each variable in the neural network model is expanded according to the number of value groups in the value batch.
[0104] As an example, when performing simulation and running of the LIF model, the data dimensions of each variable and hyperparameter are as follows: membrane potential V ∈ [4, D], fired pulse signal F ∈ [4, D], membrane potential V upd ∈ [4, 1] = [-65mV, -60mV, -65mV, -60mV], firing threshold V th ∈ [4, 1] = [10mV, 10mV, 20mV, 20mV], where D is the number of neurons.
[0105] S65 uses a neural network model to simulate the computation graph and the sample data of the batch size, and obtains the calculation results corresponding to each different set of model hyperparameter values for the sample data.
[0106] S66. Based on the calculation results, determine the model performance evaluation index, and take the optimal value of the model hyperparameters for the sample data as the target value of the model hyperparameters.
[0107] Through steps S61-S66 above, one or more processing cores in each cluster of the many-core system merge the attempted values of each hyperparameter of the neural network model into a batch of data. Each set of values of each hyperparameter in this batch of data is compiled together through a graph programming model. The neural network model is simulated based on the computation graph obtained from the compilation and the sample data of the number of value sets. Through a single brain-like simulation run of the neural network model, the calculation results corresponding to each different set of values of the model hyperparameters under the sample data can be obtained. Thus, based on the calculation results, the target values of the model hyperparameters corresponding to the sample data are obtained.
[0108] According to the model hyperparameter value acquisition method of this disclosure, by performing a single brain-like simulation run on the neural network model in parallel by each cluster in the many-core system, the calculation results corresponding to different values of the model hyperparameters can be obtained for a single input sample data and a set of values of each model hyperparameter. The simulation run is time-saving and efficient, thereby improving the efficiency of obtaining the target values of the model hyperparameters.
[0109] Figure 7 This diagram illustrates the composition of the model hyperparameter value acquisition device provided in an embodiment of the present disclosure.
[0110] Reference Figure 7 This disclosure provides a model hyperparameter value 700 based on a brain-like simulation system. The brain-like simulation system is a many-core system including at least one cluster, and the cluster in the many-core system contains one or more processing cores. The model hyperparameter value 700 includes the following modules.
[0111] The batch determination module 710 is used to obtain at least one batch of model hyperparameter values from the set of model hyperparameter values of the neural network model through the processing kernel of each cluster, wherein each batch of values includes at least two different sets of model hyperparameter values.
[0112] The simulation execution module 720 is used to broadcast the sample data of the neural network model to the processing cores of each cluster, so that the processing cores of each cluster can perform brain-like simulation of the neural network model according to the sample data and at least one batch of values of the model hyperparameters, and obtain the calculation results of at least one batch of values of each cluster.
[0113] The parameter value determination module 730 is configured to determine the target value of the model hyperparameter corresponding to the sample data according to the calculation result of each cluster of at least one value batch.
[0114] In some embodiments, the simulation running module 720 is specifically configured to, for any value batch of the model hyperparameter, perform a single brain simulation running on each group of values of the value batch and the sample data by using the neural network model, to obtain the calculation result of each cluster corresponding to each group of values of the value batch.
[0115] Figure 8 A specific component block diagram of the simulation running module in the embodiments of the present disclosure is shown. As shown in Figure 8 In some embodiments, the simulation running module 720 includes: a graph programming unit 801 configured to compile each group of values of the value batch by using a graph programming model to obtain a calculation graph of the value batch; a data replication unit 802 configured to replicate the sample data according to the number of value groups of the value batch to obtain sample data of the number of value groups; and the simulation running module 720 is specifically configured to perform a single brain simulation running on the calculation graph and the sample data of the number of value groups by using the neural network model, to obtain the calculation result of each cluster corresponding to each group of values of the value batch.
[0116] Figure 9 A specific component block diagram of the parameter value determination module in the embodiments of the present disclosure is shown. As shown in Figure 9 In some embodiments, the parameter value determination module 730 includes: an evaluation unit 901 configured to determine a model effect evaluation index of each group of values of the model hyperparameter in each cluster for the sample data according to the calculation result of each cluster; and a value unit 902 specifically configured to determine the value of the model hyperparameter with the optimal model effect evaluation index as the target value of the model hyperparameter.
[0117] In some embodiments, the sample data is any sample data in a sample data set of the neural network model; the parameter value determination module 730 is further configured to determine the target value of the model hyperparameter corresponding to each of a plurality of sample data in the sample data set; and determine the final value of the model hyperparameter from the target values corresponding to the plurality of sample data.
[0118] In some embodiments, the model hyperparameter is a single hyperparameter or a combination of hyperparameters; and the number of values of at least one model hyperparameter is greater than 1.
[0119] In some embodiments, when the model hyperparameter is a combination of hyperparameters, the combination of hyperparameters includes a shared hyperparameter, and the number of elements in the value set of the shared hyperparameter is 1.
[0120] According to the model hyperparameter value determination apparatus provided in the embodiments of the present disclosure, when the processing core performs simulation running of the neural network model, the sample data undergoes different values of the model hyperparameters and the same calculation logic. Through the simulation running, the calculation results corresponding to different values of the model hyperparameters can be obtained at one time, that is, the calculation results corresponding to different values of the model hyperparameters can be obtained through one simulation, and the simulation running process is time-saving and efficient. Furthermore, the time length for determining the target value of the model hyperparameter can be shortened, and the acquisition efficiency of the target value of the model hyperparameter can be improved.
[0121] Figure 10 A constituent block diagram of an electronic device is shown to illustrate the embodiments of the present disclosure.
[0122] With reference to Figure 10 The embodiments of the present disclosure provide an electronic device, which includes a plurality of processing cores 1001 and a network on chip 1002, wherein the plurality of processing cores 1001 are connected with the network on chip 1002, and the network on chip 1002 is configured to interact data between the plurality of processing cores 1001 and external data.
[0123] The one or more instructions are stored in one or more of the processing cores 1001, and the one or more instructions are executed by the one or more processing cores 1001, so that the one or more processing cores 1001 can perform the model hyperparameter value determination method.
[0124] In some embodiments, the electronic device can be a brain-like chip. Since the brain-like chip can use vectorized calculation, and the weight information and other parameters of the neural network model need to be loaded from an external memory such as a double data rate (DDR) synchronous dynamic random access memory, the batch processing operation efficiency of the embodiments of the present disclosure is relatively high.
[0125] The embodiments of the present disclosure also provide a processing core, which is configured to load a neural network model to complete deep learning processing, wherein the value of the model hyperparameter of the neural network model is obtained according to the model hyperparameter value determination method.
[0126] The embodiments of the present disclosure also provide a brain-like computing chip, which includes a many-core system, and the many-core system includes at least one cluster, and each cluster includes one or more processing cores described in the above embodiments.
[0127] In addition, the embodiments of the present disclosure also provide a computer readable medium, which stores a computer program, wherein the computer program is configured to implement the model hyperparameter value determination method when executed by the processing core.
[0128] Those of ordinary skill in the art will realize and understand that all or some of the steps in the methods disclosed above and the functional modules / units in the systems and devices can be implemented as software, firmware, hardware, and appropriate combinations thereof. In hardware implementation, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, one physical component can have multiple functions, or one function or step can be performed by several physical components in cooperation. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on computer-readable media, which can include computer storage media (or non-transitory media) and communication media (or transitory media). As is well known to those of ordinary skill in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information and which can be accessed by a computer. Furthermore, it is common technical knowledge that communication media typically embodies computer readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism and includes any information delivery media.
[0129] Example embodiments have been disclosed herein and, although the use of specific terms is expressly used herein, they are intended in a generic sense only and are not intended to limit the scope of the present disclosure. In some instances, it will be apparent to those skilled in the art that features, characteristics or / and elements described in connection with a particular embodiment can be used in conjunction with other embodiments unless otherwise explicitly stated. As such, those skilled in the art will appreciate that various changes can be made in form and detail without departing from the scope of the disclosure as set forth in the appended claims.
Claims
1. A method for determining hyperparameter values in a brain-like simulation system, characterized in that, The neuromorphic simulation system includes a many-core system with at least one cluster, wherein each cluster in the many-core system contains one or more processing cores, and the method includes: The processing kernels of each cluster obtain at least one batch of values for the model hyperparameters from the set of values for the model hyperparameters of the neural network model, and each batch of values includes at least two different sets of values for the model hyperparameters. The same set of sample data of the neural network model is broadcast to the processing cores of each cluster, so that the processing cores of each cluster can perform brain-like simulation on the neural network model according to the sample data and at least one batch of values of the model hyperparameters, and obtain the calculation results of the at least one batch of values of each cluster. Based on the calculation results of at least one batch of values for each cluster, determine the target values of the model hyperparameters corresponding to the sample data; The step of performing a brain-like simulation on the neural network model based on the sample data and at least one batch of values for the model hyperparameters to obtain the calculation results for each cluster in at least one batch of values includes: For any batch of values of the model hyperparameters, the neural network model is used to perform a single brain-like simulation run on each group of values in the batch and the sample data to obtain the calculation results of each cluster corresponding to each group of values in the batch.
2. The method according to claim 1, wherein, The process involves using the neural network model to perform a single brain-like simulation run on each group of values in the batch and the sample data, obtaining the calculation results for each cluster corresponding to each group of values in the batch, including: The computational graph of the value batch is obtained by compiling each group of values in the value batch using a graph programming model. Based on the number of value groups in the batch, the sample data is copied to obtain the sample data for the specified number of value groups; Using the neural network model, a single brain-like simulation is performed on the computation graph and the sample data of the number of value groups to obtain the calculation results of each cluster corresponding to each group of values in the batch of values.
3. The method according to claim 1, wherein, The step of determining the target values of the model hyperparameters corresponding to the sample data based on the calculation results of at least one batch of values for each cluster includes: Based on the calculation results of each cluster, determine the model performance evaluation index for each group of hyperparameter values in each cluster for the sample data; The optimal value of the model hyperparameter for evaluating model performance is taken as the target value of the model hyperparameter.
4. The method according to claim 1, wherein, The sample data is any one sample data from the sample data set of the neural network model; the method further includes: Determine the target values of the model hyperparameters corresponding to multiple sample data in the sample dataset; The final values of the model hyperparameters are determined from the target values corresponding to the multiple sample data respectively.
5. The method according to any one of claims 1-4, wherein, The model hyperparameters can be a single hyperparameter or a combination of hyperparameters; At least one of the model hyperparameters has a value greater than 1.
6. The method according to claim 5, wherein, When the model hyperparameters are a combination of hyperparameters, the combination of hyperparameters includes shared hyperparameters, and the set of values for the shared hyperparameters contains 1 element.
7. A device for determining hyperparameter values of a model based on a brain-like simulation system, characterized in that, The neuromorphic simulation system is a many-core system comprising at least one cluster, wherein the cluster in the many-core system contains one or more processing cores, and the device includes: The batch determination module is used to obtain at least one batch of values of the model hyperparameters from the set of values of the model hyperparameters of the neural network model through the processing kernel of each cluster, wherein each batch of values includes at least two different sets of values of the model hyperparameters. The simulation running module is used to broadcast the same sample data of the neural network model to the processing cores of each cluster, so that the processing cores of each cluster can perform brain-like simulation running on the neural network model according to the sample data and at least one batch of values of the model hyperparameters, and obtain the calculation results of the at least one batch of values of each cluster. The parameter value module is used to determine the target value of the model hyperparameter corresponding to the sample data based on the calculation results of at least one batch of values for each cluster. The step of performing a brain-like simulation on the neural network model based on the sample data and at least one batch of values for the model hyperparameters to obtain the calculation results for each cluster in at least one batch of values includes: For any batch of values of the model hyperparameters, the neural network model is used to perform a single brain-like simulation run on each group of values in the batch and the sample data to obtain the calculation results of each cluster corresponding to each group of values in the batch.
8. A processing core, said processing core being used to load a neural network model to perform deep learning processing, wherein, The values of the hyperparameters of the neural network model are obtained by the hyperparameter value method according to any one of claims 1-6.
9. A neuromorphic computing chip, the neuromorphic computing chip comprising a many-core system, the many-core system comprising at least one cluster, each cluster comprising one or more processing cores as described in claim 8.
10. An electronic device, comprising: Multiple processing cores; as well as The on-chip network is configured to interact with data between the multiple processing cores and external data; One or more processing cores store one or more instructions, and the one or more instructions are executed by one or more processing cores to enable one or more processing cores to perform the model hyperparameter value method according to any one of claims 1-6.
11. A computer-readable medium having a computer program stored thereon, wherein, When the computer program is executed by the processing kernel, it implements the model hyperparameter value method as described in any one of claims 1-6.
Citation Information
Patent Citations
Hyper-parameter determination method and device, computer equipment and storage medium
CN112529211A