Learning method with differential entropy of the synaptic weights of a neural network, processing method, computer program, calculator and associated processing system

The learning method addresses the challenge of reducing neural network memory footprint by incorporating an entropic term based on differential entropy into the cost function, achieving efficient data processing and reduced memory usage.

FR3156954A1Pending Publication Date: 2025-06-20COMMISSARIAT A LENERGIE ATOMIQUE ET AUX ENERGIES ALTERNATIVES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
FR2023014326
Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-15
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Existing methods for learning neural networks do not effectively reduce the memory footprint without compromising data processing efficiency, particularly in embedded systems.

Method used

A learning method that determines the synaptic weights of a neural network by minimizing a cost function incorporating an error term and an entropic term based on the total differential entropy of the quantized weights, allowing for simpler and more efficient gradient calculation.

Benefits of technology

This method reduces the memory footprint of neural networks while maintaining data processing efficiency, facilitating the inference of neural networks in resource-constrained environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Method for learning with differential entropy the synaptic weights of a neural network, processing method, computer program, calculator and associated processing system The invention relates to a method for learning synaptic weights of an artificial neural network (RN) configured to process, in particular to classify, data, the artificial neural network being configured to be implemented by an electronic calculator (20) connected to a sensor (15), for the processing of at least one object originating from the sensor (15).The method is computer-implemented and comprises determining the weights of the neural network from training data, each determined weight being a quantized value belonging to a predefined set of quantized values; the weights being determined by minimizing a cost function, the cost function depending on an error term corresponding to a prediction error and an entropic term. The entropic term depends on a total differential entropy of the quantized weights. Figure for abstract: Figure 1.
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Learning method with differential entropy of the synaptic weights of a neural network, processing method, computer program, calculator and associated processing system

[0001] The present invention relates to a method for learning synaptic weights of at least one layer of an artificial neural network, each artificial neuron of a respective layer being able to perform a weighted sum of input value(s), then to apply an activation function to the weighted sum to deliver an output value, each input value being received from a respective element connected as input to said neuron and multiplied by a synaptic weight associated with the connection between said neuron and the respective element, the respective element being an input variable of the neural network or a neuron of a previous layer of the neural network.

[0002] The method is implemented by computer, and comprises learning the weights of the neural network from training data, each weight resulting from said learning being a quantized value belonging to a set of quantized values.

[0003] The invention also relates to a method for processing data, in particular for classifying data, the method being implemented by an electronic computer implementing such an artificial neural network.

[0004] The invention also relates to a computer program comprising software instructions which, when executed by a computer, implement such a learning method.

[0005] The invention also relates to an electronic computer for processing data, in particular for classifying data, via the implementation of such an artificial neural network; as well as an electronic system for processing object(s), comprising a sensor and such an electronic computer connected to the sensor, the computer being configured to process each object originating from the sensor.

[0006] The invention relates to the field of learning artificial neural networks, also known as ANN (Artificial Neural Networks), also called neural networks. Artificial neural networks are, for example, convolutional neural networks, also known as CNN (Convolutional Neural Networks), recurrent neural networks, such as recurrent networks with short-term and long-term memory, also known as LTSM (Long Short-Term Memory), or transformer neural networks (LTSM). Transformera), typically used in the field of natural language processing (NLP).

[0007] The invention also relates to the field of integrated electronic computers, also called chips, for implementing such neural networks, these electronic computers making it possible to use the neural network during an inference phase, after a prior phase of learning the neural network from learning data, the learning phase typically being implemented by computer.

[0008] A neural network is generally composed of a succession of layers of neurons, each of which takes its inputs from the outputs of the previous layer. More precisely, each layer includes neurons and synapses taking their inputs from the outputs of the neurons in the previous layer. Each layer is connected to the next by a plurality of synapses. A synaptic weight is associated with each synapse. It is a number or a distribution, which takes positive as well as negative values. In the case of a dense layer, the input of a neuron is the weighted sum of the outputs of the neurons in the previous layer, the weighting being done by the synaptic weights and followed by activation via an activation function.

[0009] Such neural networks consume significant memory and computational resources for their storage and inference. The resources required may be too large to allow their use in embedded systems.

[0010] A known technique for reducing a memory footprint during the learning phase is based on network quantization. Quantization consists of reducing the number of bits used to encode each synaptic weight, so that the total memory footprint is reduced by the same factor. This quantization is defined by a synaptic weight quantization function.

[0011] The article “Entropy-Constrained Training of Deep Neural Networks” by S. Wiedemann et al describe a learning method of the above type, with a fixed precision quantization of the synaptic weights. To this end, the learning method includes a minimization of a network entropy which makes it possible to measure in particular the explicit complexity of the network in terms of memory footprint in bits, the entropy being measured with respect to an empirical probability of the distribution of the synaptic weights.

[0012] In order to minimize the entropy of the network weights, it is necessary to derive a gradient of the entropy with respect to the network parameters. S. Wiedemann et al then propose to use a continuous relaxation of the entropy in order to make it differentiable.

[0013] However, such a method does not give complete satisfaction.

[0014] The aim of the invention is then to propose a method for learning a neural network which then makes it possible to reduce the memory footprint of the network without loss of data processing efficiency and to facilitate the inference of said neural network by a computer.

[0015] For this purpose, the subject of the invention is a method for learning synaptic weights of at least one layer of an artificial neural network configured to process, in particular to classify, data, the artificial neural network being configured to be implemented by an electronic computer connected to a sensor, for the processing of at least one object from the sensor, each artificial neuron of a respective layer being capable of performing a weighted sum of input value(s), then applying an activation function to the weighted sum to deliver an output value, each input value being received from a respective element connected to the input of said neuron and multiplied by a synaptic weight associated with the connection between said neuron and the respective element, the respective element being an input variable of the neural network or a neuron of a previous layer of the neural network,

[0016] the method being implemented by computer and comprising the following step:

[0017] - determination of the weights of the neural network from training data, each determined weight being a quantified value belonging to a predefined set of quantified values;

[0018] the weights being determined by minimizing a cost function, the cost function depending on an error term corresponding to a prediction error and an entropic term,

[0019] the entropic term depending on a total differential entropy of the quantized weights.

[0020] With the state-of-the-art method, the function obtained after continuous relaxation is unusable in practice, the state-of-the-art method then performs Bayesian approximations to obtain the derivative of the function obtained. These approximations introduce a random factor in the quantification of the weights.

[0021] With the method according to the invention, the cost function also includes an entropic term which depends on the total differential entropy of the quantized weights. This then makes it possible to calculate the gradient of the cost function more simply and more efficiently. The total differential entropy forms a good differentiable approximation of the discrete entropy. The differential entropy then avoids introducing a randomness linked to a Bayesian approximation while reducing the memory footprint of the network.

[0022] The learning method according to the invention then makes it possible to determine the weights of the neural networks by taking into account the entropy of the quantified weights, the entropy being for example measured with respect to an empirical probability of the distribution of synaptic weights.

[0023] In addition, minimizing the entropy makes it possible to increase the potential for compressing the weights during a subsequent optional compression step, which provides additional leverage for reducing the network memory footprint.

[0024] According to other advantageous aspects of the invention, the learning method comprises one or more of the following characteristics, taken in isolation or in all technically possible combinations:

[0025] - the neural network comprises groups of network weights, each group of weight comprising at least one quantized weight belonging to the set, and the total differential entropy is a sum of products for said groups, each product being a product of a cardinality of a respective group and the differential entropy of said group;

[0026] - the cost function is the sum of the error term and the entropic term;

[0027] the entropic term preferably being the product of a multiplicative factor and of an entropic objective function, the entropic objective function depending on the total differential entropy of the quantized weights;

[0028] the cost function preferably still verifying the equation:

[0029] L = L cls + K L entropy

[0030] where L represents the cost function,

[0031] Lcls represents the error term,

[0032] k represents the multiplicative factor, and

[0033] Lentropy represents |the entropic objective function;

[0034] - the entropic objective function further depends on a total entropic objective predefined;

[0035] - the entropic objective function depends on a difference between the entropy total differential and total entropic objective;

[0036] - the minimization of the cost function includes the calculation of a gradient of the cost function;

[0037] - each weight group is associated with a precision parameter equal to the number of bits used for the coding of the at least one weight of the group, and when each precision parameter is an integer and when the distribution of the quantized weights of each group follows a rounded normal law, the derivative - with respect to the weights of a respective group - of the entropic objective function verifies the following equation:

[0038] _ Wg(W) ÛW nln(2)V(W)H

[0039] where Lentropy represents the entropic objective function,

[0040] W represents the respective group of weights,

[0041] V(W) represents the variance of the normal law,

[0042] E(W) represents the expectation of the normal law,

[0043] H represents the total entropic objective, and

[0044] n is the number of quantized weights of the group W;

[0045] - each weight group is associated with a precision parameter equal to the number of bits used for the coding of the at least one weight of the group, and when each precision parameter is a positive real, the entropic objective function depending on the precision parameters, the derivative - with respect to the precision parameter of a respective group - of the entropic objective function verifies the equation:

[0046] “j

[0047] where Lentropy represents the entropic objective function,

[0048] Xw represents the precision parameter of the respective group,

[0049] W represents the respective group of weights,

[0050] Hw represents the differential entropy of the group W,

[0051] H represents the total entropic objective,

[0052] |_. J represents the lower integer part, and

[0053] Tl represents the upper integer part;

[0054] - the predefined set of quantized values ​​includes an odd number of values;

[0055] - the set includes the null value;

[0056] the positive quantized values ​​of the set being preferably symmetrical, negative quantized values ​​of the set with respect to the zero value;

[0057] - the method further comprises, after the step of determining the weights, a step compression of determined weights;

[0058] - compression is achieved via entropy coding; and

[0059] - entropic coding is a coding with asymmetric digital systems;

[0060] the entropic coding preferably being a coding with asymmetric digital systems in an array.

[0061] The invention also relates to a method for processing data, in particular for classifying data, the method being implemented by an electronic computer implementing an artificial neural network, the method comprising:

[0062] - a learning phase of the artificial neural network, and

[0063] - an inference phase of the artificial neural network, during which data, received as input from the electronic calculator, are processed, in particular classified, via the artificial neural network, previously trained during the learning phase,

[0064] the learning phase being carried out by implementing a learning method as defined above.

[0065] The invention also relates to a computer program comprising software instructions which, when executed by a computer, implement a learning method as defined above.

[0066] The invention also relates to an electronic computer for processing data, in particular for classifying data, via the implementation of an artificial neural network, the electronic computer being intended to be connected to a sensor, for the processing of at least one object originating from the sensor, each artificial neuron of a respective layer of the neural network being capable of performing a weighted sum of input value(s), then applying an activation function to the weighted sum to deliver an output value, each input value being received from a respective element connected to the input of said neuron and multiplied by a synaptic weight associated with the connection between said neuron and the respective element, the respective element being an input variable of the neural network or a neuron of a previous layer of the neural network,

[0067] the calculator comprising:

[0068] - an inference module configured to infer the artificial neural network previously trained, for the processing, in particular the classification, of data received as input from the electronic calculator,

[0069] the artificial neural network previously trained during a learning phase being derived from a computer program as defined above.

[0070] According to other advantageous aspects of the invention, the calculator comprises one or more of the following characteristics, taken in isolation or in all technically possible combinations:

[0071] - the inference module is configured to perform the weighted sum of value(s) input, then to apply an activation function to the weighted sum;

[0072] - during the learning phase the weights of the artificial neural network are compressed, and the inference module further comprises a decompression unit configured to decompress said compressed weights;

[0073] the compression preferably being carried out via entropy coding, and the decompression unit then carrying out entropy decoding of the compressed weights; and

[0074] - entropic coding is a coding with asymmetric digital systems in a table, and the decompression unit includes a decoding table for asymmetric digital array coding.

[0075] The invention also relates to an electronic system for processing object(s), comprising a sensor and an electronic computer connected to the sensor, the computer being configured to process at least one object originating from the sensor, the computer being as defined above.

[0076] The invention will appear more clearly on reading the description which follows, given solely by way of non-limiting example, and made with reference to the drawings in which:

[0077] [Fig-1] [Fig.l] is a schematic representation of an electronic system object processing according to the invention, comprising a sensor, and an electronic computer connected to the sensor, the computer being configured to process, via the implementation of an artificial neural network, at least one object from the sensor;

[0078] [Fig.2] [Fig.2] is a schematic representation of a set of values quantified on 4 bits, used for learning the weights of the neural network according to the invention;

[0079] [Fig.3] [Fig.3] is a schematic representation of a decompression unit included in the calculator of [Fig.l], the decompression unit configured to decompress quantized weights compressed by asymmetric array digital systems (tANS) coding;

[0080] [Fig.4] [Fig.4] is a flowchart of a method, according to the invention, of processing of data, in particular data classification, the method being implemented by the electronic computer of [Fig.l], implementing the artificial neural network;

[0081] [Fig.5] [Fig.5] is a representation of the weight distribution of a layer of two separate neural networks, namely a first neural network trained during learning without an entropic term and a second neural network trained during learning with an entropic term according to the invention;

[0082] [Fig.6] [Fig.6] is a view of several curves obtained with a network of neuron trained according to the invention for different quantification accuracies, each curve representing the evolution of a performance of the network as a function of the entropy of the weights of the network for a respective quantification precision;

[0083] [Fig.7] [Fig.7] is a histogram representing the distribution of the number of weights distinct quantized for each layer of a neural network; and

[0084] [Fig.8] [Fig.8] is a representation of the distribution of the memory footprint per layer for two distinct networks, namely for an uncompressed neural network on the one hand and for a neural network compressed by asymmetric digital systems (ANS) coding on the other hand.

[0085] In [Fig.l], an electronic system 10 for processing object(s) is configured to process one or more objects, not shown, and comprises a sensor 15 and an electronic computer 20 connected to the sensor 15, the computer 20 being configured to process at least one object from the sensor 15.

[0086] The electronic processing system 10 is for example an electronic system for detecting object(s), the sensor 15 then being an object(s) detector and the computer 20 being configured to process at least one object detected by the object(s) detector.

[0087] The electronic processing system 10 forms, for example, a face detector capable of recognizing the faces of previously identified persons and / or of detecting faces of unknown persons, i.e. faces of persons who have not been previously identified. The computer 20 then makes it possible to learn the identities of the detected persons, and also to identify unknown persons.

[0088] The sensor 15 is known per se. The sensor 15 is for example an object detector configured to detect one or more objects, or an image sensor configured to take one or more images of a scene, and transmit them to the computer 20.

[0089] Alternatively, the sensor 15 is a sound sensor, an object detection sensor, such as a lidar sensor, a radar sensor, an infrared sensor, a capacitive proximity sensor, an inductive proximity sensor, a Hall effect proximity sensor or even a presence sensor, configured to acquire a characteristic signal as a function of the presence or absence of object(s), then to transmit it to the computer 20.

[0090] The computer 20 is configured to process a set of data(s), the set of data(s) typically corresponding to one or more signals captured by the sensor 15. The computer 20 is then typically configured to interpret a scene captured by the sensor 15, i.e. to identify and / or to recognize a type of one or more elements - such as people or physical objects - present in the captured scene and corresponding to the signal or signals captured by the sensor 15.

[0091] The computer 20 is configured to perform data processing, in particular data classification, via the implementation of an artificial neural network RN, the latter typically comprising several successive processing layers CTi, where i is an integer index greater than or equal to 1. In the example of [Fig.l], the index i is for example equal to 1, 2 and respectively 3, with first CTI, second CT2 and third CT3 processing layers represented in this [Fig.l]. Each respective processing layer CTi comprises, as known per se, one or more artificial neurons 22, also called formal neurons.

[0092] The CTi processing layers are typically arranged successively within the RN neural network, and the artificial neurons 22 of a given processing layer are typically connected at their input to the artificial neurons 22 of the previous layer, and at output to the artificial neurons 22 of the following layer. The artificial neurons 22 of the first layer, such as the first CTI processing layer, are connected at input to the input variables, not represented, of the neural network RN, and the artificial neurons 22 of the last processing layer, such as the third processing layer CT3, are connected at output to the output variables, not represented, of the neural network RN. In the example of [Fig.l], the second processing layer CT2 then forms an intermediate layer whose artificial neurons 22 are connected at input to the artificial neurons 22 of the first processing layer CTI, and at output to the artificial neurons 22 of the third processing layer CT3.

[0093] As known per se, each artificial neuron 22 is associated with an operation, i.e. a type of processing, to be performed by said artificial neuron 22 within the corresponding processing layer. Each artificial neuron 22 is typically capable of performing a weighted sum of input value(s), then applying an activation function to the weighted sum to deliver an output value, each input value being received from a respective element connected to the input of said neuron 22 and multiplied by a synaptic weight associated with the connection between said neuron 22 and the respective element.The respective element connected as input to said neuron 22 is an input variable of the RN neural network when said neuron belongs to a first layer, also called input layer, of said RN neural network; or is a neuron of a previous layer of the RN neural network when said neuron belongs to an intermediate layer or to a last layer, also called output layer, of the RN neural network. As known per se, the activation function, also called thresholding function or transfer function, makes it possible to introduce non-linearity into the processing performed by each artificial neuron. Classic examples of such an activation function are the sigmoid function, the hyperbolic tangent function, and the Heaviside function, and the linear rectification unit function, also called ReLU (from the English Rectified Linear Unit).

[0094] The neural network RN is for example a convolutional neural network, and the processing layers CTI, CT2, CT3 are then typically each chosen from the group consisting of: a convolution layer, a batch normalization layer, a pooling layer, a correction layer and a fully connected layer.

[0095] According to a first aspect of the invention, in the example of [Fig.l], a learning module 25, external to the computer 20, is configured to carry out learning, also called training, of the neural network RN.

[0096] In the example of [Fig.l], the computer 20 then comprises only one inference module 30 configured to infer the previously trained neural network RN, for the processing, in particular the classification, of data received as input from the computer 20. The computer 20 is then configured to use the neural network RN, previously trained, to calculate new output values ​​from new input values. In other words, it is configured to perform only RN neural network inference.

[0097] The computer 20 is preferably an on-board computer, and is typically implemented in the form of a processor or a microcontroller.

[0098] In the example of [Fig.l], the learning module 25 is produced in the form of software, that is to say in the form of a computer program. It is furthermore capable of being recorded on a medium, not shown, readable by a computer. The computer-readable medium is for example a medium capable of storing electronic instructions and of being coupled to a bus of a computer system. By way of example, the readable medium is an optical disk, a magneto-optical disk, a ROM memory, a RAM memory, any type of non-volatile memory (for example EPROM, EEPROM, FLASH, NVRAM), a magnetic card or an optical card. A computer program comprising software instructions is then stored on the readable medium.

[0099] In the example of [Fig.l], the inference module 30 is produced in the form of a programmable logic component, such as an FPGA (Field Programmable Gate Array) or in the form of a dedicated integrated circuit, such as an ASIC (Application Specific Integrated Circuit).

[0100] As a variant, not shown, the computer 20 comprises both the learning module 25 and the inference module 30. According to this variant, the learning module 25 and the inference module 30 are each produced in the form of a programmable logic component, such as an FPGA, or in the form of a dedicated integrated circuit, such as an ASIC. According to this variant, the computer 20 is then configured to perform both the learning and the inference of the neural network RN.

[0101] The learning module 25 is configured to carry out the learning of the RN neural network, in particular the synaptic weight values ​​- also called synaptic weights according to shorter terminology - of at least one CTI, CT2, CT3 layer of the RN neural network from learning data, and preferably of each CTI, CT2, CT3 layer of said RN neural network.

[0102] For example, the learning module 25 is configured to perform a gradient back-propagation algorithm.

[0103] The learning module 25 is further configured to perform the quantification of said synaptic weights by a function, called quantifier, each synaptic weight resulting from said learning being a quantified value belonging to a set EQ of quantified values.

[0104] For example, the quantifier is defined such that the quantized value of a weight Wij, the quantized value being denoted verifies the following equation:

[0105] [1]

[0106] = Quantize(| x + )

[0107] where Wij is the j-th weight of the i-th CTi processing layer,

[0108] is the j-th quantized weight of the i-th CTi processing layer, and

[0109] maxr,s| represents the maximum value among the absolute values ​​of the synaptic weights.

[0110] The quantizer according to equation [1] is the one used in the S AT method, described in the article “Towards Efficient Training for Neural Network Quantization” by Q. Jin et al.

[0111] In the remainder of the description, unless otherwise specified, the term “weight” means a synaptic weight value and the term “quantified weight” means a quantified synaptic weight value.

[0112] The neural network RN comprises a plurality of groups W; of quantized weights of the network, each group Wi comprising at least one quantized weight belonging to the set EQ. A group W; is for example a layer CTi of the neural network RN.

[0113] In the remainder of the description, the notation W corresponds to any group W;. In other words, the index i is not specified when the characteristic described is applicable to each of the groups Wi.

[0114] In each group W, the distribution of quantified weights follows a discrete random variable LX "| with empirical probability law.

[0115] Each group W of quantized weights is associated with a precision parameter Xw equal to the number of bits used for the coding of at least one quantized weight of the respective group W.

[0116] The training data comprises a set of example data used during the training process. The data comprises a plurality of inputs Xi, each input Xi being associated with a label Yi, which is for example a binary number or a real number. An input Xi is for example a vector comprising n values ​​corresponding to known characteristics.

[0117] The aim of learning is to determine synaptic weights which make it possible to best predict the value of the labels Yi from the corresponding input values ​​Xi.

[0118] For this purpose, the learning module 25 is configured to determine the weights by minimizing a cost function L. The cost function L depends on an error term Lcls corresponding to a prediction error and an entropic term.

[0119] The learning module 25 is further configured to calculate the gradient of the cost function L.

[0120] For example, the cost function L is the sum of the error term Lcls corresponding to a prediction error and the entropic term.

[0121] The error term Lcls measures the deviation between predictions Y'i of the neural network RN from the determined weights and inputs Xi and the true value of labels Yi of the training data.

[0122] The error term Lcls is for example the cross entropy between the predictions Y'i and the labels Yi.

[0123] The entropic term depends on a total differential entropy Htot of the quantized weights.

[0124] The entropic term is for example the product of a multiplicative factor K and an entropic objective function Lentropy, the entropic objective function Lentropy depending on the total differential entropy Htot of the quantized weights.

[0125] The cost function L then typically verifies the following equation:

[0126] [2]

[0127] L = Lcls+KLentr°l,y

[0128] where L represents the cost function,

[0129] Lcls represents the error term,

[0130] k represents the multiplicative factor, and

[0131] Lentropy represents |the entropic objective function.

[0132] For example, the K factor takes a fixed value throughout the training, preferably substantially equal to 0.1.

[0133] Alternatively, the K factor takes a progressive value that increases linearly during a first part of the training up to a threshold value, then keeps the threshold value during the rest of the training. The linear growth is for example configured to take place for the first twenty learning iterations, the learning iterations also being called epochs. The linear growth is for example configured to take place for the first twenty epochs of the backpropagation algorithm.

[0134] The entropic objective function Lentropy preferably further depends on a predefined total entropic objective H.

[0135] In particular, the entropic objective function Lentropy depends on a difference between the total differential entropy Htot and the total entropic objective H.

[0136] According to an exemplary embodiment, the precision parameters Xw are integers. The total differential entropy Htot is then a sum of products for the groups W, each product being a product of a cardinal IWI of a respective group W and the differential entropy Hw of said group.

[0137] The total differential entropy Htot then verifies the following equation:

[0138] [3]

[0139] H u , = Y. w ÏWiH w

[0140] where Htot is the total differential entropy of the quantized weights,

[0141] IWI is the cardinal of the group W, i.e. the number of weights included in the group W,

[0142] Hw is the differential entropy of the group W.

[0143] For example, the entropic objective function Lentropy satisfies the equation:

[0144] [4]

[0145] entropy _ A ~ H

[0146] where Lentropy is the entropic objective function,

[0147] | W | is the cardinal of the group W,

[0148] Hw is the differential entropy of the group W, and

[0149] H is the total entropic objective.

[0150] The differential entropy Hw extends the concept of Shannon entropy to continuous probability laws. The differential entropy Hw is for example the differential binary entropy of the weights of the group W.

[0151] The empirical probability law of the discrete random variable L of the quantified weights for each group W being discrete, the differential entropy Hw therefore corresponds to an approximation of the Shannon entropy of the discrete random variable of the quantified weights for each group W. The demonstration of the validity of this approximation will be described later.

[0152] The discrete random variable L of the quantized weights actually corresponds to the rounding of a continuous random variable X which follows a continuous probability law. This continuous probability law is used for the calculation of the differential entropy Hw-

[0153] In the following description, the term “entropy” means Shannon entropy.

[0154] In the particular case where the weights of each group W follow a discrete random variable L corresponding to the rounding of the continuous random variable X following a normal law of variance V(X) and expectation E(X), the learning module 25 is configured to calculate the derivative with respect to the weight(s) of each respective group W of the entropic objective function Lentropy, as defined by equation [4], for example according to the following equation:

[0155] [5]

[0156] rHv 2Zn(2)V(W)â

[0157] ...lltro„ dA r«V <w “ 2ln(2W(W)H

[0158] _ W-Æ(W) nln(l)V(W)H

[0159] where Lentropy is the entropic objective function,

[0160] V(W) is the variance of the second continuous random variable X of the respective group W,

[0161] E(W) is the expectation of the second continuous random variable X of the respective group W,

[0162] H is the total entropic objective, and

[0163] n is the number of quantized weights of the group W.

[0164] Alternatively, the precision parameters Xw are positive real numbers. The function entropic objective Lentropy then further depends on the precision parameters Xw.

[0165] In particular, the entropic objective function Lentropy verifies the following equation:

[0166] [6]

[0167] entropy l-Xw-f]À jJ) W - jjÀ vJÀ njTî]

[0168] where Lentropy is the entropic objective function,

[0169] Xwest is the precision parameter of the respective W group,

[0170] Hestl'objectif total entropy,

[0171] Hw(k) is the entropy of the weights quantized to precision k,

[0172] [J represents the whole part, also called the lower whole part (from the English floor, or parquet in French), and

[0173] LJ represents the upper whole part (from the English ceiling, or plafond in French).

[0174] This variant makes it possible to provide the optimal quantification precisions per group W minimizing the entropy of the weights.

[0175] Advantageously, the learning module 25 is further configured to calculate the derivative with respect to the precision parameter Xw of the entropic objective function Lentr°py5 for example according to the following equation:

[0176] [7]

[0177]

[0178] where Lentropy represents the entropic objective function,

[0179] Xw represents the precision parameter of the respective group,

[0180] W represents the respective group of weights,

[0181] Hw represents the differential entropy of the group W,

[0182] H represents the total entropic objective,

[0183] LJ represents the lower integer part, and

[0184] LJ represents the upper whole part.

[0185] According to an optional addition, the learning module 25 is configured to carry out the learning of the weights from learning data, each weight resulting from said learning being a quantified weight value belonging to the set EQ of quantified values, the set EQ including an odd number of values.

[0186] As a variant of this optional complement, with reference to [Fig.2], the set EQ includes the zero value.

[0187] In the example of [Fig.2], the distribution 50 of the quantized values ​​is such that the positive quantized values ​​55 of the set EQ are symmetrical to the negative quantized values ​​60 with respect to the zero value.

[0188] The null value captures a very large portion of the weights due to the natural distribution of weights. Thus, adding the null value helps reduce the entropy of the quantized weights. The symmetry with respect to the null value allows for a distribution close to a rounded normal law.

[0189] According to a second aspect complementary and independent of the first aspect, the electronic system further comprises a compression module 32, illustrated in [Fig.l], configured to compress the quantified weights determined by the learning module 25.

[0190] In the example of [Fig.l], the compression module 32 is external to the computer 20 and to the learning module 25.

[0191] The compression module 32 is for example produced in the form of software, that is to say in the form of a computer program. It is furthermore capable of being recorded on a medium, not shown, readable by a computer. The computer-readable medium is for example a medium capable of storing electronic instructions and of being coupled to a bus of a computer system. By way of example, the readable medium is an optical disk, a magneto-optical disk, a ROM memory, a RAM memory, any type of non-volatile memory (for example EPROM, EEPROM, FLASH, NVRAM), a magnetic card or an optical card. A computer program comprising software instructions is then stored on the readable medium.

[0192] Alternatively, the compression module 32 is produced in the form of a programmable logic component, such as an FPGA (Field Programmable Gate Array) or in the form of a dedicated integrated circuit, such as an ASIC (Application Specific Integrated Circuit).

[0193] As a variant, not shown, the compression module 32 is included in the computer 20 and / or in the learning module 25.

[0194] For example, the compression module 32 is configured to perform entropy coding of the quantized weights.

[0195] The compression module 32 is configured to provide as output a bit string 68 representing the quantized weight(s) of a group W.

[0196] For example, entropy coding is an asymmetric numeric systems coding, also called ANS coding (from the English Asymmetric Numerical Systems), the operation of which is known per se, for example from the article entitled “Asymmetric Numerical Systems: entropy coding combining speed of Huffman coding with compression rate of arithmetic coding” by J. Duda, published in 2014.

[0197] However, such entropic coding has never been used for the compression of parameters, such as synaptic weights, of neural networks.

[0198] According to this second aspect of the invention, the symbols to be encoded are the quantization values ​​of the weights, i.e. the quantized weights.

[0199] Preferably, the entropy coding is a tabled asymmetric numeric systems coding, also called tANS coding (from the English tabled Asymmetric Numerical Systems). The compression module 32 is configured to associate at least one state E with each quantized weight value.

[0200] For example, the number 1 of state(s) E encoding n quantified values ​​verifies the following equation:

[0201] [8]

[0202] / =

[0203] where 1 is the number of state(s) E,

[0204] n the number of quantized values, and

[0205] b is an integer.

[0206] Preferably, the number of state(s) E assigned to a quantized value depends on the probability of this quantized value. In other words, the more a quantized value is determined among a group W, the more states E will be associated with it.

[0207] According to this second aspect of the invention, the inference module 30 comprises a decompression unit 35 configured to decompress the compressed weights. For example, the decompression unit 35 is configured to perform an entropy decoding of the compressed weights.

[0208] The decompression unit 35 takes as input the bit string 68, provided as output from the compression module 32, and an initial state Eb

[0209] In the case where the entropic coding is the tANS coding, the decompression unit 35, illustrated in [Fig.3], comprises a look-up table 70, also called a LUT (Look-Up Table). The look-up table 70 forms a decoding table for the tANS coding.

[0210] In the example of [Fig.3], the decompression unit 35 further comprises an accumulator 72 and an adder 74.

[0211] The lookup table 70 includes 1 lines 76 corresponding to the 1 encoded states E by the compression module 32, each line 76 corresponding to a state E;, and three columns 78. A first column 78i comprises a symbol S, that is to say the quantified weight corresponding to the state E;. A second column 782 comprises the number of bits Nb to be read in the bit string 68 to be decoded according to the state E;. A third column 783 corresponds to an integer X' necessary to calculate the following state Ei+i.

[0212] The look-up table 70 is configured to take as input the state E; and emit as output three signals 80;. A first signal 80; is representative of the symbol S contained in the first column 78; of the state E;. A second signal 802 is representative of the number of bits Nb to be read in the bit string 68, contained in the second column 782. A third signal 803 is representative of the integer X' necessary to calculate the following state Ei+i according to the third column 783.

[0213] The accumulator 72 is configured to transform the second signal 802 into a signal 82 representative of the bits read in the bit string 68, the number of bits read depending on the second signal 802.

[0214] The adder 74 is configured to take as input the third signal 803 and the signal 82, and to then calculate the following state Ei+i.

[0215] As a variant, not shown, the adder 74 is not present. The look-up table 70 then comprises more columns 78. The additional columns 78 directly comprise the different state values ​​following Ei+i. In this case, the memory footprint necessary for coding the look-up table 70 is greater than in the presence of the adder 74.

[0216] An example of operation of the computer 20 according to the invention will now be explained with regard to [Fig.4] representing a flowchart of a method for processing data from the sensor 15, in particular for classifying data, via the implementation of the artificial neural network RN with the computer 20, the method comprising a learning phase 100 in which a method, according to the invention, for learning synaptic weights of at least one layer CTI, CT2, CT3 of the neural network RN is implemented; then an inference phase 150 in which the artificial neural network RN, previously trained, is used to calculate output values, in order to process, in particular to classify, i.e. classify, said data.

[0217] As described above, the learning phase 100 is preferably implemented by a computer, this learning phase 100 being carried out by the learning module 25, and optionally the compression module 32, which are typically software modules. The subsequent inference phase 150 is, for its part, preferably implemented by the computer 20, and more precisely by its inference module 30. In particular, the learning phase 100 according to the invention then allows an implementation of the inference phase 150.

[0218] The learning phase 100 comprises a step 110 of learning the neural network RN, in particular quantified weights of said network.

[0219] The learning step 110 is performed by the learning module 25.

[0220] The learning step 110 is performed from the learning data. This learning is performed via a gradient back-propagation algorithm.

[0221] The backpropagation algorithm aims to converge towards an optimal configuration of the synaptic weights, obtained by minimizing the cost function L.

[0222] According to the invention, the cost function L depends on an error term Lcls corresponding to the prediction error and the entropic term, the entropic term depends on the total differential entropy Htot of the quantized weights.

[0223] The learning step 110 then comprises the quantification of the weights obtained by applying the quantifier. Each determined weight is a quantized value belonging to the predefined set EQ of quantized values.

[0224] During the learning phase 100, at the end of the learning step 110, the compression module 32 optionally then performs a compression step 120 of the quantized weights.

[0225] The compression is for example carried out via entropic coding, preferably via ANS coding. Advantageously, the compression is carried out via tANS coding.

[0226] During the inference phase 150, the inference module 30 infers the artificial neural network RN previously trained to process, in particular classify, the data received as input from the electronic computer 20, the neural network RN having been previously trained and compressed during the learning phase 100.

[0227] The data received by the computer 20 corresponds to the data captured by the sensor 15.

[0228] The inference phase 150 comprises a first step 160 of decompression of the compressed weights, carried out by the decompression unit 35 included in the inference module 30.

[0229] For example, the decompression unit 35 performs entropy decoding, preferably asymmetric digital systems decoding, also called ANS decoding. Advantageously, the entropy decoding performed by the decompression unit 35 is a table asymmetric digital systems decoding, also called tANS decoding, using the look-up table 70.

[0230] At the end of the decompression step 160, the inference module 30 performs a neural calculation step 170.

[0231] In particular, the neural calculation step 170 corresponds to the execution of the weighted sum of input value(s) from the data of the sensor 15. The weighting is done by the quantified weights determined during the learning step 110, which have been optionally decompressed during the decompression step 160.

[0232] Alternatively, subsequently, the inference module 30 applies a weighted activation function to the weighted sum.

[0233] The learning method according to the invention makes it possible to carry out the inference phase 150 with the inference module 30 requiring a reduced memory footprint while maintaining good performance.

[0234] The reduction of the memory footprint is obtained by two main aspects described above, these aspects being complementary and independent. The first aspect is the addition of the entropic term dependent on the total differential entropy Htot, to the cost function L. The second aspect is the compression of the quantized weights obtained during the learning step 110, and advantageously the entropic coding, such as the tANS coding, of said quantized weights.

[0235] The two aspects are complementary, because the addition of the entropic term allows the reduction of the entropy of the quantized weights, which has the effect of increasing the compression potential of the quantized weights.

[0236] The impact of this first aspect on reducing the memory footprint will now be described.

[0237] First, we will justify the use of differential entropy as an approximation of Shannon entropy for the distribution of quantized weight(s) of a given group W.

[0238] Subsequently, the group W considered is for example a CTi processing layer of a neural network RN.

[0239] The quantized weights included in the given group W are defined by the discrete random variable L where X is the continuous random variable that follows a continuous probability law with finite support on the interval [-N, 2V] with N a natural integer, and L . 1 is the rounded function.

[0240] We assume that the continuous random variable X is symmetric and of zero expectation, and consequently the discrete random variable LX ] is symmetric and of zero expectation.

[0241] The continuous random variable X is a random variable with density f(x). The differential binary entropy H(X) of the continuous random variable X typically satisfies the equation:

[0242] [9]

[0243]

[0244]

[0245]

[0246]

[0247]

[0248]

[0249]

[0250]

[0251]

[0252]

[0253]

[0254]

[0255]

[0256]

[0257]

[0258]

[0259]

[0260]

[0261]

[0262]

[0263]

[0264]

[0265] For example, if the continuous random variable X follows a centered normal distribution of variance V(X), the differential binary entropy H(X) verifies the following equation:

[10] ff(X)=41og2(V(X)27œ) By Chasles' relation, equation [9] is equivalent to the following equation: [H] H ( X ) = - L* JJ f ( t) log^f ( t ) Moreover, the entropy of the discrete random variable [ l] verifies the equation:

[12] h(Lxl) = L*! ='W2(r( L^l =i) ) "H(LX1 ) = <X2»+î)log2(p(i-4 cXSid) ) • 1 / • ? \ fï+o / fî+o 1 "LX1) = The discrete random variable LX"| has the same entropy as a continuous random variable X of density g(t), the density g(t) verifying the following equation:

[13] / \ rLd+4 gkl = i^f(u)du In the rest of the description we will indifferently note the entropy [ X~| ) and h(x)- The difference between the entropy h(X) and the differential entropy H(X) is then defined by the following equation:

[14] H(X)-H(X) âA-\ dl 2 (f(ù)du I tf(t} \ ~ff(ï)-H(X)=l w / (Olog,h^ Æ “ \ JLfl4 J'uxht J Since equation

[14] corresponds to a Kullback-Leibler divergence, the following equation is then obtained:

[15] h(x) - h(x) = d k £x I x)

[0266]

[0267]

[0268]

[0269]

[0270]

[0271]

[0272]

[0273]

[0274]

[0275]

[0276]

[0277]

[0278]

[0279]

[0280]

[0281]

[0282] According to the intermediate value theorem, the density f(x) takes at least once the value d+4 on the interval [ ; 1 / .1 ], otherwise the density f would be LïfWdU t' 2. »+2J '*2 always strictly greater or less than its average, which is absurd. Let c; such that p+ÿ, and is tz#x £(O=iog / (t) \. According to the inequality finite increments applied to L(t), for t included in the interval [f - 2 j we obtain the following equation:

[16] Now by definition of Ci, L(c;) = 0. The difference between the entropy h(x) and the differential entropy H(X) is then bounded and verifies the following equation:

[17] ■i+i . v jpi » |h(x) -H(X) I ^-^1* In the above integral, both t and c; belong to the interval _ so K-Ci| - 1. Finally, the difference between the entropy H(X) and the differential entropy H(X) is then bounded and verifies the following equation:

[18] For example, if the continuous random variable X follows a centered normal distribution of variance o2 and cut on [ - M AH, assuming N is large enough so that the probability of the tail of the distribution is negligible, the density f(x) then verifies the following equation:

[19] where o is the standard deviation of the continuous random variable X. The difference between the entropy H(X) and the differential entropy H(X) then verifies the following equation:

[20]

[0283]

[0284]

[0285]

[0286]

[0287]

[0288]

[0289]

[0290]

[0291]

[0292]

[0293] where - I f'(t) I. Now ti' belongs to the interval [ f ~ ï + 4 ] ' so t - 1 < < t + 1- By Chasles' relation, the following equation is obtained:

[21] \H(X) -H(X) I U. jV(-t+ l)f(t)dt+\0 (r+ ï)f(t)dt] । / r°\ “\H(X)-H(X)\<1^[\^-tf(t)dt + \ o tf(t)dt + ]_ N f(ndt] The third integral above is equal to 1 by property of the density f(x). The sum of the first two integrals above is the expectation of a half-normal distribution, that is, of the absolute value of a normal distribution, cut on the interval [0, N]. The expectation of the absolute value of a normal distribution is equal to / 2”, where o is the standard deviation of TT the normal law. Assuming that N is sufficiently large, the effect of the cutting on the interval [0, N] of the half-normal law on the value of the expectation is negligible. Thus, we consider that the sum of the first two integrals is equal to this which is a good approximation. The difference between the entropy L^l) and the differential entropy H(X) then verifies the following equation:

[22] \H(]_X1) -H(X) \ <^(4^ + i) The person skilled in the art will thus observe that if the standard deviation o is sufficiently large, the differential entropy H(X) is a good approximation of the entropy / / ([ xl) ■ The above equation

[22] is consistent with the experimental results obtained after training the Resnet-18 convolutional neural network on the cifarlO training dataset, with the Resnet-18 network being trained to classify data, as shown in Table 1 below. [Tables 1] Quantization Precision Entropy Memory Footprint (in MB) Entropy Target 1 Entropy Target 2 Entropy Target 3 4 bits Discrete Entropy 0.120 0.099 0.062 Differential Entropy 0.12 0.10 0.06 3 bits Discrete Entropy 0.077 0.069 0.063 Differential entropy! e 0.07 0.06 0.05 2 bits Discrete entropy 0.078 0.067 0.059 Differential entropy! e 0.08 0.06 0.05

[0294] In Table 1 above, the memory footprint of the trained Resnet-18 neural network is calculated from the entropy of the discrete random variable [X], also called discrete entropy, and the differential entropy of the continuous random variable X, also called differential entropy. The memory footprint required to store the quantized weights is compared for several quantization precisions, the quantized values ​​being on 2 bits, 3 bits or even 4 bits in Table 1, and for several values ​​of total entropic objective H, called entropic objective in Table 1. The total entropic objectives H are ranked from the least aggressive to the most aggressive, that is to say in decreasing order in terms of memory footprint.

[0295] The person skilled in the art will then observe first of all that the differential entropy H(X) is a good approximation of the entropy xl) ' and then that at the lowest accuracies and when the total entropic objective H is more aggressive, the quality of the approximation by the differential entropy tends to deteriorate, in other words to deviate from its real value calculated with the discrete entropy. This observation is consistent with equation

[22] . Indeed, the lower the quantification precision and the total entropic objective H, the lower the standard deviation o of the continuous probability law, and consequently the higher the limit of the difference between the entropy J / (LX1 ) and the differential entropy H(X), the latter being a function of the inverse of the standard deviation o.

[0296] In a second step, we will highlight the effect of the first aspect of the invention on the reduction of the memory footprint, in order to then minimize the use of memory resources of the computer 20 during the inference of the neural network RN.

[0297] The impact of the entropic objective function Lentropy on the distribution of quantized weights is for example represented in [Fig.5].

[0298] [Fig.5] shows two histograms 200, 250, namely a first histogram 200 and a second histogram 250 represent the distribution of the quantized and 4-bit coded weights of a layer of the Resnet-18 neural network of table 1.

[0299] On the first histogram 200 at the top of [Fig.5], the state-of-the-art learning module determines the quantized weights in the absence of an entropic objective function in the cost function L. The weights are dispersed over 15 values ​​numbered from 0 to 14.

[0300] On the second histogram 250 at the bottom of [Fig.5], the learning module 25 determines the quantified weights in the presence of the entropic objective function Lentropy described above. The weights are dispersed over 6 values ​​numbered from 0 to 5.

[0301] The entropic objective function Lentropy therefore allows a reduction in the variability of the quantified weights, and consequently a reduction in the memory footprint of the Resnet-18 neural network.

[0302] Indeed, the layer of the Resnet-18 neural network trained without an entropic objective function requires a memory footprint of 48.59 kB compared to 28.46 kB for the layer of the Resnet-18 neural network trained with the Lentropy entropic objective function.

[0303] The entropic objective function Lentropy therefore allows a reduction in the memory footprint. In addition, the learning method according to the invention makes it possible to obtain good performance as shown in Table 2 below.

[0304] [Tables2] Quantization accuracy Entropy target 1 Entropy target 2 Entropy target 3 4 bits Memory footprint (in MB) 0.229 0.100 0.085 Accuracy (%) 90.63 89.72 86.98 3 bits Memory footprint (in MB) 0.132 0.077 0.069 Accuracy (%) 90.63 90.19 89.25 2 bits Memory footprint (in MB) 0.088 0.078 0.067 Accuracy (%) 89.63 89.74 89.38

[0305] In Table 2, the performance is evaluated by the classification accuracy of the Resnet-18 neural network in Table 1. Here, the classification accuracy, called Precision in Table 2, is measured by an accuracy metric known as Top-1 accuracy, i.e., the accuracy is the proportion of training data for which the prediction Yi' is equal to the corresponding label Yi.

[0306] The skilled person will then observe that the classification accuracy is close to 90% in most cases, thus the entropic term allows good performance to be obtained. Furthermore, the skilled person will observe that for a quantization accuracy of 2 bits or 3 bits, the total entropic objective H has little influence on classification accuracy, while the total entropy objective H has a lot of influence on the memory footprint.

[0307] The advantage of the optional addition of the invention is also investigated, according to which the EQ set of quantized values ​​includes an odd number of quantized values. The classification accuracy and memory footprint of the Resnet-18 neural network of Table 1 whose quantized weights are included in such an EQ set are shown in Table 3 below.

[0308] [Tables3] EQ Size Entropic Objective 1 Entropic Objective 2 Entropic Objective 3 5 Memory Footprint (in MB) 0.07 0.051 0.043 Accuracy (%) 90.21 87.8 82.78 3 Memory Footprint (in MB) 0.008 Accuracy (%) 81.75

[0309] Table 3 shows that, for an EQ set including 5 values ​​and the entropic objective 1, a classification accuracy is obtained higher than that obtained for a quantization accuracy of 2 bits, i.e. an EQ set including 4 values, while the memory footprint is lower. An odd number of quantized values ​​therefore makes it possible to reduce the memory footprint of the Resnet-18 neural network, while gaining in performance.

[0310] Table 3 further shows that a ternary quantization, in other words an EQ set including 3 values, results in a significant drop in classification accuracy. Those skilled in the art will nevertheless observe that the ternary quantization, with a memory footprint equal to 0.008 MB in this example, results in a decrease of approximately 90% compared to the memory footprint for a quantization accuracy of 2 bits, this being then equal to 0.088 MB in the example of Table 2.

[0311] The results of Table 2 and Table 3 are represented on curves 300, 305, 310 and 315 in [Fig.6].

[0312] Each curve 300, 305, 310, 315 corresponds to a distinct precision parameter, or number of values ​​included in the EQ set, for which the total memory footprint of the Resnet-18 neural network is represented as a function of the performance of the Resnet-18 neural network, in particular the Top-1 precision metric.

[0313] Thus, a first curve 300 corresponds to the results obtained for quantized weights coded with a precision parameter equal to 4 bits, in other words the EQ set includes 16 quantization values.

[0314] A second curve 305 corresponds to the results obtained for quantized weights coded with a precision parameter equal to 3 bits, in other words the EQ set includes 8 quantization values.

[0315] A third curve 310 corresponds to the results obtained for quantized weights coded with a precision parameter equal to 2 bits, in other words the EQ set includes 4 quantization values.

[0316] A fourth curve 315 corresponds to the results obtained for quantified weights belonging to the set EQ including 5 quantification values.

[0317] An isolated point 320 corresponds to the result obtained for quantified weights belonging to the set EQ including 3 quantification values.

[0318] Each curve 300, 305, 310, 315 has an inflection 330 corresponding to the Pareto frontier, i.e. the best compromise between performance and memory footprint. Below the inflection 330, a small decrease in memory footprint results in a large reduction in classification accuracy. Above the inflection 330, a large increase in memory footprint results in a small increase in classification accuracy.

[0319] [Fig.6] also highlights the effect of the optional complement according to which the EQ set includes an odd number of values. Indeed, for the Resnet-18 neural network whose quantized weights take 5 distinct values, with reference to the fourth curve 315, we note that for performances similar to the Resnet-18 neural networks whose quantized weights are coded by 4 bits, 3 bits, and respectively 2 bits, with reference to the first 300, second 305, and respectively third 310 curves, the memory footprint is lower.

[0320] The impact of the second aspect of the invention, i.e. the compression of the quantized weights on the reduction of the memory footprint will now be described.

[0321] The results described below take into account only this second aspect. In other words, the RN neural networks described below were trained by a learning module configured to minimize a cost function without an entropic objective function Lentropy.

[0322] Compression is preferably performed by tANS coding. tANS coding has the advantage of being easily decodable by inference due to the presence of the look-up table 70 included in the decompression unit 35.

[0323] In addition, tANS coding and tANS decoding make it possible to greatly reduce the memory footprint required to store the quantized weights.

[0324] For example, the histogram 350 of [Fig.7] represents the number of distinct quantized values ​​of each CTi processing layer of the RN neural network, the RN neural network being a Resnet-50 network trained on the data ImageNet training. The Resnet-50 neural network includes 52 CTi processing layers.

[0325] Histogram 350 suggests that the quantization precision Xw, where each group W is a CTi processing layer, is optimized depending on the layer concerned.

[0326] Further, each CTi processing layer has an odd number of distinct quantized values ​​included in a respective set EQi of quantized values.

[0327] Preferably, each set EQi includes the zero value.

[0328] Preferably, the positive quantized values ​​of each set EQi are symmetrical to the negative quantized values ​​of each set EQi with respect to the zero value.

[0329] On Thistogram 350, the maximum number of distinct values ​​for a layer reaches 31 values; and the minimum number is 13 values.

[0330] The two histograms 400, 450 of [Fig.8], namely a first histogram 400 and a second histogram 450, represent the memory footprint necessary to store the quantized weights of the Resnet-50 neural network of [Fig.7].

[0331] In the first histogram 400 at the top of [Fig.8], the memory footprint in MB required to store the quantized weights of each CTi layer before compression is shown. Some layers reach a memory footprint of 1.3 MB.

[0332] In the second histogram 450 at the bottom of [Fig.8], the memory footprint in MB required to store the quantized weights of each CTi layer after compression by tANS coding is shown. The maximum memory footprint achieved is 185 KB, so some CTi layers have been compressed up to 20 times.

[0333] Thus, the reduction of the memory footprint is done by keeping the RN neural network with good performance even after decoding, as shown in Table 4 below.

[0334] [Tables4] RN Compression Accuracy (in %) Memory footprint (in MB) 1 State of the art 1 68.8 3.6 State of the art 2 70.0 2.78 State of the art 3 68.5 2.01 tANS coding 69.0 1.67 2 State of the art 1 71.3 5.5 State of the art 2 73.5 5.91 State of the art 4 73.7 5.25 tANS ​​coding 75.6 4.52

[0335] The RN neural networks in Table 4 are trained such that the quantized weights are included in an EQ set including the zero value.

[0336] The RN 1 neural network in Table 4 corresponds to a Resnet-18 network trained on ImageNet training data. Before compression, the Resnet-18 network requires a memory footprint of 46.8 MB and achieves a classification accuracy of 69.8%.

[0337] The RN 2 neural network in Table 4 corresponds to a Resnet-50 network trained on ImageNet training data. Before compression, the Resnet-50 network requires a memory footprint of 102.5 MB and achieves a classification accuracy of 76.1%.

[0338] State of the art 1 corresponds to the values ​​obtained by the LZMA compression method, notably used for weight compression in the article “HEMP: High-order Entropy Minimization for neural network comPression” by Tartanglione et al. in 2021.

[0339] State of the art 2 corresponds to the values ​​obtained by the method described in the article “Scalable model compression by entropy penalized reparameterization” by Oktay et al. in 2020. This method includes training of the decoder, in addition to training of the network parameters.

[0340] State of the art 3 corresponds to the FracBits method, applied to the weights only and not to the activations, described in the article “FracBits: Mixed Precision Quantization via Fractional Bit-Widths” by L.Yang and Q. Jin in 2020.

[0341] State of the art 4 corresponds to the Golomb coding method coupled with arithmetic coding with context, described in the article “DeepCABAC: A Universal Compression Algorithm for Deep Neural Networks” by Wiedemann et al. in 2020.

[0342] For both RN 1 and RN 2 neural networks, Table 4 shows a significant reduction in the memory footprint required to store all the quantized weights of each layer, when they are compressed via tANS coding compared to the various methods of the state of the art. The Top-1 accuracy of the RN 1 and RN 2 neural networks compressed by tANS coding is substantially equal to the accuracy of the neural networks of the state of the art or even higher for the RN 2 neural network. Those skilled in the art will then observe that compression via tANS coding allows this significant reduction in the memory footprint, while not altering the accuracy obtained.

[0343] Thanks to the two aspects of the learning method according to the invention, it is thus possible to train the neural network RN configured to be implemented by the electronic computer 20, while having a memory footprint necessary for the reduced operation of the RN network, while maintaining the performance of the RN network, in particular its classification accuracy. Thus, the invention makes it possible to have smaller on-board electronic computers 20 to carry out data processing, in particular data classification, via the implementation of the artificial neural network RN.

[0344] It is thus understood that the learning method according to the invention makes it possible to reduce the memory footprint of the RN network without loss of data processing efficiency and to facilitate the inference of said RN neural network by the computer 20.

Claims

Claims

1. Method for learning synaptic weights of at least one layer (CTI, CT2, CT3) of an artificial neural network (RN) configured to process, in particular to classify, data, the artificial neural network (RN) being configured to be implemented by an electronic computer (20) connected to a sensor (15), for the processing of at least one object from the sensor (15), each artificial neuron of a respective layer (CTI, CT2, CT3) being capable of performing a weighted sum of input value(s), then applying an activation function to the weighted sum to deliver an output value, each input value being received from a respective element connected to the input of said neuron and multiplied by a synaptic weight associated with the connection between said neuron and the respective element, the respective element being an input variable of the neural network (RN) or a neuron of a previous layer of the neural network (RN),the method being implemented by computer and comprising the following step: - determining (110) the weights of the neural network (RN) from learning data, each determined weight being a quantized value belonging to a predefined set (EQ) of quantized values; the weights being determined by minimizing a cost function (L), the cost function (L) depending on an error term (Lcls) corresponding to a prediction error and an entropic term, characterized in that the entropic term depends on a total differential entropy (Htot) of the quantized weights.,

2. The method of claim 1, wherein the neural network (RN) comprises groups (W) of weights of the network, each group (W) of weights comprising at least one quantized weight belonging to the set (EQ), and the total differential entropy (Htot) is a sum of products for said groups, each product being a product of a cardinality (IWI) of a respective group and the differential entropy (H(W)) of said group.

3. A method according to claim 1 or 2, wherein the cost function (L) is the sum of the error term (Lcls) and the entropic term; the entropic term preferably being the product of a multiplicative factor (k) and an entropic objective function (Lentr°py), the entropic objective function (Lentropy) depending on the total differential entropy (Htot) of the quantized weights; the cost function (L) preferably still satisfying the equation: L = Lds + KLentr°py where L represents the cost function, Lcls represents the error term, k represents the multiplicative factor, and Lentropy represents the entropic objective function.

4. Method according to claim 2 or 3, wherein the entropic objective function (Lentropy) further depends on a predefined total entropic objective (H).

5. The method of claim 4, wherein the entropic objective function (Lentropy) depends on a difference between the total differential entropy (Htot) and the total entropic objective (H).

6. A method according to any preceding claim, wherein minimizing the cost function (L) comprises calculating a gradient of the cost function (L).

7. Method according to claim 6, wherein each group (W) of weights is associated with a precision parameter (Xw) equal to the number of bits used for coding the at least one weight of the group (W), and when each precision parameter (Xw) is an integer and when the distribution of the quantized weights of each group (W) follows a rounded normal distribution, the derivative - with respect to the weights of a respective group (W) - of the entropic objective function (Lentropy) satisfies the following equation: _ wg(W) dW ~ nln(2)V(yV)H where Lentropy represents the entropic objective function, W represents the respective group of weights, V(W) represents the variance of the normal distribution, E(W) represents the expectation of the normal distribution, H represents the total entropic objective, and n is the number of quantized weights of the group W.

8. Method according to claim 6, in which each group (W) of weights is associated with a precision parameter (Xw) equal to the number of bits used for the coding of the at least one weight of the group (W), and when each precision parameter (Xw) is a positive real, the entropic objective function (Lentropy) depending on the precision parameters (Xw), the derivative - with respect to the precision parameter (Xw ) of a respective group (W) - of the entropic objective function (L entr°py) verifies the equation: - H \H MM 'H MM / where Lentropy represents the entropic objective function, Xw represents the precision parameter of the respective group, W represents the respective group of weights, Hw represents the differential entropy of the group W, H represents the total entropic objective, LJ represents the lower integer part, and [J represents the upper integer part.

9. A method according to any preceding claim, wherein the predefined set (EQ) of quantized values ​​includes an odd number of values.

10. A method according to any preceding claim, wherein the set (EQ) includes the zero value; the positive quantized values ​​of the set (EQ) preferably being symmetrical, negative quantized values ​​of the set (EQ) with respect to the zero value.

11. Method according to any one of the preceding claims, in which the method further comprises, after the step (110) of determining the weights, a step (120) of compressing the determined weights.

12. The method of claim 11, wherein the compression (120) is performed via entropy coding.

13. The method of claim 12, wherein the entropy coding is asymmetric digital systems (ANS) coding; the entropy coding preferably being table asymmetric digital systems (tANS) coding.

14. Method for processing data, in particular for classifying data, the method being implemented by an electronic computer (20) implementing an artificial neural network (RN), the method comprising: - a learning phase (100) of the artificial neural network (RN), and - an inference phase (150) of the artificial neural network (RN), during which data, received as input from the electronic computer (20), are processed, in particular classified, via the artificial neural network (RN), previously trained during the learning phase (100), characterized in that the learning phase (100) is carried out by implementing a learning method according to any one of claims 1 to 13.

15. A computer program comprising software instructions which, when executed by a computer, implement a learning method according to any one of claims 1 to 13.

16. Electronic computer (20) for processing data, in particular for classifying data, via the implementation of an artificial neural network (RN), the electronic computer (20) being intended to be connected to a sensor (15), for the processing of at least one object from the sensor (15), each artificial neuron of a respective layer (CTI, CT2, CT3) of the neural network (RN) being capable of performing a weighted sum of input value(s), then applying an activation function to the weighted sum to deliver an output value, each input value being received from a respective element connected to the input of said neuron and multiplied by a synaptic weight associated with the connection between said neuron and the respective element, the respective element being an input variable of the neural network (RN) or a neuron of a previous layer of the neural network,the calculator (20) comprising: - an inference module (30) configured to infer the previously trained artificial neural network (RN), for the processing, in particular the classification, of data received as input from the electronic calculator (20), characterized in that the artificial neural network (RN) previously trained during a learning phase (100) comes from a computer program according to claim 15.,

17. A calculator (20) according to claim 16, wherein the inference module (30) is configured to perform the weighted sum of input value(s), then to apply an activation function to the weighted sum.

18. A calculator (20) according to claim 16 or 17, wherein, during the learning phase (100) the weights of the artificial neural network (RN) are compressed, and the inference module (30) further comprises a decompression unit (35) configured to decompress said compressed weights; the compression preferably being carried out via entropic coding, and the decompression unit (35) then carrying out entropic decoding of the compressed weights.

19. A computer (20) according to claim 18, wherein the entropy coding is a table asymmetric digital systems (tANS) coding, and the decompression unit (35) comprises a decoding table (70) of the table asymmetric digital systems (tANS) coding.

20. Electronic system for processing object(s), comprising a sensor (15) and an electronic computer (20) connected to the sensor (15), the computer (20) being configured to process at least one object from the sensor (15), characterized in that the computer (20) is according to any one of claims 16 to 19.