Multifunctional electronic component, device, methods and corresponding programs.

A multifunctional electronic component and platform enable efficient execution and flexible reconfiguration of neural networks on embedded systems, addressing the challenges of size, power, and real-time constraints in current AI implementation paradigms.

FR3145998B1Active Publication Date: 2025-06-27INST SUPÉRIEUR DE LAÉRONAUTIQUE & DE LESPACE (ISAE)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
FR2023001627
Authority / Receiving Office
FR · FR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-02-22
Publication Date
2025-06-27
Estimated Expiration
2043-02-22

AI Technical Summary

Technical Problem

Current paradigms for implementing and training neural network artificial intelligence algorithms are challenging to transpose onto embedded systems due to size, power consumption, and real-time constraints, and existing solutions are difficult to modify or reconfigure for different tasks.

Method used

The development of a multifunctional electronic component and platform that enables the serialization and parallelization of calculations, allowing for the on-the-fly configuration and reconfiguration of neural networks on embedded architectures, using a programmable logic component and orchestration software.

Benefits of technology

This solution enables efficient execution of artificial intelligence algorithms on embedded systems, allowing for flexible reconfiguration and improved performance in terms of power consumption and real-time capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000029_0000
    Figure 00000029_0000
  • Figure 00000030_0000
    Figure 00000030_0000
  • Figure 00000031_0000
    Figure 00000031_0000
Patent Text Reader

Abstract

Multifunctional electronic component, device, method and corresponding programs. The invention relates to an electronic data processing platform, comprising a general-purpose processor (PS) and a so-called level 1 memory (L1). This platform further comprises a programmable logic component (PL) comprising a so-called level 0 memory (L0), the programmable logic component being configured so as to define typologies of executable calculations, said platform further comprising orchestration software running on the general-purpose processor, the orchestration software being responsible for providing, to the programmable logic component, a set of configuration data for its execution within the limits of the typologies of executable calculations. Abstract figure: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Multifunctional electronic component, device, methods and corresponding programs. Technical field

[0001] The invention relates to the implementation of multifunctional electronic components. The invention relates more particularly to the implementation of multifunctional electronic components that can be reconfigured on the fly according to requirements. The invention more naturally finds an application in the implementation of neural networks on embedded architectures. Prior art

[0002] An artificial neural network (or ANN) can be thought of as a complex system combining and interconnecting a large number of non-linear processing units called neurons. Some artificial intelligence algorithms are well known for supervised learning, in which an artificial neural network is used to make a prediction from an input layer receiving input data to an output layer, as shown in [Fig.l]. There are many types and variants of neural network, but they can be classified into three main types: - The MultiLayer Perceptron (or MLP) is the extended version of the original perceptron that contains three or more layers. These layers are also called Fully Connected (or FC) layers and, although they are not used as stand-alone entities, they are usually integrated as subcomponents of other architectures (such as in CNNs). - Convolutional neural networks (or CNNs) are neural networks specialized primarily in image processing applications. They are notably based on convolution operations (matrix operations) commonly used in image processing (such as the Sobel filter). - Recurrent neural networks (RNNs) contain internal memory entities that allow for a feedback loop. This principle allows for a temporal relationship to be taken into account in the processed data. Therefore, the output of a cell is not only influenced by its input but also by previous calculations, which allows for capturing a certain temporal relationship inherent in the data.

[0003] The inventors identified that current paradigms for implementing and training neural network artificial intelligence algorithms are mainly based on centralized architectures with heavy software infrastructures such as server farms, usually GPU-based and even containing some FPGA / ASIC computing accelerators. The training of execution parameters is done online and once trained, the network (and its parameters) is transferred to an inference target for a dedicated task or mission.

[0004] Based on this analysis, the inventors found that: a. Artificial intelligence algorithms (especially those involving large networks) are difficult to transpose onto embedded systems that are highly constrained in terms of size, power consumption and real-time constraints. b. All current embedded approaches are based on online learning - offline execution explained above. The training of neural networks (i.e. obtaining its operating parameters), which is the part that consumes the most time, memory and computing resources, is not performed directly on embedded systems. The networks are trained online, their data is stored, and then they are implemented as is on the destination electronic board. c. Once a neural network has been designed, trained and optimally integrated into an embedded device, it is difficult to modify it (for another task, for a new configuration, a more efficient architecture, ...). Especially when using FPGAs, this reconfiguration / re-implementation of the neural network can be a long and complex task.

[0005] In other words, the inventors have found that the implementation of a neural network in an embedded architecture suffers from numerous problems, the first of which is the fixed nature of the network thus implemented. It is certainly possible, even in an FPGA, to modify the weights of the neural network, in order to take into account improvements in learning (performed online), but it is difficult to envisage modifying the functionalities of the neural network thus obtained. Among the non-modifiable functionalities, the enlargement of the size of the neural network poses problems that are often insoluble. In addition, embedded neural networks suffer from limitations in learning, which is incompatible with the requirements of low power consumption and the limited computing power of these embedded systems.

[0006] There is therefore a need to provide data processing solutions for neural networks on embedded architectures, which can to resolve the problems posed by the prior art. Summary of the invention

[0007] The present technique makes it possible to propose a solution aimed at remedying certain drawbacks of the prior art.

[0008] According to a first aspect, the invention describes the manner in which the inventors decided to formalize the operations carried out within the framework of a series of calculations. More particularly, the invention relates to the serialization and parallelization of calculations carried out on input data to produce output data. The series of calculations in question may be the implementation of one or more neural networks, or any other digital calculations carried out from input data to produce output data.

[0009] According to a second aspect, from the modeling of the operations carried out within the framework of a series of calculations, the inventors have defined a methodology for creating electronic components capable of implementing the series of calculations in question in a generic manner.

[0010] According to a third aspect, once the electronic component is materialized (i.e. for example created or melted), the invention relates to a method of configuring and using this electronic component in order to execute one or more series of calculations modeled according to the methodology defined by the inventors.

[0011] Thus, the present technique relates in particular to an electronic data processing platform, comprising a general-purpose processor and a so-called level 1 memory, said platform further comprising a programmable logic component comprising a so-called level 0 memory, the programmable logic component being configured so as to define typologies of executable calculations. Said platform further comprises orchestration software running on the general-purpose processor, the orchestration software being responsible for providing, to the programmable logic component, a set of configuration data for its execution within the limits of the typologies of executable calculations.

[0012] In a particular embodiment, the programmable logic component is an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).

[0013] In a particular embodiment, the programmable logic component is configured to perform inference calculations for artificial intelligence.

[0014] In a particular embodiment, the programmable logic component is configured to perform learning calculations for artificial intelligence.

[0015] In a particular embodiment, the orchestration software is arranged to provide the execution component with a level 1 memory address at which the configuration data set is accessible, and to provide the programmable logic component with an execution command.

[0016] In a particular embodiment, the programmable logic component is further configured to extract the data structures and functional characteristics of the calculations to be executed and configured to determine how to read and write data in level 1 (L1) memory and determine an execution logic to be implemented.

[0017] In a particular embodiment, the execution logic to be implemented belongs to the group comprising:

[0018] - a multi-layer perceptron type neural network;

[0019] - a convolutional type neural network;

[0020] - a recurrent neural network.

[0021] In a particular embodiment, the programmable logic component is configured to copy, within its level 0 memory, said configuration data set, and to perform the calculations required by the configuration data of the configuration data set.

[0022] In a particular embodiment, the programmable logic component is configured to recursively load, into the level 0 memory, input data from the level 1 memory, then to calculate, from said input data, output data and to copy said output data into the level 1 memory.

[0023] The proposed technique also aims at a method for obtaining an electronic platform as described previously in any one of its different embodiments, in which the creation of the programmable logic component comprises the following steps:

[0024] - obtaining a parameterization file for the programmable logic component, comprising, for at least one characteristic of said programmable logic component, a value representative of the capacity of the programmable logic component to have this characteristic;

[0025] - compilation of the parameterization file of the programmable logic component, delivering on the one hand a binary file for implementing the programmable logic component, and on the other hand an executable file for orchestrating the operation of the programmable logic component;

[0026] - the creation of the programmable logic component based on said binary file;

[0027] - loading said executable file by the general processor.

[0028] According to another aspect, the proposed technique also relates to a method of using of the programmable logic component of an electronic platform as described previously in any of its different embodiments, this method comprising a step of configuring the programmable logic component using configuration parameters, and at least one iteration of the following steps:

[0029] - loading a set of execution parameters;

[0030] - loading a set of execution variables;

[0031] - the calculation of at least one execution result as a function of the set of pa execution parameters and execution variables;

[0032] - recording, within a memory area of ​​the pro logic component grammable, of said at least one execution result.

[0033] Consequently, the present technique also relates to programs, capable of being executed by a computer or by a data processor, these programs comprising instructions for controlling the execution of the steps of the methods as mentioned above.

[0034] A program may use any programming language, and be in the form of source code, object code, or code intermediate between source code and object code, such as in a partially compiled form, or in any other desirable form.

[0035] The present technique also aims at an information medium readable by a data processor, and comprising instructions of a program as mentioned above.

[0036] The information medium may be any entity or terminal capable of storing the program. For example, the medium may comprise a storage means, such as a ROM, for example a CD ROM or a microelectronic circuit ROM, or a magnetic recording means, for example a mobile medium (memory card) or a hard disk or an SSD.

[0037] On the other hand, the information carrier may be a transmissible carrier such as an electrical or optical signal, which may be conveyed via an electrical or optical cable, by radio or by other means. The program according to the present technique may in particular be downloaded over an Internet-type network.

[0038] Alternatively, the information carrier may be an integrated circuit in which the program is incorporated, the circuit being adapted to execute or to be used in the execution of the method in question.

[0039] According to one embodiment, the present technique is implemented by means of software and / or hardware components. In this regard, the term "module" may correspond in this document to a software component, a hardware component or a set of hardware and software components.

[0040] A software component corresponds to one or more computer programs, one or several subroutines of a program, or more generally to any element of a program or software capable of implementing a function or a set of functions, as described below for the module concerned. Such a software component is executed by a data processor of a physical entity (terminal, server, gateway, set-top-box, router, etc.) and is likely to access the hardware resources of this physical entity (memories, recording media, communication buses, electronic input / output cards, user interfaces, etc.).

[0041] Similarly, a hardware component corresponds to any element of a hardware assembly capable of implementing a function or a set of functions, as described below for the module concerned. It may be a programmable hardware component or one with an integrated processor for executing software, for example an integrated circuit, a smart card, a memory card, an electronic card for executing firmware, etc.

[0042] Each component of the system described above of course implements its own software modules.

[0043] The different embodiments mentioned above can be combined with each other for the implementation of the present technique. Figures

[0044] Other aims, characteristics and advantages of the invention will appear more clearly on reading the following description, given as a simple illustrative, and non-limiting, example, in relation to the figures, among which:

[0045] [Fig-1] illustrates the general principle of a method for creating a programmable logic component, in a particular embodiment of the proposed technique;

[0046] [Fig.2] illustrates the general principle of a method of using a programmable logic component, in a particular embodiment of the proposed technique;

[0047] [Fig.3] presents an example of a configuration file of a programmable logic component, in a particular embodiment of the proposed technique;

[0048] [Fig.4] schematically describes the result of a parameterization operation of a programmable logic component, in a particular embodiment of the proposed technique;

[0049] [Fig.5] describes an example architecture of a data processing platform in a particular embodiment of the proposed technique;

[0050] [Fig.6] shows examples of interfaces of a programmable logic component having drive capabilities, in a particular embodiment of the proposed technique;

[0051] [Fig.7] shows examples of interfaces of a pro logic component grammable having inference capabilities, in a particular embodiment of the proposed technique;

[0052] [Fig.8] describes an example of interactions between Manager and Executor, in a particular embodiment of the proposed technique;

[0053] [Fig.9] illustrates an example of use of a programmable logic component for the implementation of a multi-layer MLP neural network, in a particular embodiment of the proposed technique;

[0054] [Fig. 10] illustrates an example of use of a programmable logic component for the implementation of a Lenet-5 type CNN neural network, in a particular embodiment of the proposed technique;

[0055] [Fig. 11] illustrates an example of use of a programmable logic component for the implementation of an LSTM type RNN neural network, in a particular embodiment of the proposed technique;

[0056] [Fig. 12] describes the steps of calculating a first output data element by a programmable logic component, in a particular embodiment of the proposed technique;

[0057] [Fig. 13] describes the steps of calculating a second output data element by a programmable logic component, in a particular embodiment of the proposed technique. Detailed description of the invention 1. Reminders of the principles

[0058] As explained above, the invention aims to provide a computing platform, in particular for the execution of artificial intelligence algorithms, which is versatile and energy-efficient. To do this, the inventors have implemented a method for creating a pseudo-generic computing component, this component being associated, within a platform, with a general-purpose processor, to provide a flexible solution for executing artificial intelligence algorithms, in particular. The pseudo-generic component is the result of work to generalize calculations, in particular those carried out in the case of artificial intelligence algorithms, work which is presented subsequently.The general principle of the invention consists in designing a pseudo-generic component (also called a programmable logic component or execution component), this "pseudo-generic" design consisting in defining the calculations that this execution component is able to perform. Once the pseudo-generic component has been parameterized (i.e. once the types of calculations that this component can perform have been delimited), this execution component is configured on the fly (i.e. at the moment when its use is required) to make it perform specific calculations. The confi . On-the-fly configuration of the component consists of transmitting a configuration file. This configuration file includes the parameters (functions, values, coefficients, etc.) necessary for the implementation of the calculations, within the genericity limits determined at the time of design. The organization of the implementation of the pseudo-generic component is carried out by a general-purpose processor, associated, within an execution platform, with this pseudo-generic component. This platform, overall, can be in the form of a SOC (from the English for "system on chip"). The execution platform thus includes a general-purpose processor, with the capacity to access a so-called level 1 (L1) memory. The general-purpose processor also has access to one or more external interfaces, such as network interfaces, USB, UART, SPI, etc.The general-purpose processor executes software responsible for orchestrating the executions of the pseudo-generic component. It is this software that provides the configuration file "on the fly" and the data on which the calculations must be performed. Once this configuration and data have been received, the pseudo-generic component (or programmable logic component) performs the calculations requested by the configuration file. To do this, the programmable logic component jointly implements its own memory (called level 0 memory, or LO) and the platform's level 1 memory. Thus, the programmable logic component also has instructions allowing read and write access to this level 1 memory.

[0059] Thus, in general, on the basis of the generalization of the calculations carried out by the inventors, the method for creating the programmable logic component which is the subject of the invention comprises, as illustrated in [Fig.l] in a particular embodiment: - A step 11 of obtaining a parameterization file of the programmable logic component, comprising, for at least one characteristic of said programmable logic component, a value representative of the capacity of the programmable logic component to have this characteristic; - A compilation step 12 of the parameterization file of the programmable logic component, delivering, on the one hand, a binary file (of the “bitstream” type) for implementing the programmable logic component and, on the other hand, an executable file for orchestrating the operation of the programmable logic component; - A step 13 of creating the programmable logic component based on the binary file (for example, operation of flashing the “bitstream” binary file onto an FPGA, or even creating an ASIC based on the binary file); and - A step 14 of loading the executable file by the processor ge- generalist.

[0060] These steps can lead to the manufacturing (initialization) of an integrated execution platform (of the SOC type, from the English for “System On Chip”) comprising both the general processor, the level 1 memory accessible by the general processor, the programmable logic component from the parameterization file (this component having resources determined from the parameterization file and having a level 0 memory).

[0061] Equipped with this programmable logic component and more generally with this implementation platform, the invention then relates to the manner in which this component is used within the platform, in particular to configure and execute the various artificial intelligence functionalities which are implemented therein. The invention relates in particular to the manner in which the programmable logic component receives the configuration data and executes the artificial intelligence algorithms according to the directives provided by the software running on the general-purpose processor, by performing level 0 and level 1 memory writes and reads.

[0062] The following sections detail the various characteristics of the pseudo-generic component and the platform associated with it. 2. Generalization of calculations

[0063] First, the inventors have defined a formalization of operations carried out within the framework of a series of calculations. The calculations in question may be matrix calculations, but not only. More particularly, it is possible to describe the methodology implemented by the inventors based on the particular example of neural networks. Thus, on the basis of the findings of inefficiency, particularly energy inefficiency, detailed previously, the inventors have defined a generic description of neural networks (a) having online training capabilities, (b) allowing on-the-fly configuration and reconfiguration for any type of neural network, (c) portable on any embedded device but particularly targeting FPGA / ASIC devices and (d) allowing the possibility of integration into a federated or distributed learning architecture.

[0064] The basis of the inventors' approach is the assertion that all neural networks are, in fact, transformations of one multidimensional space (the input dimensions) into another (the output dimensions). Generally speaking, any computation is a transformation of a multidimensional input space into a multidimensional output space. Optionally, these input and output spaces may be one-dimensional (a polynomial transformation with a single variable, for example).

[0065] More specifically, two main global operations are used for any neural network:

[0066] - inference: this is the prediction operation, it is generally the part which runs on the embedded electronic component once the neural network has been configured and trained; this inference can generally be defined as the resolution of an equation from a plurality of input variables;

[0067] - learning: this is the learning operation composed of two phases. The gradient descent in which the goal is to propagate the error between the output (generated by the neural network) and the expected output (label) to all layers and the second phase where the goal is to correct the transformation in order to make it more optimal for the next inference step (to simplify, updates of the coefficients of the transformation).

[0068] These two phases are described below, and it can be seen that they are both equivalent to a multidimensional matrix calculation, comprising a set of input and output parameters (input dimension, output dimensions) and internal parameters (dimensionality of the matrices of the hidden layers, weights - i.e. values ​​- of these matrices, sizes and values ​​of the biases and activation functions selected).

[0069] 2.1 Inference operations (“feedforward” pass)

[0070] For the complete inference phase (i.e. the "feedforward" pass) of a complex neural network, considering that Output = Y, Input = X, and W the internal data (internal to the complete transformation, generally called the weights), it is possible to write that a neural network is a function f(..) from X and W to Y:

[0071] Output = NeuralN etworkilnpiiï^ InternalData) “Y = f(X, W)

[0072] Or more formally,

[0073] (x f'Y with XWG and YG

[0074] X represents a tuple of real numbers in the set Rn (with n dimensions), W represents a tuple of real numbers in the set Rk (with k dimensions) Y a tuple of real numbers in the set Rm (with m dimensions). The complete inference transformation can be decomposed as a combination of N unitary transformations and it is possible to write: 100751 y = j(x,W) ~ Y = ( / ,(X Wû. «4

[0076] The results provided by these functions fN, fN_],... are stored in memory as a hidden layer denoted H. Therefore: 100771 r = / dAd ■ ■ • W2\ W3) ~ Y = HN = wv) 100781 Y = fw,v) » y = »4

[0079] Thus, for example, an MLP layer (extended and complex version of the perceptron original) is a transformation from a 1-dimensional vector to a 1-dimensional vector via matrix multiplication, vector addition, and the use of a nonlinear activation function. Considering that this transformation is denoted MLP and is applied from a vector X of N values ​​to a vector Y of M values. The MLP layer consists of multiplying X by a weight matrix W, adding a bias vector b of M values ​​(note that W and b are considered as the internal data globally denoted W for the transformation), and passing all the results into an activation function denoted q>.

[0080] Y = MLP{X, W) F = O' x X + b)

[0081] For example, if we combine two MLP functions, it is possible to write that:

[0082] Y = NN(X, W) Y = H2 = MLP2(MLP / X, W2)

[0083] MLP^X, x X + b^

[0084] Y = H2 = MLP2(H{, W2) H2 = <P2(W2 x H} + b^

[0085] By describing neural networks in this way, any type of neural network can be built. For example, the Lenet5 architecture which, using: C for the convolution layer transformation, SS for the downsampling transformation and FC for "fully connected" (i.e. MLP), X for example an MNIST image as input and Y the prediction of handwritten digits can be written:

[0086] Y = MLP / MLP^MLP^SS / C / SS^C / X', W»)))

[0087] The transformations (i.e. all the different functions f) can be any type of neural network cell, for example: - MLP a 1-dimensional to 1-dimensional transformation. - RNN cells such as LSTM, GRU, MGU, STAR or SRU which are a transformation from 1 dimension to 2 dimensions (for illustration, we can see an MLP with a time dimension). - CNN cells (multidimensional to multidimensional) or combined CNN-RNN such as ConvLSTM. - Simple operations: Buffer (example 1D to 2D such as from an MLP layer to an RNN), Addition or aggregation layer (1D to 1D, 2D to 2D,...), Subsampling (Average, Max,...) - Others: Extended Kalman Filter (EKF) or Uncentered Kalman Filter (UKF) Self-organizing maps (SOM), Markov decision processes (MDP), POMDP,... 2.2 Backpropagation operations

[0088] On the other hand, to enable the neural network to learn, we need the inverse operations to update all the functions / based on the error between the generated output Y (from the neural network inference pass) and the output expected Y*, error noted ôY:

[0089] ôY = Error(Y*, Y)

[0090] The goal of backpropagation is to update the internal data W for each function / , in order to obtain a more efficient neural network at the next iteration (i.e. with a lower error). The very first step is to calculate the gradient of all hidden layers, the inverse of inference from output to input, for example:

[0091] 5X = f\5Y,W)

[0092] ÔX = ÔH^ Wj)

[0093] f'2(ÔH2,W2)

[0094] ÔHNA = f\(ôY,WN)

[0095] These gradients are calculated from the description using a propagation rule (Daisy Chain Rule), then used to update the internal transformation data using an optimization function, noted optim, such as Stochastic Gradient Descent (SGD) normal and with momentum, ADAM or ADAMAX. These algorithms are based on a set of parameters noted P (e.g.: the learning rate, these parameters can be different for each layer).

[0096] SW — optimÇW, 5H, P)

[0097] dWN = optim (WN, ÔY, PN)

[0098] Thus, in this gradient descent calculation, the objective is to update the internal data W, which correspond to the weights of the matrices as well as to the biases. The learning (in which the backpropagation operation is carried out) is therefore again a matrix calculation.

[0099] Consequently, for the example of neural networks, the implementation of these consists of fixing a certain number of parameters (number of matrices, dimensions thereof, weights, biases, activation functions, optimization functions, etc.) which make it possible to completely implement a network in the form of a particular component. On this basis, the inventors worked to formalize a certain genericity of these parameters in order to implement, within a component, a set of generic functions allowing the implementation of a much larger number of neural networks, and more generally a much larger number of mathematical functions. 2.3 Memory Management

[0100] However, the conceptual view of multiple dimensions is not applicable to the memory layout of an architecture based on an FPGA, an ASIC, a CPU or a GPU. All dimensions must be converted into a single dimension. Thus, a 3-dimensional entity (matrix) (DIM1, DIM2 and DIM3) is stored in a simple format (all elements are stored in a row, and addressable in a row) with for example DIM1 first, then DIM2 and finally DIM3. This layout could have been implemented in another order DIM3 first, then DIM2 and finally DIM1. The choice of this layout of the data to be stored in memory is important to consider in order to access it in the most efficient way possible, especially in the case of a multi-level memory architecture.

[0101] Thus, on a (physical) device, according to the technique developed by the inventors, this memory is used to read the available data (inputs X and weights W) according to their location in the memory for a transformation f and store (i.e. write) the output Y to a specific memory location. The logic of the transformation (all the necessary operators) is implemented on the pseudo-generic component. The latter accesses the memory according to access parameters provided to it, loads the data into memory, performs the requested operations on this data according to the execution parameters provided, and writes the results obtained to a specific address in its internal memory.

[0102] In other words, for a given pseudo-generic component, the implementation of a matrix operation consists of loading, row by row or column by column, the matrices involved in the matrix operation (for example a multiplication), then performing the operation on each term of the matrix and finally recording each of these terms in a resulting memory area of ​​the pseudo-generic component. Other mathematical operations can also be applied to the terms of the matrix before their recording if necessary. Thus, according to the invention, the implementation (use) of a pseudo-generic component to manage several different mathematical operations consists of executing a method comprising as illustrated in relation to [Fig.2], in a particular embodiment: - a configuration step 21 of the component using configuration parameters; the configuration parameters define how the component must run; for a neural network, for example, the configuration parameters define the sizes of the different layers, the activation functions to be used, the order of the layers, etc.

[0103] and at least one iteration of the following steps: - a step 22 of loading a set of execution parameters: these are typically the parameters of the different functions, such as for example the numerical constants, factors, etc.; in the case of a neural network, this is typically the loading of the weights of a matrix, the activation function, and any associated biases for the layer being executed; - a step 23 of loading a set of execution variables; this typically involves loading the input values; in the case of the neural network, this is the input vector or matrix of the layer being executed; - at least one calculation step 24 of at least one execution result as a function of the set of execution parameters and execution variables; for a conventional mathematical function, this involves calculating the image of the set of execution variables using the execution parameters; for the case of a neural network, this involves calculating the result matrix of the current layer; - a recording step 25, within a memory area of ​​the component, of said at least one execution result.

[0104] The different steps of this method are implemented as many times as necessary to complete all the calculations. In the case of the neural network, these steps are implemented until, for example, one or more output vectors are obtained as a function of the input variables. 2.4 General implementation

[0105] Based on these developments, the inventors have therefore developed a method for creating a pseudo-generic component, within a platform, allowing the implementation of several types of calculations by this single pseudo-generic component. It is then possible to carry out several different types of calculations (for example, implement several different neural networks) without it being necessary to modify the pseudo-generic component itself. This creation method is based on the implementation (manual or automated) of a configuration file which makes it possible to frame (define) the possible functionalities of the component. This configuration file is used to generate, from a compiler, two files: a binary file of the "bitstream" type, used to define the programmable logic of the pseudo-generic component (FPGA or ASIC) and a file executable by a processor.The final implementation architecture of these two files is a platform comprising a general-purpose processor (CPU), a memory managed by this processor and a generic component, the general-purpose processor and the generic component exchanging data via at least one data bus.

[0106] The file executable by the processor is a so-called orchestration file: it manages the configuration of the component, the data necessary for the implementation of the component (for example the training data for a neural network, which are the weights of the matrices, the variables, etc.) from the memory available for this processor. In this architecture, the component (defined by its “bitstream” file), for its part, receives, from the processor, the necessary configuration data, then once configured, accesses execution data in L1 memory (input variables, weights of the network layer(s), etc.). The component executes the operations instructed to it by the processor (matrix operations).

[0107] We thus have a “configurable” platform comprising the general processor, the memory associated with this processor (called level 1 memory - L1), a “generic” component itself having its own working memory (called level 0 memory - LO). Such a platform can for example be implemented within an “intelligent” vehicle. Thanks to the “generic” component, it is possible to configure new functions to be fulfilled, for example new artificial intelligence algorithms (which are matrix calculation algorithms as explained previously), which were not originally planned during vehicle manufacture.

[0108] In the following sections, an example of the implementation of such a platform oriented towards the processing of artificial neural networks (i.e. matrix calculation) is described. The inventors have developed the method described to have, within a “pseudo-generic” component, the functionalities necessary for the implementation of several typologies of neural networks while limiting the impact of this genericity.

[0109] 3. Initial parameterization of the pseudo-generic component and the platform associate

[0110] The creation of a combined platform includes a first parameterization step discussed in this section. The very first and initial step of this parameterization defines the "genericity limits" of the component. Since genericity has a cost, designers must define what the pseudo-generic component is capable of doing. An example of a parameterization file is illustrated in relation to [Fig.3], in a particular embodiment.

[0111] This file is completed by a user to generate a pseudogeneric component and its associated software. The user determines which set of operators, from a predetermined set of available operators, should be included in the component. He also sets maximum values ​​for certain parameters.

[0112] The parameters in this parameterization file are used to generate synthesizable VHDL code from a completely generic description of a component. By determining the operators and parameters, the user can thus adjust the resulting pseudo-generic component to his specific needs and remove unnecessary logic. The logical size (in an FPGA or ASIC) is then reduced. The parameterization part, as illustrated in [Fig.4], allows two entities to be defined:

[0113] (a). The COB ANN SW entity which is the orchestrator;

[0114] (b). The COBANN HW entity which contains the main invention. This is the coprocessor artificial neural network that performs the calculations of the expected neural network based on the method described previously.

[0115] The inventors thus obtain a dual architecture comprising a “Processing System” (PS), i.e. the processor, associated with the logic (PL running on an FPGA or an ASIC) as described in [Fig.5]. - PS: processor-based processing system (CPU); - MGR: Manager (processing software) for configuring the programmable logic component; - DH: Data Manager; - CH: Configuration Manager; - TH: Training Manager; - PER: Peripherals; - PL: programmable logic component; FPGA or ASIC component; - El: executor 1 within the programmable logic component; - E2: executor 2 within the programmable logic component; - FS: access to the file system; - MC: Memory controller;

[0116] The component described in [Fig.5] includes two entities. A Manager entity (also called "software component") which is in charge of the overall management (configuration, data provision) as well as the connections with the different PER peripheral interfaces (for example cameras or sensors on a real embedded system). Another entity is the Executor component El, E2 (also called "hardware component") which can be unitary or multiple (two are shown in [Fig.5]). It is the Executor which is configured by the Manager thanks to the configuration described previously. These components are illustrated in [Fig.5] which describes a Manager connected to two Executors. One of these executors has the ability to do online training (on the component) so it is able to do back-propagation (i.e. BP capable). The other is only able to do inference (feedforward i.e.FF only) and therefore assumes offline training (according to the offline training paradigm). As a reminder, the training capabilities of a component are defined during the parameterization step, described previously.

[0117] Thus, once configured and synthesized, the pseudo-generic hardware component is then implemented in an FPGA, or in an ASIC and is controlled by the PS processor (the whole being able to be integrated within a SOC). This hardware component is considered as having at least two memory levels, namely LO and LL As explained previously, the LO memory is the internal memory of the hardware component such as block memories (BlockRAM - BRAM) on FPGA. The L1 memory is the external memory accessible via the PS processor such as DDR memory. Thanks to this extended memory, and the management (configuration) of this memory by the software component (running on the PS processor), the inventors allow the use of larger neural networks. It is important to note that, due to its hardware construction, all PS and PL entities access the same DDR memory denoted Ll. This implies that the data accessible to the Manager on the PS are also accessible to the Executors of the PL. This is important later in explaining when considering the interactions between the Manager (on the PS) and the Executor (on the PL).

[0118] It is also possible to implement this architecture on FPGAs (and ASICs) only. In this type of architecture, the Manager component is directly integrated into the logic of the FPGA (or ASIC) as well as the memory controller. This component is then not necessarily interfaced with peripherals, nor with a file system. All transactions for configuration and data management can be carried out via an Ethernet interface (in 1 Gbits for example), thus opening the possibilities towards distributed neural networks.

[0119] The pseudo-generic component includes several interfaces, as explained in [Fig.6] and [Fig.7]. An Executor entity has several interfaces in order to function. A component with training capabilities has more interfaces than the component that is only capable of inference. A BP capable executor ([Fig.6]) has thirteen different interfaces. There are nine used for inference (Feed-Forward i.e. FF, [Fig.7]): (1) config_in, (2) data_in, (3) result_out, (4) transf_data_inout, (6) layers_data_inout, (7) opcode_in, (8) lo-calstate_out and (9) memory_layout_out. There are 4 used for learning (Back-Propagation i.e. BP): (10) label_in, (11) dtransf_data_inout, (12) dlayers_data_inout and (13) random_table_in. Several interfaces are only used for training neural networks (when these features are implemented).Therefore, they are not present when the component is set up only for inference (calculating predictions).

[0120] [config_in] contains the entire configuration of the deep learning entity (neural network). It specifies the list of transformations to be applied from input to output. It defines the characteristics of each of the operators (i.e. transformations) such as the activation function and how and where to read the different active operands for this transformation (i.e. transformation-specific data and layer data).

[0121] [data_in] This is the data that is presented to the HW pseudo-generic component in order to be processed. This can be a single sample (i.e. an image, a flow, etc.) or several sample inputs if the network allows backpropagation for training (see parameterization section) (noted X in the generalization of the calculations).

[0122] [result_out] this is the output of the feedforward calculation of the neural network (denoted Y in the generalization of the calculations).

[0123] [transf_data_inout] This contains all the transformation data (weights) of the ANN (denoted W in the generalization of the calculations).

[0124] [layers_data_inout] These are the intermediate results of each layer (denoted H in the generalization of the calculations).

[0125] [report_out] Various indicators used to determine network performance.

[0126] [opcode_in] sent by the PS to execute the HW pseudo-generic component. This can be (non-exhaustive): configuration, inference, learning.

[0127] [localstate_out] returns the state of the pseudo-generic component. It can be (non-exhaustive): parameterized, Target error reached, Target error not reached.

[0128] [memory_layout_out] the format and dimensions of the outputs. Since the memory layout of each layer can be deduced from the transformations, the programmable logic component provides the layout of the results to the PS processor.

[0129] The training interfaces of neural networks are:

[0130] [label_in] This is the label of the data for training, i.e. it is the expected output for the neural network. It must be compared with the result above to generate an internal optimization error of the network via backpropagation (denoted Y in the generalization of the calculations).

[0131] [dtransf_data_inout] These are the delta coefficients, used for calculating the gradient (noted W in the generalization of the calculations).

[0132] [dlayers_data_inout] These are the delta layers, used for calculating the gradient (noted H in the generalization of the calculations)

[0133] [random_table_in] an array of random values ​​(generated by the software running on the PS processor), to shuffle the training samples.

[0134] Thus, the initial parameterization makes it possible to have a pseudogeneric component associated with a general processor.

[0135] The Executor interfaces described are, in operational conditions, buses for accessing memory areas of the external memory L1 (therefore the DDR). It is recalled that, on SoC implementations, this data is accessible via the memory manager of the PS (the CPU). As illustrated in [Fig.8], the role of the Manager (MGR) is to provide the Executor (E2, for the example) with the memory base addresses (Adr_Bse) on which it can read / write the data linked to these different interfaces. The structure of the data from these base addresses is determined from the configuration as explained later. 4. Configuring the pseudo-generic component

[0136] As described in the generalization of calculations section, a neural network is considered as a combination (usually a succession) of transformations applied to an input in order to generate an output. The configuration of the programmable logic component allows to define this succession of transformation from input to output. The memory layout of the input must also be defined. For each transformation, the user can define the type of transformation (see a list below), the type of the activation function, the data layout of the input (of the layer). The different parameters of the backpropagation layers can also be defined, if they were included during the parameterization phase.The configuration file includes for example an identifier, a neural network size, a structuring of the input data (size and dimension for example), the structuring of the layers (for each layer) [type, format of the data - internal, input, output -, activation function to use, dimension of the vectors or matrices], the structuring of the back-propagation, for each back-propagation layer (type of error, parameters, learning rate, moment, regularization), the number of cycles ("epoch"), etc.

[0137] In summary, the configuration is used to define the types of operators (i.e. transformations) to be applied and its characteristics, and to which operands they apply (internal transformation operands like weights and data from the previous layer). The configuration can be stored in a file. The PS processor performs an initialization of the pseudo-generic component with this configuration before the calculations. This configuration is defined by the user, it can be determined using popular artificial intelligence tools (framework) such as Pytorch or TensorFlow.

[0138] Non-exhaustive list of inference parameters (feedforward): - Transformation type: conv, sub max, sub min, sub avg, sub max, conv w sub, FC, LSTM, STAR, GRU, MGU, memory transformation, aggregation - Activation function: none, sigmoid, tanh, tanh opt, sig int, tanh int, relu, softmax

[0139] Non-exhaustive list of backpropagation parameters: - Descending gradient type: Stochastic Gradient Descent (SGD), Stochastic Gradient Descent with Momentum (SGDM), adam, adamax, etc.

[0140] Usually, the internal memory of an FPGA device (or even an ASIC) is limited and / or extremely expensive. As indicated above, the approach adopted is not limited to local memory because it is possible to use the external memory components that allow via write / read operations to load and export the data necessary for each operation. The burst access of the LO to L1 memory is defined in the setup phase.

[0141] 5. Memory exchange processing / Component configuration

[0142] 5.1. Generation of a configuration and data in L1 memory

[0143] The configuration of the component consists, once the component has been "instantiated" (i.e. once the genericity limits of this component have been defined and a functional component is available), in allowing the component to interact correctly with the L1 memory on each of its memory bus interfaces (as defined previously). A configuration, within the genericity limits of the parameterization, is composed of three main elements: inputs, labels and transformation lists. If we only consider neural networks performing inference, then the configuration does not include labels. We are interested here in the data presented on the config_in interface. This configuration is for example generated from an XML file whose formalism, structure and elements that compose it have been defined.This XML file is given as input to the Manager component which generates the correct data structure(s) for the config_in interface. It is quite possible to take into account ONNX (Open Neural Network Exchange) type configurations or even specific files coming from the most used frameworks (development frameworks) such as Pytorch (and its models saved in "*.pth"). If we consider offline training components, the model descriptions integrate the data of the different transformations specific to the neural network. Thus, the transf_data_inout interface is also constituted in accordance with the configuration of the config_in interface.In other words, the manager, based on a configuration file (config_in interface), informs the hardware component (Executor) of the type as well as the structure and location (memory addresses) of the transformation data (the data of the different layers present on the transf_data_inout interface) in the LL memory. This configuration also allows the executor to know the structure and location of the input data (inputs_in and possibly labels_in if the component is capable of training). Thus, the executor can implement the functions planned for the neural network from the planned configuration while correctly accessing the necessary data from the data structure and base addresses also provided by this configuration.Then, the executor can perform the neural network calculations by writing / reading the intermediate data for the different layers in the L1 memory (via the layers_data_inout interface) and provide the final result on the result_out interface. The executor can also provide the data structure (addresses and dimensions) for the data it generates via the memory_layout_out interface. 5.2 Configuring the Executor

[0144] A configuration is thus provided to the Executor component from a given base address. The Manager provides it with the different addresses on which these different interfaces will be connected and used. The Manager also transmits an opcode (operation code) to indicate to it that it must configure itself from this data present on its config_in interface at the base address provided. The Executor therefore determines from the configuration provided (containing the description of the inputs and transformations, as well as the labels) how to correctly read and write the data and results in the DDR (the L1 memory). The component also provides on the memory_layout_out interface the structure of the data that it generates (typically the outputs and the data for the intermediate layers, so that the MGR manager can read and decode them).At the end of this configuration phase, the Executor returns its configuration state and if this state is correct, then the component is ready to run and is configured from the provided description.

[0145] In other words, the Executor reads the configuration from the addresses provided by the Manager, then it extracts, from the addresses provided, in memory L1, the data structures to be used and the appropriate functional characteristics, then determines: how it performs the readings and writings in memory L1 and what is the functional logic to be applied (MLP, CNN, activation functions, etc.). 5.3 Examples

[0146] The implementation of the method of the invention is illustrated using a concrete example. To do this, we assume that we have three different neural networks that we wish to apply from the same input data (here, the historical case study of MNIST). These networks are (1) ANN-1 is a multi-layer MLP network; (2) ANN-2 is a Lenet-5 type CNN network and (3) ANN-3 is an LSTM type RNN (recurrent) network. The following hypotheses are formulated: - These three networks are capable of solving the MNIST classification problem. They take as input a gray image of size 28x28 and provide as output a classification vector of size 10 which corresponds to the probability of the digit recognized on the image provided as input (between 0 and 9); - The considered DDR (thus the L1 memory) is large enough to contain all the data of the three networks. This means that the configuration data, the data for the intermediate levels and the transformation data for each network are available in L1 memory (this assumption is purely rhetorical since many DDRs for embedded cards can reach 16, 32 or even 64 Gbits, which is quite sufficient to deal with this illustrative problem). - The three networks were trained offline using a tool (Pytorch, Tensorflow,...) and were saved in ONNX format in the files annl.onnx, ann2.onnx and ann3.onnx. These files were provided to the Manager entity in order to place the data and associated configurations in the Ll memory. The Executor was correctly configured with the logical functionalities required for the execution of these three types of networks. - The Executor has limited logic, so it can only handle the execution of one neural network at a time (one can configure and execute ANN-1 or ANN-2 or ANN-3).

[0147] Thus, after a step of generating the configuration in memory L1 and copying the associated data, as illustrated in relation to [Fig.9] (for the ANN-1 network), [Fig. 10] (for the ANN-2 network) and [Fig. 11] (for the ANN-3 network) the Manager can configure the Executor so that it executes each of the networks available in memory LL

[0148] The Executor provides, successively, for each of the three networks, on the interfaces "config_in", "transf_data_inout" and "rents_data_inout", "data-in" the data necessary for the implementation of the network. Once configured with this data, the Executor performs the calculations for which it was configured, including in particular, the matrix calculations and the activation functions of the implemented logic. The data are read / written in LO memory to perform the calculations "internal" to the Executor and when these calculations are completed (or when the memory must be purged (for example for a possible following calculation cycle), the data are written in Ll memory (on the interfaces "result_out" or "memory_layout_out").

[0149] The Executor reads / writes its operating data at different addresses in Ll, the associated configuration allows it to interpret the memory structure linked to these inputs / outputs. It can be noted that in this specific case, the component reads the input data (data_in), will write its output data (result_out) and provides the structure of its memory interfaces (i.e. memory layout on the memory_layout_out interface) from the same addresses (therefore the same memory areas). It may be relevant to use different areas in particular if the neural networks used (and to be configured on the component) have different objectives / goals. 5.4 Execution of L0 / L1 interaction management

[0150] Once configured as described above, the Executor must implement the operations for which it has been configured and for this it performs several “copy” operations in order to acquire all the data (operands) necessary for carrying out the operations configured in L1 in order to have them available in memory L0. Conversely, once the operations have been carried out, it performs copy operations of the results obtained (by its calculations) from its local memory L0 to external memory LL. It is possible to illustrate this operation by considering the elementary operation of an MLP network in the form of expression 1 considering the matrices or in the form of expression 2 considering the elements (the bias b is a particular weight forming part of the transformation data).

[0151]

[0152] y = MLP(X, IV, b) = (p(W x X + b) (1) ax 4+ M(2)

[0153] The vector X is a vector of size N accessible on the “layers_data_inout” interface of the L1 memory and likewise the vector resulting from this operation noted Y (of size M). The transformation data W and b are accessible on the interface "transf_data_inout" of LL memory If we consider that N=1000 and M =500, we can determine that the weight matrix W is of size M *N =500* 1000 and the bias vector is of size M = 500. A possible division is therefore to copy from L1 memory to L0 local memory the data necessary to calculate the elements of Y in a unitary way. Thus step by step, each element y, is calculated using the following method, as illustrated in [Fig. 12] and [Fig. 13] which respectively present the calculation of the element Yi then of the element Y2:

[0154] (1) Copy data from L1 to L0,

[0155] (2) Calculation and generation of the result from the implemented logic; and

[0156] (3) Copying the result from memory L0 to external memory LL

[0157] It should be noted that the copy of the input vector X is done only once for all calculations of Y (because this input vector X does not change).

Claims

Claims

1. Electronic data processing platform, platform comprising a general-purpose processor (PS) and a so-called level 1 memory (Ll), said platform further comprising a programmable logic component (PL) comprising a so-called level 0 memory (LO), the programmable logic component being parameterized, by means of a binary file for implementing said programmable logic component, so as to define typologies of executable calculations, said platform further comprising orchestration software implemented by means of an executable file running on the general-purpose processor, the orchestration software being responsible for providing, to the programmable logic component, a set of configuration data for its execution within the limits of the typologies of executable calculations defined by said binary file,said configuration data comprising said binary file and said executable file being generated by compiling a configuration file defining a set of operators included in said programmable logic component from a predetermined set of available operators.,

2. Electronic platform according to claim 1, characterized in that the programmable logic component is an FPGA or an ASIC.

3. Electronic platform according to claim 1, characterized in that the programmable logic component is configured so as to perform inference calculations for artificial intelligence.

4. Electronic platform according to claim 1, characterized in that the programmable logic component is configured so as to execute learning calculations for artificial intelligence.

5. Electronic platform according to claim 1, characterized in that the orchestration software is arranged to provide the execution component with a level 1 memory address at which the configuration data set is accessible, and to provide the programmable logic component with an execution command.

6. Electronic platform according to claim 5, characterized in that the programmable logic component is further configured to extract the data structures and the functional characteristics of the calculations to be executed and configured to determine how to read and write data in level 1 memory (L1) and determine an execution logic to be implemented.

7. Electronic platform according to claim 6, characterized in that the execution logic to be implemented belongs to the group comprising: - a multi-layer perceptron type neural network; - a convolutional type neural network; - a recurrent neural network.

8. Electronic platform according to claim 5, characterized in that the programmable logic component is configured to copy, within its level 0 memory (LO), said configuration data set and to perform the calculations required by the configuration data of the configuration data set.

9. Electronic platform according to claim 6, characterized in that the programmable logic component is configured to recursively load, into the level 0 memory (LO), input data from the level 1 memory (L1), then to calculate, from said input data, output data and to copy said output data into the level 1 memory (L1).

10. Method for obtaining an electronic platform according to claim 1, characterized in that the creation of the programmable logic component comprises the following steps: - obtaining (11) a parameter file for the programmable logic component, said parameter file defining a set of operators to be included in said programmable logic component from a predetermined set of available operators, said parameter file comprising, for at least one characteristic of said programmable logic component, a value representative of the capacity of the programmable logic component to have this characteristic; - compiling (12) the parameter file for the programmable logic component, delivering on the one hand a binary file for implementing the programmable logic component, and on the other hand an executable file for orchestrating the operation of the programmable logic component;- creation (13) of the programmable logic component in function; said binary file; - loading (14) of said executable file by the general processor, delivering software for orchestrating the operation of the programmable logic component.

11. Method of using the programmable logic component of an electronic platform according to claim 1, characterized in that it comprises a step of configuring (21) the programmable logic component using configuration parameters, and at least one iteration of the following steps: - loading (22) a set of execution parameters; - loading (23) a set of execution variables; - calculation (24) of at least one execution result as a function of all the execution parameters and the execution variables; - recording (25), within a memory area of ​​the programmable logic component, said at least one execution result.