INTEGRATED CIRCUIT CONFIGURED TO RUN AN ARTIFICIAL NEURAL NETWORK

The integrated circuit design addresses the energy consumption and structural complexity issues of existing neural network execution circuits by using dedicated memories and rotation shift circuits, enabling faster and more energy-efficient execution with improved parallelization capabilities.

FR3141543B1Active Publication Date: 2025-06-27STMICROELECTRONICS FRANCE +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
FR2022011288
Authority / Receiving Office
FR · FR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-10-28
Publication Date
2025-06-27
Estimated Expiration
2042-10-28

AI Technical Summary

Technical Problem

Existing integrated circuits used to execute artificial neural networks are energy-consuming, have complex and bulky structures, and lack flexibility in parallelizing the execution of neural networks.

Method used

An integrated circuit design that includes separate memories for storing neural network parameters and input/output data, along with rotation shift circuits for efficient data manipulation, and a computing unit capable of parallelizing operations.

Benefits of technology

This design reduces energy consumption, simplifies the circuit structure, and enables faster execution of neural networks while allowing for parallelization of operations, thereby improving efficiency and flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000012_0000
    Figure 00000012_0000
  • Figure 00000012_0001
    Figure 00000012_0001
Patent Text Reader

Abstract

According to one aspect, an integrated circuit is provided comprising: - a first memory (WMEM) configured to store parameters of a neural network, - a second memory (DMEM) configured to store data provided as input to the neural network or generated by this neural network, - a computing unit (PEBK) configured to execute the neural network, - a first rotation shift circuit (BS1), the first rotation shift circuit being configured to transmit the data from the second memory to the computing unit, - a second rotation shift circuit (BS2), the second rotation shift circuit being configured to deliver the data generated by the execution of the neural network by the computing unit, - a control unit (CTRL) configured to control the computing unit (PEBK) and the first and second rotation shift circuits (BS1, BS2). Figure for abstract: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: INTEGRATED CIRCUIT CONFIGURED TO EXECUTE AN ARTIFICIAL NEURAL NETWORK

[0001] Embodiments and implementations relate to artificial neural networks.

[0002] Artificial neural networks are used to perform given functions when executed. For example, one function of a neural network may be classification. Another function may be generating a signal from a signal received as input.

[0003] Artificial neural networks generally comprise a succession of layers of neurons.

[0004] Each layer takes as input data to which weights are applied and outputs output data after processing by activation functions of the neurons of said layer. This output data is transmitted to the next layer in the neural network. The weights are configurable parameters to obtain good output data.

[0005] Neural networks can for example be implemented by final hardware platforms, such as microcontrollers integrated into connected objects or in specific dedicated circuits.

[0006] Neural networks are generally trained during a learning phase before being integrated into the final hardware platform. The learning phase can be supervised or unsupervised. The learning phase allows the weights of the neural network to be adjusted to obtain good output data from the neural network. To do this, the neural network can be executed by taking as input already classified data from a reference database. The weights are adapted according to the data obtained at the output of the neural network compared to expected data.

[0007] Running a neural network trained by an integrated circuit requires manipulation of a large amount of data.

[0008] This data manipulation can result in significant energy consumption, particularly when the integrated circuit must perform numerous write or read memory accesses.

[0009] These integrated circuits used to implement neural networks are therefore generally energy-consuming and have a complex and bulky structure. In addition, these integrated circuits are not very flexible in terms of parallelizing the execution of the neural network.

[0010] There is therefore a need to provide an integrated circuit that can quickly execute a neural network while reducing the energy consumption required to execute an artificial neural network. There is also a need to provide such an integrated circuit that has a simple structure so as to reduce its dimensions.

[0011] According to one aspect, there is provided an integrated circuit comprising: - a first memory configured to store parameters of a neural network to be executed, - a second memory configured to store data provided as input to the neural network to be executed or generated by this neural network, - a computing unit configured to execute the neural network, - a first rotation shift circuit between an output of the second memory and the calculation unit, the first rotation shift circuit being configured to transmit data from the output of the second memory to the calculation unit, - a second rotation shift circuit between the computing unit and the second memory, the second rotation shift circuit being configured to deliver the data generated during the execution of the neural network by the computing unit, - a control unit configured to control the computing unit and the first and second rotational shift circuits.

[0012] Such an integrated circuit has the advantage of integrating memories for storing the parameters of the neural network (these parameters including the weights of the neural network but also its topology, i.e. the number and type of layers), input data of the neural network and data generated at the output of the different layers of the neural network. Thus, the memories can be accessed directly by the calculation unit of the integrated circuit, and are not shared through a bus. Such an integrated circuit therefore makes it possible to reduce the movement of the parameters of the first memory and the data of the second memory. This makes it possible to make the execution of the artificial neural network faster.

[0013] The use of a memory to store the parameters of the neural network allows the adaptability of the circuit to the task to be carried out (the weights as well as the topology of the neural network being programmable)

[0014] Furthermore, the use of rotation shift circuits allows for low-power data manipulation. In particular, the first rotation shift circuit allows for simple reading of data stored in the second memory when this data is required for the execution of the neural network by the computing unit. The second rotation shift circuit allows for simple writing of data generated by the computing unit during the execution of the neural network into the second memory. The rotation shift circuits are dimensioned so that that, for the execution of the neural network, useful data can be written into these circuits on data, as soon as the latter data is no longer useful for the execution of the neural network.

[0015] The data and weights being placed in the memories of the integrated circuit can be accessed at each clock stroke of the integrated circuit.

[0016] Such an integrated circuit has a simple and compact structure, and consumes little energy, in particular thanks to the use of rotational shift circuits instead of using a cross interconnection circuit (also known by the English term "crossbar").

[0017] In an advantageous embodiment, the calculation unit comprises a bank of processing elements configured to parallelize the execution of the neural network, the first rotation shift circuit being configured to transmit the data from the second memory to the different processing elements.

[0018] Such an integrated circuit allows parallelization of operations during the execution of the neural network.

[0019] Preferably, the integrated circuit further comprises a first multiplexer stage, the first rotation shift circuit being connected to the second memory via the first multiplexer stage, the first multiplexer stage being configured to deliver to the first rotation shift circuit a data vector from the data stored in the second memory, the first rotation shift circuit being configured to shift the data vector from the first multiplexer stage.

[0020] Advantageously, the integrated circuit further comprises a second multiplexer stage, the calculation unit being connected to the first rotation shift circuit via the second multiplexer stage, the second multiplexer stage being configured to deliver the data vector shifted by the first rotation shift circuit to the calculation unit.

[0021] In an advantageous embodiment, the integrated circuit further comprises a buffer memory, the second rotation shift circuit being connected to the calculation unit via the buffer memory, the buffer memory being configured to temporarily store the data generated by the calculation unit during the execution of the neural network before the second rotation shift circuit delivers this data to the second memory. For example, this buffer memory can be implemented by a physical memory or by a temporary storage element (flip-flop).

[0022] Preferably, the integrated circuit further comprises a pruning stage between the buffer memory and the second rotation shift circuit, the pruning stage being configured to delete data, in particular unnecessary data, among the data generated by the computing unit.

[0023] Advantageously, the second memory is configured to store data matrices provided as input to the neural network to be executed or generated by this neural network, each data matrix being able to have several data channels, the data of each data matrix being grouped in the second memory in at least one data group, the data groups being stored in different banks of the second memory, the data of each data group being intended to be processed in parallel by the different processing elements of the calculation unit.

[0024] The data matrices may be images received as input to the neural network, for example. The position of the data then corresponds to pixels of the image. The data matrices may also correspond to a feature map generated by the execution of a layer of the neural network by the computing unit (also known by the English terms “feature map” and “activation map”).

[0025] Placing the data and parameters of the neural network in the first memory and the second memory of the integrated circuit allows access to the data necessary for the execution of the neural network at each clock stroke of the integrated circuit.

[0026] In an advantageous embodiment, each data group of a data matrix comprises data from at least one position of the data matrix for at least one channel of the data matrix.

[0027] Such an integrated circuit is thus adapted to parallelize the execution of the neural network in width (on the different positions of the data in the data matrix) and in depth (on the different channels of the data matrices). In particular, the calculation unit may comprise a bank of processing elements configured to parallelize the execution of the neural network in width and in depth.

[0028] According to another aspect, there is provided a system on chip comprising an integrated circuit as described above.

[0029] Such a system on chip has the advantage of being able to execute an artificial neural network using only the integrated circuit. Such a system on chip therefore does not require interventions from a microcontroller of the system on chip for the execution of the neural network. Such a system on chip also does not require the use of a common bus of the system on chip for the execution of the neural network. Thus, the artificial neural network can be executed more quickly, more simply while reducing the energy consumption required for its execution.

[0030] Other advantages and characteristics of the invention will appear on examining the detailed description of embodiments, which are in no way limiting, and the drawings. annexed on which:

[0031] [Fig.l]

[0032] [Fig.2]

[0033] [Fig.3] illustrate embodiments and implementations of the invention.

[0034] [Fig.l] illustrates an embodiment of a system on chip SOC. The system on The chip typically includes a microcontroller MCU, a data memory Dat_MEM, at least one code memory C_MEM, a time measurement circuit TMRS (in English "timer"), general purpose input-output ports GPIO, and an I2C communication port.

[0035] The system-on-chip SOC also includes an NNA integrated circuit for implementing artificial neural networks. Such an NNA integrated circuit may also be referred to as a "neural network acceleration circuit."

[0036] The system on chip SOC also includes buses for interconnecting the various elements of the system on chip SOC.

[0037] [Fig.2] illustrates an embodiment of the NNA integrated circuit for implementing neural networks.

[0038] This NNA integrated circuit comprises a PEBK calculation unit. The PEBK calculation unit comprises a bank of at least one PE processing element. Preferably, the PEBK calculation unit comprises several processing elements PE#0, PE#1, ..., PE#N-1. Each PE processing element is configured to perform elementary operations for the execution of the neural network. For example, each PE processing element is configured to perform elementary operations of convolution, dimension reduction (in English "pooling"), scaling (in English "scaling"), activation functions of the neural network.

[0039] The NNA integrated circuit further comprises a first WMEM memory configured to store parameters of the neural network to be executed, in particular weights and a configuration of the neural network (in particular its topology). The first WMEM memory is configured to receive the parameters of the neural network to be executed before the implementation of the neural network from the data memory Dat_MEM of the system on chip. The first WMEM memory may be a volatile memory.

[0040] The integrated circuit further comprises an SMUX shift stage having inputs connected to the outputs of the first WMEM memory. The SMUX shift stage is thus configured to receive the parameters of the neural network to be executed stored in the first WMEM memory. The SMUX shift stage also comprises outputs connected to inputs of the PEBK calculation unit. In this way, the PEBK calculation unit is configured to receive the parameters of the neural network in order to be able to execute it. In particular, the SMUX shift stage is configured to select weights and configuration data from memory to deliver them to the PEBK calculation unit, and more specifically to the various PE processing elements.

[0041] The NNA integrated circuit also comprises a second DMEM memory configured to store data provided to the neural network to be executed or generated during its execution by the PEBK calculation unit. Thus, the data may for example be input data of the neural network or data (also designated by the English term “activation”) generated at the output of the different layers of the neural network. The second DMEM memory may be a volatile memory.

[0042] The NNA integrated circuit further comprises a first multiplexer stage MUX1. The first multiplexer stage MUX1 comprises inputs connected to the second memory DMEM. The first multiplexer stage MUX1 is configured to deliver a data vector from the data stored in the second memory DMEM.

[0043] The NNA integrated circuit further comprises a first rotation shift circuit BS1 (also known by the English expression “barrel shifter”). The first rotation shift circuit BS1 has inputs connected to the outputs of the first multiplexer stage MUX1. The first rotation shift circuit BS1 is thus configured to be able to receive the data transmitted by the first multiplexer stage MUX1. The first rotation shift circuit BS1 is configured to shift the data vector of the first multiplexer stage MUX1. The first rotation shift circuit BS1 has outputs configured to deliver the data of this first rotation shift circuit BS1.

[0044] The NNA integrated circuit also comprises a second multiplexer stage MUX2. This second multiplexer stage MUX2 has inputs connected to the outputs of the first rotation shift circuit BS1. The second multiplexer stage MUX2 also comprises outputs connected to inputs of the PEBK calculation unit. Thus, the PEBK calculation unit is configured to receive the data from the first rotation shift circuit BS1. The second multiplexer stage MUX2 is configured to deliver the data vector shifted by the first rotation shift circuit BS1 to the PEBK calculation unit, so as to transmit the data of the data vector to the various processing elements PE.

[0045] The NNA integrated circuit further comprises a WB buffer memory at the output of the PEBK calculation unit. The WB buffer memory therefore comprises inputs connected to an output of the PEBK calculation unit. The WB buffer memory is thus configured to receive the data calculated by the PEBK calculation unit. In particular, the WB buffer memory may be a storage element making it possible to store a single word of data.

[0046] The NNA integrated circuit also comprises a pruning stage PS (also referred to as a “pruning stage”). The pruning stage PS comprises inputs connected to outputs of the buffer memory WB. This pruning stage PS is configured to delete unnecessary data delivered by the PEBK calculation unit. The pruning stage PS is configured to delete certain unnecessary data generated by the PEBK calculation unit. In particular, data generated by the PEBK calculation unit is useless when the execution of the neural network has a stride greater than one.

[0047] The NNA integrated circuit also comprises a second rotation shift circuit BS2. The second rotation shift circuit BS2 has inputs connected to outputs of the pruning stage PS. The second rotation shift circuit BS2 has outputs connected to inputs of the second memory DMEM. The second rotation shift circuit BS2 is configured to shift the data vector delivered by the pruning stage PS before storing them in the second memory DMEM.

[0048] The integrated circuit NNA further comprises a control unit CTRL configured to control the various elements of the integrated circuit NNA, i.e., the shift stage SMUX, the first multiplexer stage MUX1, the first rotation shift circuit BS1, the second multiplexer stage MUX2, the calculation unit PEBK, the buffer memory WB, the pruning stage PS, the second rotation shift circuit BS2 as well as the accesses to the first memory WMEM and to the second memory DMEM. In particular, the control unit CTRL only accesses the useful data of the first memory WMEM and of the second memory DMEM.

[0049] Figure 3 illustrates an embodiment of the arrangement of the second DMEM memory. The DMEM memory comprises several data banks. Here, the DMEM memory comprises three data banks. The number of banks is greater than or equal to a parallelization capacity in width by the processing elements of the calculation unit (i.e., a parallelization on a number of positions of the same channel of a data matrix). Each bank is represented in a table having a predefined number of rows and columns. Here, the memory is configured to record data from a data matrix, for example an image or a feature map, having several channels. The data of the matrix is ​​stored in the different banks of the DMEM memory. Here, the data matrix has four rows, five columns and ten channels.Each data in the matrix has a value where c ranges from 0 to 9 and indicates the channel of that data in the matrix, x and y indicate the position of the data in the matrix, x ranges from 0 to 3 and corresponds to the row of the matrix, and y ranges from 0 to 4 and corresponds to the column of the matrix.

[0050] The matrix data is stored in groups in the different banks of the DMEM memory. In particular, each data group of a data matrix comprises data from at least one position of the data matrix and from at least one channel of the data matrix. The maximum number of data in each group is defined according to a depth-first parallelization capacity (i.e., parallelization on a certain number of channels of the matrix) of the execution of the neural network by the computing unit.

[0051] The number of processing elements PEKB corresponds to a maximum parallelization for the execution of the neural network, that is to say a parallelization in width multiplied by a parallelization on the different channels of the data. Thus, the number of processing elements PE can be equal to the number of banks of the DMEM memory multiplied by the number of channels of each bank of the DMEM memory. This maximum parallelization is generally not used all the time when executing a neural network, in particular due to the reduction of the dimensions of the layers in the depth of the neural network.

[0052] The groups are formed according to the parallelization capacity in width and depth of the computing unit. For example, the group GO includes the data at 4«, in the bank BCO, the group G1 includes the data Aqj to A^ in the bank BC1, the group G2 includes the data at Aq^ in the bank BC2.

[0053] More particularly, the data of the different channels of the matrix having the same position in the matrix are stored on the same line of the same bank. If the number of channels is greater than the number of columns of a bank, then it is not possible to store on the same line of a bank, and therefore in the same group, all the data of the different channels having the same position in the matrix. The remaining data are then stored in free lines at the end of each bank. For example, the data Aq0 to of the group G0 are stored on line #0 of the bank BCO, and the data A^o and Aæ are stored on line #6 of the bank BC2.

[0054] The first rotation shift circuit has a number of inputs equal to the number of banks of the DMEM memory and the second rotation shift circuit has a number of outputs equal to the number of banks of the DMEM memory. In this way, the first rotation shift circuit is configured to receive the data from the different banks.

[0055] The use of rotation shift circuits BS1 and BS2 allows for low-power data manipulation. Indeed, the first rotation shift circuit allows the data stored in the second memory to be simply read when this data is necessary for the execution of the neural network by the computing unit. The second rotation shift circuit allows the data generated by the computing unit to be simply written into the second memory when of the execution of the neural network. The rotation shift circuits are sized so that, for the execution of the neural network, useful data can be written into these circuits on data, as soon as the latter data is no longer useful for the execution of the neural network.

[0056] Such an arrangement of the data of the matrix in memory therefore makes it possible to simplify the manipulation of the data using the first rotation shift circuit and the second rotation shift circuit. Furthermore, such an arrangement of the data of the matrix in memory allows simple access to the DMEM memory in reading and writing.

Claims

Claims

1. Integrated circuit comprising: - a first memory (WMEM) configured to store parameters of a neural network to be executed, - a second memory (DMEM) configured to store data provided as input to the neural network to be executed or generated by this neural network, - a calculation unit (PEBK) configured to execute the neural network, - a first rotation shift circuit (BS1) between an output of the second memory (DMEM) and the calculation unit (PEBK), the first rotation shift circuit being configured to transmit the data from the output of the second memory to the calculation unit, - a second rotation shift circuit (BS2) between the calculation unit (PEBK) and the second memory (DMEM), the second rotation shift circuit being configured to deliver the data generated during the execution of the neural network by the calculation unit,- a control unit (CTRL) configured to control the calculation unit (PEBK), the first and second rotation shift circuits (BS1, BS2) as well as accesses to the first memory (WMEM) and to the second memory (DMEM).,

2. Integrated circuit according to one of claims 1, in which the calculation unit (PEBK) comprises a bank of processing elements (PE#0, ..., PE#N-1) configured to parallelize the execution of the neural network, the first rotation shift circuit being configured to transmit the data from the second memory to the different processing elements.

3. Integrated circuit according to one of claims 1 or 2 further comprising a first multiplexer stage (MUX1), the first rotation shift circuit (BS1) being connected to the second memory (DMEM) via the first multiplexer stage (MUX1), the first multiplexer stage (MUX1) being configured to deliver to the first rotation shift circuit (BS1) a data vector from the data stored in the second memory (DMEM), the first rotation shift circuit (BS1) being configured to shift the data vector of the first multiplexer stage (MUX1).

4. Integrated circuit according to claim 3, comprising a second stage multiplexer (MUX2), the computing unit (PEBK) being connected to the first rotation shift circuit (BS1) via the second multiplexer stage (MUX2), the second multiplexer stage (MUX2) being configured to deliver the data vector shifted by the first rotation shift circuit (BS1) to the computing unit (PEBK).

5. Integrated circuit according to one of claims 1 to 4, further comprising a buffer memory (WB), the second rotation shift circuit (BS2) being connected to the calculation unit (PEBK) via the buffer memory (WB), the buffer memory (WB) being configured to temporarily store the data generated by the calculation unit (PEBK) during the execution of the neural network before the second rotation shift circuit (BS2) delivers this data to the second memory (DMEM).

6. An integrated circuit according to claim 5, further comprising a pruning stage (PS) between the buffer memory (WB) and the second rotation shift circuit (BS2), the pruning stage (PS) being configured to delete certain data from among the data generated by the calculation unit (PEBK).

7. Integrated circuit according to one of claims 1 to 6 in which the second memory (DMEM) is configured to store data matrices provided as input to the neural network to be executed or generated by this neural network, each data matrix being able to have several data channels, the data of each data matrix being grouped in the second memory (DMEM) in at least one data group, the data groups being stored in different banks of the second memory (DMEM), the data of each data group being intended to be processed in parallel by the different processing elements of the calculation unit (PEBK).

8. An integrated circuit according to claim 7 wherein each data group of a data matrix comprises data from at least one position of the data matrix for at least one channel of the data matrix.

9. System on chip comprising an integrated circuit according to one of claims 1 to 8.