A DIGITAL ELECTRONIC CIRCUIT CONTAINING A MICROCONTROLLER AND SYSTOLIC SEQUENCE ACCELERATOR FOR THE USE OF ARTIFICIAL NEURAL NETWORKS IN EDGED EMBEDDED SYSTEMS.

TR202419548BActive Publication Date: 2026-06-22HAVELSAN HAVA ELEKTRONIK SANAYI VE TICARET ANONIM SIRKETI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
TR · TR
Patent Type
Patents
Current Assignee / Owner
HAVELSAN HAVA ELEKTRONIK SANAYI VE TICARET ANONIM SIRKETI
Filing Date
2024-12-17
Publication Date
2026-06-22

Smart Images

  • Figure 00000012_0000
    Figure 00000012_0000
  • Figure 00000013_0000
    Figure 00000013_0000
Patent Text Reader

Abstract

The invention relates to a systolic array accelerator digital electronic circuit that can be added as a peripheral to microcontrollers to accelerate artificial neural network computations in edge devices. Since the digital electronic circuit in question is specialized for performing computations used in artificial neural networks, it enables these operations to be performed faster and facilitates the use of real-time artificial intelligence in microcontroller-based edge embedded applications.
Need to check novelty before this filing date? Find Prior Art

Description

1 TARIFF USE OF ARTIFICIAL NEURAL NETWORKS IN END-EMBROIDERED SYSTEMS A MICROCONTROLLER AND SYSTOLIC SEQUENCE ACCELERATOR FOR DIGITAL ELECTRONIC CIRCUIT Technical Area 5 The invention is designed to accelerate artificial neural network computations in edge devices. systolic array accelerator that can be added as a peripheral to microcontrollers It is related to a digital electronic circuit. The invention is a digital electronic circuit used in artificial neural networks. Because it is specialized to perform these calculations, these processes are faster. 10 enabling this and in microcontroller-based embedded applications It provides the opportunity for the use of real-time artificial intelligence. Previous Technique Systolic arrays and microcontrollers are combined in edge neural network applications. Complex processes can be performed quickly and efficiently. Thanks to this, wearable 15 innovative solutions across a wide range of areas, from devices to autonomous unmanned vehicles Improvements can be made. Real-time performance and low power consumption thanks to the parallel structure of the systolic arrays. by ensuring consumption, artificial neural networks can be run in embedded systems. This can be achieved. Processing data at end devices enhances data security. 20 and reducing bandwidth by decreasing the need for cloud connectivity. It provides. As an application of artificial neural networks, health monitoring devices, smart sensors, and It can be applied in many fields, such as image processing. For example, wearable 2 Devices can monitor heart health in real time, smart farming sensors. They can predict diseases in advance. Autonomous drones, on the other hand, use systolic series to predict illnesses. They can overcome obstacles and pursue their goals. This technology encourages the use of artificial intelligence in edge devices, shaping the future. It lays the groundwork for smart applications. 5 Systolic arrays and microcontrollers in cutting-edge artificial intelligence applications. Before it could be used, it performed data processing and complex computational tasks. more centralized, less efficient, and generally slower methods to bring it about or, without artificial intelligence, rule-based solutions are used. Device and The large amount of data collected from sensors is sent to data centers or 10 for analysis. It has to be sent to the cloud and this data is processed on servers. This is necessary, which slows down the process and causes data security problems. It constitutes. Many processes are performed manually, and the computational models used, It is insufficient for processing today's complex datasets. Systolic 15 The integration of series and local AI applications provides a solution to these problems. By offering these features, it makes data processing faster, more efficient, and more secure. Thanks to local computing, there is no need to transfer data to central servers. This reduces processing time and bandwidth usage. Since the data is processed on the device, the risk to privacy and security is also reduced. Manual 20 Automation is replacing manual labor, resulting in faster and more accurate results. This enables systolic series to process complex datasets. It offers advanced computational models. Patent number US2024069971A1, which is included in the prior art. The document lists 25 microcontrollers to be used in edge artificial intelligence applications. 3 By integrating systolic series technology, complex procedures can be performed faster and more efficiently. It refers to a practice that has been carried out. Patent No. WO2020190807A1, which falls under the prior art. the document discusses microcontrollers to be used in edge artificial intelligence applications By integrating systolic series technology, complex procedures can be performed faster and more efficiently. 5 It refers to a practice that has been carried out. Patent number US2022413803A1, which is included in the prior art. the document discusses microcontrollers to be used in edge artificial intelligence applications By integrating systolic series technology, complex procedures can be performed faster and more efficiently. It refers to a practice that was carried out. 10 Patent number US2022413851A1, which falls under the prior art. the document discusses microcontrollers to be used in edge artificial intelligence applications By integrating systolic series technology, complex procedures can be performed faster and more efficiently. It refers to a practice that has been carried out. When existing systems and methods in technology are examined, systolic sequences are analyzed using ASIC 15. implemented as fixed MAC sequences in base-based accelerators and generally It has been understood that this is implemented in microprocessors. In microcontrollers, however, it is more common. specialized or vector tools for statistical machine learning, such as "decision trees". It has been observed that peripheral devices are used to speed up calculations. In current systolic series applications, the co-processing units are structurally 20 These processing units are used even if they are not required in the artificial intelligence application. It operates idle, consuming power. Therefore, to accelerate artificial neural network calculations on edge devices, microcontrollers can be added as peripherals and are structurally identical processing devices. A digital electronic circuit of systolic series accelerators in which units are not used 25 The need to realize the design has arisen. 4 Purposes of the Invention The goal of this invention is to accelerate artificial neural network computations in edge devices. for, systolic arrays, which can be added as a peripheral to microcontrollers. It is the implementation of an accelerator digital electronic circuit. Another purpose of this invention is to create 5 components that, by design, do not use coprocessing units. This is the implementation of a digital electronic circuit for a systolic series accelerator. Detailed Description of the Invention The electronic circuits created to achieve the purpose of this invention are shown in the attached figures. It has been shown. This figure; 10 Figure 1: Schematic of the architecture of a sample microcontroller containing a systolic array. It is the appearance. Figure 2: Schematic view of the data router architecture. The parts shown in the figure are individually numbered, and these numbers correspond to... The corresponding answers are given below. 15 1. Processor core 2. Processor bus interface 3. System data bus interface 4. External memory Block 5 SRAM 20 6. Systolic series control and configuration unit 7. Data router 8. Processing units 9. SRAM memory unit 10. Control unit 11. Switching matrix 12. Buffer system 13. Systolic series The subject of the invention is a digital electronic circuit 5. - A processor core (1), - A data bus interface (2) connected to the processor core (1), - A system bus interface (3) connected to the processor bus interface (2), - External memory (4), Block SRAM (5) and connected to the system bus interface (3) systolic series control and configuration unit (6), 10 - Block SRAM (5) and systolic array control and configuration unit (6) connected data router (7), - Processing units (8) connected to the data router (7), - SRAM memory unit (9) connected to each processing unit (8) It includes. 15 The invention concerns a digital electronic circuit, data router (7), - Block SRAM (5) and systolic array control and configuration unit (6) Control unit of connected data router (10), - Switching matrix (11) connected to the control unit (10) of the data router, 20 - buffer systems connected to the switching matrix (11) and the processing unit (8) (12), It includes. The processor core (1) used in the electronic circuit that is the subject of the invention is low power 25 using RISC-V architecture optimized for power consumption and high performance It is the microcontroller core. General purpose tasks, integer and floating point. 6 for managing processes and neural network (NN) computations It functions as a central processing unit and runs applications. Processor bus interface (2), processor core (1), system bus interface (3) acts as a bridge between them and enables data transfer. The system bus interface (3) serves as the main communication path of the system. It sees the processor core (1) external memory (4) and systolic array control and by connecting to the configuration unit (6) control, configuration, communication and data It provides the transfer. External memory (4) receives instructions from the CPU, operating system data and general It is a directly accessible memory for storing computational data. It meets the needs of 10 people. fast and moderate memory requirements with internal block SRAM (5) while being met with; SRAM memory unit (9) requires a lot of memory space in situations requiring high-capacity storage for applications It is used. Block SRAM (5) is high-speed memory reserved for the systolic array (13). This 15 Memory provides fast access to NN model weights, inputs, and outputs. It enables calculations to be performed efficiently, including network weights and intermediate settings. It stores results and input / output data. Block SRAM (5) is used to store NN parameters for data router control. It works together with the control unit (10). The control unit (10) controls the weights, biases and 20 Block SRAM (5) for holding and storing other critical information It uses. The control unit (10) performs intermediate calculations or multiple calculations as needed. It temporarily writes data requiring excessive processing back to block SRAM (5). This collaboration ensures that NN operations (7) in the data router run smoothly. It provides storage and access to the data necessary for its execution. 25 7 Systolic array control and configuration unit (6), loading of NN models and systolic series (13) including the initiation of calculations It manages the initial setup and operation. These management processes include: configuration, data flow control, and synchronization of processing units It includes processor core 5, which also receives commands and configuration data. It provides an interface with (1). Accordingly, the operation of the systolic series (13) It starts or stops. Systolic array control and configuration unit (6), control of data router It manages data flow and operations by communicating with unit (10). Systolic series Control and configuration unit (6), data router control unit (10), 10 It sends configuration commands. These commands route the data. It defines the schemas, operating modes, and execution initiation signals. data Control unit of the router (10), systolic series control and configuration to the unit (6), status updates, completion signals and sometimes the CPU It responds by sending processed data that requires intervention. In short; systolic 15 array control and configuration unit (6), configuration to data router (7) By sending commands, it controls the process of processing and directing data. is doing. The data router (7) also includes the data path of the systolic series (13). Systolic array data bus, systolic array (13) block SRAM (5) main system data bus band 20 efficient and high-speed systolic series without competing for width (13) It connects in coordination. This critical data router (7) component input in the efficient distribution of data and weights to processing units (8) and intermediaries are involved in guiding the final results. Processing units (8) direct data between the processing units (8) and the SRAM memory unit (9) 25 It manages the flow. Each processing unit (8) receives the correct data (input) at the right time. because it takes the values ​​and weights) and the data coming out of each processing unit (8) 8 the next processing unit (8) is correctly directed to the SRAM memory He is responsible for ensuring that it is recorded back to the unit (9). The processing units (8) are the cores of the systolic series (13). Basic calculations They do this. Each of them processes frequently accessed small data, such as intermediate tensor values. It has its own SRAM memory unit (9) for storage. NN’s basic 5 to speed up and make more efficient one of the operations, matrix multiplications. They are designed to bring about processing units via the systolic series data bus. (8) faster data transfer is provided between buffer systems (12) by sending intermediate data, these can be sent to other processing units (8) or later They enable the data to be redirected back to external memory (4) for calculations. 10 Buffer systems (12) require the calculation of each processing unit (8). Processing neatly ordered data packets containing inputs and weights. In short, buffer systems (12) transmit to the units (8). by feeding the units (8) with accurate data in a timely manner, the data router (7) They ensure that calculation processes are carried out smoothly. 15 SRAM memory unit (9) enables processing units (8) to operate at high speed and efficiency. by guaranteeing the operation of systolic series (13) in artificial intelligence applications They ensure its success. Input data, partial sums, result data. Block SRAM (5) is frequently used to store frequently accessed data such as They reduce the need for data retrieval and increase efficiency. 20 The control unit of the data router (10), control and data in the data router (7) It is the unit that undertakes the processing tasks. This unit controls the systolic series and... by receiving commands from the configuration unit (6), operating mode and data flow Routing paths and switching matrix with buffer according to requirements It configures the system. At the same time, it makes the processing units (8) able to operate. basic prerequisites such as formatting or scaling the incoming data to retrieve it. It also performs processing operations. This single-unit approach handles complex tasks. 9 simplifies tasks and enables the data router (7) to work efficiently. It enables it to function. The switching matrix (11) is the dynamic switching of different layers and operations of the NN. By directing the data flow according to their requests, the entire system works efficiently. This matrix, consisting of dynamically configurable keys, provides 5 By offering flexible routing schemes, data can be efficiently routed to processing units (8) It guarantees that it will reach the complex way. Switching matrix (11) Even in these scenarios, managing the loops allows for the re-processing of intermediate data. This makes it possible to redirect operations. In this way, NN operations can be performed smoothly. This is done in the following way. 10 The buffer system (12) provides synchronized data flow, enabling efficient systolic series (13) It supports parallel processing. It acts as a temporary storage unit. Buffer systems (12) keep data packets until they reach the processing units (8) It holds. In this way, all processing units (8) receive data at the same time and synchronize. It processes data in a certain way. To ensure efficient data flow, every 15 The processing unit (8) has its own special buffer systems (12). This system, It guarantees the smooth functioning of the systolic series (13). The invention describes a digital electronic circuit that uses four different coprocessors instead of coprocessors. Types of processing units (8) were used. Each type of processing unit (8) has a different tensor. It is customized to perform the following types of operations: 20 - Singular units such as logarithms, their multiplicative inverses, and activation functions. transactions, - Binary operations such as addition and multiplication, - Reducers that change tensor dimensions such as total, maximum, minimum. transactions, 25 - Permutations, transpositions, expansions, etc., that modify the tensor data shape. transactions It is defined as such. Data router (7) and systolic series control and configuration unit in the invention. (6) the number and type of processing units required for artificial intelligence application (8) by activating and appropriately directing the data flow to the processor core (1) According to a calculation to be made on it, more efficient and faster processing is possible. 5 It provides different units of operation (8) for the types of operations. This also saves physical space. The processor takes the processing load onto itself. obtaining, application latency and memory access bandwidth This reduces their requirements and enables lower power consumption.

Claims

11 REQUESTS 1. The invention aims to accelerate artificial neural network computations in edge devices. systolic array that can be added as a peripheral to microcontrollers It relates to an accelerator digital electronic circuit, - A processor core (1), 5 - A data bus interface (2) connected to the processor core (1), - A system bus interface (3) connected to the processor bus interface (2), - External memory (4), Block SRAM (5) and connected to the system bus interface (3) systolic series control and configuration unit (6), - Block SRAM (5) and systolic array control and configuration unit (6) 10 connected data router (7), - Processing units (8) connected to the data router (7), - SRAM memory unit (9) connected to each processing unit (8) includes and data router (7), - Block SRAM (5) and systolic array control and configuration unit (6) 15 Control unit of connected data router (10), - Switching matrix (11) connected to the control unit (10) of the data router, - buffer connected to the switching matrix (11) and processing units (8) systems (12), It is characterized by its inclusion. 20 2. A digital electronic circuit like the one in Claim 1, capable of four different types of processing. units (8) are used and each type of processing unit (8) has different tensors. It is characterized by being customized for performing specific types of operations. is being done.