Configurable cascade circuit for efficient in-memory processing

By designing a configurable cascade circuit, the problems of high configuration delay and insufficient resource coordination in the in-memory computing circuit are solved, and the flexible configuration of computing logic and data flow paths are realized, and an efficient computing-cache-transmission integrated pipeline is built, suitable for edge-end multi-scenario applications.

CN120234288APending Publication Date: 2025-07-01XIDIAN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510270557.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing in-memory computing circuits have problems such as high configuration instruction analysis and hardware response delays, the cascade relationship between computing units needs to be pre-statically defined, the lack of deep coordination between storage and computing resources, and the frequent write-back of off-chip storage in intermediate cached data, resulting in bottlenecks in real-time multi-scenario applications at edge-end high-level real-time.

Method used

A configurable cascade circuit for efficient in-memory processing is designed, including a state machine module, a cascade instruction allocator unit, a configuration register unit and a configurable cascade calculation engine. The cascade instructions are dynamically allocated through bit domain segmentation and address mapping mechanisms. The decoder unit parses the operation block type and generates control words. The bus drives the in-memory computing unit is distributed, and the intermediate data is stored using dual BRAM dynamic cache strategy.

Benefits of technology

It realizes flexible configuration of computing logic and data flow paths, and builds an integrated computing-cache-transmission pipeline, overcomes the problems of hardware resource solidification and poor task adaptability, eliminates the frequent handling of intermediate data, and provides a high-adaptive hardware foundation for edge-end multi-scenario applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234288A_ABST
    Figure CN120234288A_ABST
Patent Text Reader

Abstract

The invention discloses a configurable cascade circuit for efficient in-memory processing, which comprises an integrated state machine module, a cascade instruction distributor unit, a configuration register unit, a configurable cascade calculation engine and an on-chip storage unit, and forms a complete in-memory calculation processing chain. The state machine dynamically adjusts the circuit operation state; the cascade instruction distributor unit decomposes a cascade instruction into discrete operation blocks and loads the discrete operation blocks into a plurality of registers in the configuration register unit; the decoder unit generates a control word containing the serial number of the in-memory computing unit and the channel configuration parameters of the in-memory computing unit; and the configurable cascade computing engine enables the corresponding in-memory computing unit according to the control word, and after the in-memory computing unit is driven to complete a specified computing task, intermediate cache data is written into the on-chip storage unit in real time by adopting a double-BRAM dynamic cache strategy. And circularly controlling the iteration work of the calculation unit in the memory until the cascade termination signal triggers and outputs a final calculation result. According to the invention, a calculation-caching-transmission integrated assembly line is constructed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of in-memory computing circuits, and particularly relates to a configurable cascaded circuit for efficient in-memory processing. Background Art

[0002] With the rapid development of artificial intelligence and edge computing technologies, the "memory wall" problem caused by frequent data movement in traditional computing circuits has become increasingly prominent. In-memory computing technology embeds computing units into the memory array and directly performs operations at the data storage location, becoming an important direction to break through the energy efficiency bottleneck.

[0003] In response to the above problems, some configurable in-memory computing architectures have been proposed in the industry, but there are still significant deficiencies in their configuration processes: First, the configuration instruction parsing and hardware response delays are relatively high, resulting in low efficiency of dynamic task switching; Second, the cascading relationship between computing units needs to be statically defined in advance and cannot construct a computing pipeline according to instructions; Finally, there is a lack of deep coordination between storage and computing resources, and the intermediate cached data still needs to be frequently written back to off-chip storage, failing to fully utilize the advantages of in-memory computing. These defects make the existing circuits face severe bottlenecks in high-real-time and multi-scenario applications at the edge. Summary of the Invention

[0004] In order to solve the above problems existing in the prior art, the present invention provides a configurable cascaded circuit for efficient in-memory processing. The technical problems to be solved by the present invention are realized through the following technical solutions:

[0005] An embodiment of the present invention provides a configurable cascaded circuit for efficient in-memory processing, and the configurable cascaded circuit includes:

[0006] On-chip memory units, including a first BRAM, a second BRAM, and a third BRAM; the first BRAM is used to store externally input cascading instructions, and the cascading instructions include a cascading process termination operation block and / or several in-memory computing unit operation blocks; the second BRAM and the third BRAM are used to alternately store intermediate cached data; the second BRAM is also used to store externally input initial data;

[0007] A state machine module, which is used to make the state machine module transfer from the waiting state to the cascading instruction pre-configuration state according to an externally input start signal;

[0008] A cascading instruction distributor unit, which is used to read the cascading instructions from the first BRAM in the cascading instruction pre-configuration state, divide the cascading instructions into multiple sub-instructions through a bit field segmentation mechanism, and allocate each sub-instruction to a corresponding configuration register unit through an address mapping mechanism. After all sub-instructions are allocated, an instruction transfer completion signal is generated to make the state machine module transfer from the cascading instruction pre-configuration state to the computing engine working state;

[0009] A decoder unit, which, when the computing engine is in the working state, reads the cascade process termination operation block or the in-memory computing unit operation block corresponding to the sub-instruction from each configuration register unit respectively: when the read result is the in-memory computing unit operation block, it parses and outputs the first serial number corresponding to the in-memory computing unit operation block and the channel configuration parameters; when the read result is the cascade process termination operation block, it outputs a cascade termination signal to cause the state machine module to transfer from the working state of the computing engine to the working completion state;

[0010] A configurable cascade computing engine, which, when the computing engine is in the working state, determines the enabled in-memory computing units according to the first serial number in the control word, configures the enabled in-memory computing units according to the channel configuration parameters in the control word, and uses the configured in-memory computing units to calculate and output intermediate cache data for the data read from the second BRAM or the third BRAM, so as to store the intermediate cache data in the second BRAM or the third BRAM; wherein, the data read from the second BRAM for the first time is the initial data, and the data read from the second BRAM or the third BRAM later is the intermediate cache data.

[0011] In an embodiment of the present invention, the cascade instruction allocator unit includes an instruction address counter and an instruction allocator; wherein,

[0012] The instruction address counter is used to accumulate and count the second serial number of the register allocated with the sub-instruction, and generate an instruction transmission completion signal when the allocated second serial number reaches a preset allocation serial number threshold;

[0013] The instruction allocator is used to divide the cascade instruction into multiple sub-instructions according to a preset bit threshold through a bit field segmentation mechanism, and allocate each sub-instruction to the register corresponding to the second serial number through an address mapping mechanism.

[0014] In an embodiment of the present invention, the configuration register unit includes a register counter and a plurality of registers; wherein,

[0015] Each register is used to store the cascade process termination operation block or the in-memory computing unit operation block in the corresponding sub-instruction;

[0016] The register counter is used to accumulate and count the second serial number of the read register.

[0017] In an embodiment of the present invention, the decoder unit includes an operation block parsing module and a parameter discrimination module; wherein,

[0018] The operation block parsing module is used to read the cascaded process termination operation block or the in-memory computing unit operation block from the register corresponding to the second serial number read. When the read result is the in-memory computing unit operation block, the in-memory computing unit operation block is separated into an operation code and a parameter field;

[0019] The parameter discrimination module is used to judge the operation block type according to the operation code. When the operation block type is the in-memory computing unit operation block, the corresponding parameter field is parsed to separate the input channel parameter and the output channel parameter, and a control word including the first serial number of the corresponding in-memory computing unit operation block, the input channel parameter, and the output channel parameter is output. When the operation block type is the cascaded process termination operation block, a cascaded termination signal is generated.

[0020] In an embodiment of the present invention, the configurable cascaded computing engine includes a distributed configuration bus and a plurality of in-memory computing units; wherein,

[0021] The distributed configuration bus is used to receive the control word output by the decoder unit and broadcast it to each in-memory computing unit;

[0022] Each in-memory computing unit is used to determine the enabled in-memory computing unit according to the first serial number in the control word, configure the enabled in-memory computing unit according to the input channel parameter and the output channel parameter, and calculate the data read from the second BRAM or the third BRAM by using the configured in-memory computing unit to output intermediate cache data, so as to store the intermediate cache data in the second BRAM or the third BRAM.

[0023] In an embodiment of the present invention, each in-memory computing unit is further used to generate an in-memory computing completion signal after calculating and outputting the intermediate cache data, so as to control the register counter to continue to accumulate and count the second serial number of the read register.

[0024] In an embodiment of the present invention, the configurable cascaded computing engine further includes a first multiplexer, a second multiplexer, and a first two-way selector; wherein,

[0025] The first multiplexer is used to select and output the intermediate cache data calculated by the currently determined enabled in-memory computing unit from all in-memory computing units;

[0026] The second multiplexer is used to select and output the in-memory computing completion signal generated by the currently determined enabled in-memory computing unit from all in-memory computing units;

[0027] The first two-way selector is used to select to store the intermediate cache data output by the first multiplexer in the second BRAM or the third BRAM.

[0028] In one embodiment of the present invention, when the storage selection signal is 1, intermediate cache data is stored in the second BRAM and intermediate cache data is read from the third BRAM; when the storage selection signal is 0, intermediate cache data is read from the second BRAM and intermediate cache data is stored in the third BRAM;

[0029] The first two-way selector is further configured to select to store the intermediate cache data output by the first multiplexer in the second BRAM or the third BRAM according to the storage selection signal.

[0030] In one embodiment of the present invention, the on-chip storage unit further includes a second two-way selector and a third two-way selector; wherein,

[0031] The second two-way selector is configured to read and output the previously stored intermediate cache data from the second BRAM or the third BRAM according to the in-memory computing completion signal and the storage selection signal;

[0032] The third two-way selector is configured to read and output the last stored intermediate cache data from the second BRAM or the third BRAM according to the cascade termination signal and the storage selection signal.

[0033] In one embodiment of the present invention, the state machine module is further configured to cause the state machine module to transition from the work completion state to the waiting state according to the waiting signal.

[0034] Advantages of the present invention:

[0035] The configurable cascade circuit for efficient in-memory processing proposed by the present invention is a configurable cascade in-memory processing circuit that supports configuration and has a high processing speed. Through the innovative instruction-driven configurable cascade computing engine design, it breaks through the limitations of the hardware solidification of traditional in-memory computing circuits and realizes that the computing logic and data flow path can be configured according to tasks; the four-layer collaborative design of the configurable computing engine, cascade instruction distributor unit, configuration register unit, and on-chip storage unit constructs an "integration pipeline of computing - caching - transmission". Through the resource configuration mechanism driven by cascade instructions and in-depth collaboration with heterogeneous computing engines, all computing, caching, and transmission tasks are directly completed inside the cascade circuit, overcoming the bottlenecks such as hardware resource solidification, poor task adaptability, and low collaborative efficiency of multi-stage pipelines in existing in-memory computing circuits, eliminating the frequent transfer of intermediate data, and providing a highly adaptable hardware foundation for multi-scenario applications at the edge.

[0036] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings

[0037] Figure 1It is a schematic structural diagram of a configurable cascaded circuit for efficient in-memory processing provided by an embodiment of the present invention;

[0038] Figure 2 It is a schematic implementation diagram of a state machine module in a configurable cascaded circuit provided by an embodiment of the present invention;

[0039] Figure 3 It is a schematic structural diagram of a cascaded instruction distributor unit provided by an embodiment of the present invention;

[0040] Figure 4 It is a schematic structural diagram of a cascaded instruction structure of an external input provided by an embodiment of the present invention;

[0041] Figure 5 It is a schematic structural diagram of a configuration register unit and a decoder unit provided by an embodiment of the present invention;

[0042] Figure 6 It is a schematic structural diagram of a configurable cascaded computing engine provided by an embodiment of the present invention;

[0043] Figure 7 It is a schematic structural diagram of an on-chip storage unit in a configurable cascaded circuit provided by an embodiment of the present invention. Detailed implementation manners

[0044] The following further describes the present invention in detail with specific embodiments, but the implementation manners of the present invention are not limited thereto.

[0045] Please refer to Figure 1 , an embodiment of the present invention provides a configurable cascaded circuit for efficient in-memory processing, and the configurable cascaded circuit includes:

[0046] An on-chip storage unit, including a first BRAM, a second BRAM, and a third BRAM; the first BRAM is used to store a cascaded instruction of an external input, and the cascaded instruction includes a cascaded process termination operation block and / or a plurality of in-memory computing unit operation blocks; the second BRAM and the third BRAM are used to alternately store intermediate cache data; the second BRAM is further used to store the initial data of the external input;

[0047] A state machine module, configured to make the state machine module transition from a waiting state to a cascaded instruction pre-configuration state according to a start signal of an external input;

[0048] A cascaded instruction distributor unit, configured to, in the cascaded instruction pre-configuration state, read the cascaded instruction from the first BRAM, divide the cascaded instruction into multiple sub-instructions through a bit field splitting mechanism, and allocate each sub-instruction to a corresponding configuration register unit through an address mapping mechanism, and generate an instruction transmission completion signal after all sub-instructions are allocated, so that the state machine module transitions from the cascaded instruction pre-configuration state to a computing engine working state;

[0049] A decoder unit, which, when the computing engine is in the working state, reads the cascade process termination operation block or the in-memory computing unit operation block corresponding to the sub-instruction from each configuration register unit respectively: if the read result is the in-memory computing unit operation block, it parses and outputs a control word including the first serial number corresponding to the in-memory computing unit operation block and the channel configuration parameters; if the read result is the cascade process termination operation block, it outputs a cascade termination signal to cause the state machine module to transfer from the computing engine working state to the working completion state;

[0050] A configurable cascade computing engine, which, when the computing engine is in the working state, determines the enabled in-memory computing units according to the first serial number in the control word, configures the enabled in-memory computing units according to the channel configuration parameters in the control word, and uses the configured in-memory computing units to calculate and output intermediate cache data for the data read from the second BRAM or the third BRAM, so as to store the intermediate cache data in the second BRAM or the third BRAM; wherein, the data read from the second BRAM for the first time is the initial data, and the data read from the second BRAM or the third BRAM later is the intermediate cache data.

[0051] Next, each part will be introduced in detail.

[0052] The state machine module in the embodiment of the present invention is as Figure 2 shown, including 4 states: waiting state (STATE_IDLE), cascade instruction pre-configuration state (STATE_PRE), computing engine working state (STATE_WORK), and working completion state (STATE_DONE). In the waiting state: the circuit is initialized and transfers to the cascade instruction pre-configuration state after receiving the externally received start signal (start_i). In the instruction pre-configuration state: the first BRAM ( Figure 1 denoted as Bank I in) of the on-chip storage unit receives the externally input cascade instructions through the high-speed interface, and after being distributed by the cascade instruction distributor unit, they are stored in the configuration register unit. After all the distributions are completed, an instruction transfer completion flag (inst_done) is generated and it transfers to the computing engine working state. In the computing engine working state: the decoder unit fetches the instructions stored in the configuration register unit and parses them, and calls the enabled in-memory computing units in the configurable cascade computing engine according to the parsed channel configuration parameters, thereby completing one calculation. This process is repeated until the cascade termination signal (work_done) is detected, and then it transfers to the working completion state. In the working completion state: it is responsible for post-processing of data. For example, the hardware delay unit generates a waiting signal (wait_done) and automatically transfers to the waiting state.

[0053] Further, the cascade instruction distributor unit in the embodiment of the present invention is as Figure 1 and Figure 3As shown in the figure, it includes an instruction address counter (denoted as Icnt in the figure) and an instruction allocator (denoted as Case[i] in the figure); among them, the instruction address counter is used to accumulate and count the second serial numbers of the registers allocated by sub-instructions, and generate an instruction transmission completion signal when the allocated second serial number reaches a preset allocation serial number threshold; the instruction allocator is used to divide the cascaded instruction into multiple sub-instructions according to a preset bit threshold through a bit field segmentation mechanism, and allocate each sub-instruction to the register corresponding to the second serial number through an address mapping mechanism. In the embodiment of the present invention, the register unit is configured as Figure 1 and Figure 5 As shown in the figure, it includes a register counter (Ncnt) and several registers (REG); among them, each register is used to store the cascaded process termination operation block or the in-memory computing unit operation block in the corresponding sub-instruction; the register counter is used to accumulate and count the second serial numbers of the read registers. More specifically:

[0054] The embodiment of the present invention adopts a 128-bit cascaded instruction packet architecture as Figure 4 As shown in the figure, each cascaded instruction packet is composed of 8 16-bit operation blocks. The operation blocks adopt a dual-mode classification mechanism, that is, they can be in-memory computing unit operation blocks or cascaded process termination operation blocks, and each in-memory computing unit operation block selects a composite structure of 4-bit operation code (opcode) and 12-bit parameter field (param), where the operation code is used to identify the operation block type, and the parameter field realizes the configuration function of the in-memory computing unit operation block.

[0055] The design of the cascaded instruction in the embodiment of the present invention is shown in Table 1. Among them, in-memory computing unit 1 and in-memory computing unit 2 are in-memory computing unit operation blocks, which support multi-modal computing unit configuration based on parameters. The cascaded termination identification unit is a cascaded process termination operation block. When this operation block is recognized, it means that the entire configurable cascaded circuit has completed its work.

[0056] Table 1 Design of Cascaded Instructions

[0057]

[0058]

[0059] The embodiment of the present invention adopts a 64×16-bit register bank storage structure, but is not limited to 64×16-bit registers, and is specifically designed according to actual needs. For example, it can also be an 8×16-bit register bank. In the cascaded instruction pre-configuration state, batch configuration loading is realized through the high-speed interface of the first BRAM, and the instructions are dynamically allocated through the cascaded instruction allocator unit. The core working mechanism of the cascaded instruction allocator unit includes:

[0060] 1. Instruction Bitfield Division: Dynamically divide the 128-bit cascaded instruction, and use bitfield decoding to divide the cascaded instruction into multiple sub-instructions according to a preset bit threshold. For example, Figure 1 It is shown as follows: The preset bit threshold is 16, and a 16-bit configuration segment is used as a sub-instruction. The 128-bit cascaded instruction can be divided into 8 independent 16-bit configuration segments. For example, the 128-bit cascaded instruction is denoted as I[127:0], and the 8 independent 16-bit configuration segments after division are I[15:0], I[31:16], ……, I[127:112] respectively. Each 16-bit configuration segment corresponds to an operation block, and this operation block may be an in-memory computing unit operation block or a cascaded process termination operation block.

[0061] 2. Instruction Address Mapping: Based on the dynamic addressing logic of the instruction address counter i value, select the corresponding operation block and store it in the register in the configuration register unit. That is, after determining the value of i, that is, determining the serial number (the second serial number) of the register to which the divided sub-instruction will be allocated, accurately write the 8 groups of 16-bit configuration segments after separation into the register address range [8i, 8(i + 1)) corresponding to this serial number, so as to realize the accurate mapping between the configuration segment and the register.

[0062] 3. Send Instruction Transmission Completion Signal: The value i in the instruction address counter is incremented by 1 every cycle in the instruction pre-configuration state. When the accumulated value i reaches the preset allocation serial number threshold (i = 7), an instruction transmission completion flag (inst_done) is generated, and the calculation engine working state is entered.

[0063] It should be noted that the value i here is set according to the actual external input cascaded instruction situation. Figure 1 This shows a case where the cascaded instruction is 128 bits, and the sub-instruction after division is 16 bits. Since each register is a 16-bit register, only 8 accumulations are required. Since the value i starts from 0, the maximum value is 7, that is, the preset allocation serial number threshold is 7, and correspondingly, it occupies registers REG[0] to REG[7] in the configuration register unit. Figure 1 What is shown in [the figure] are 64 registers REG[0] to REG

[63] . After the cascaded instruction is allocated, in all registers: the operation block stored in the last register is the cascaded process termination operation block, and the operation blocks stored in other registers are all in-memory computing unit operation blocks.

[0064] Furthermore, in the embodiment of the present invention, the decoder unit is as Figure 5As shown in the figure, it includes an operation block parsing module and a parameter discrimination module. Among them, the operation block parsing module is used to read the cascaded process termination operation block or the in-memory computing unit operation block from the register corresponding to the second serial number. When the read result is the in-memory computing unit operation block, the in-memory computing unit operation block is separated into an operation code and a parameter field; the parameter discrimination module is used to determine the operation block type according to the operation code. When the operation block type is the in-memory computing unit operation block, the corresponding parameter field is parsed to separate the input channel parameter and the output channel parameter, and a control word including the first serial number corresponding to the in-memory computing unit operation block, the input channel parameter, and the output channel parameter is output. When the operation block type is the cascaded process termination operation block, a cascaded termination signal is generated. More specifically:

[0065] In the embodiment of the present invention, the operation block R[n] in the nth register is selected according to the value n of the register counter and input into the decoder unit. The decoder unit includes an operation block parsing module and a parameter discrimination module, and its core working mechanism is as follows:

[0066] 1. Operation block parsing module: Parse the 16-bit operation block (cascaded process termination operation block or in-memory computing unit operation block) read from the register corresponding to the second serial number each time. When the read result is the in-memory computing unit operation block, the operation block is separated into a 4-bit high operation code and a 12-bit low parameter field by bit field splitting.

[0067] 2. Parameter discrimination module: Determine the operation block type (in-memory computing unit operation block / cascaded process operation block) according to the operation code. In the embodiment of the present invention, when the operation code is 0x0 or 0x1, it is determined that the operation block is the in-memory computing unit operation block; when the operation code is 0xf, it is determined that the operation block is the cascaded process termination operation block.

[0068] At the same time, the parameter discrimination module parses the parameter field according to the operation block type: When the operation block is the in-memory computing unit operation block, read its parameter field from the in-memory computing unit operation block and separate the [7:4] and [3:0] bits, and respectively parse them into the input channel parameter C i and the output channel parameter C o through a selector; when the operation block is the cascaded process termination operation block, there is no need to parse the parameter field, and only a cascaded termination signal needs to be sent to indicate the completion of the cascaded operation.

[0069] Therefore, when the decoder unit parses the in-memory computing unit operation block, the decoder unit outputs the serial number OP (the first serial number) corresponding to the in-memory computing unit and its input channel parameter C i and the output channel parameter C oWhen the control word is such that the state machine continues to maintain the working state of the computing engine. When the decoder unit parses the cascade process termination operation block, it outputs a cascade termination signal (work_done), triggering the switching of the hardware-level state machine to enter the working completed state.

[0070] Furthermore, in the embodiments of the present invention, the configurable cascade computing engine is as Figure 1 and 6 shown, including a distributed configuration bus and a number of in-memory computing units; among them,

[0071] The distributed configuration bus is used to receive the control word output by the decoder unit and broadcast it to each in-memory computing unit; each in-memory computing unit is used to determine the enabled in-memory computing unit according to the first serial number in the control word, and configure the enabled in-memory computing unit according to the input channel parameters and output channel parameters in the control word, and use the configured in-memory computing unit to calculate the data read from the second BRAM or the third BRAM to output intermediate cache data, so as to store the intermediate cache data in the second BRAM or the third BRAM. At the same time, in the embodiments of the present invention, each in-memory computing unit is also used to generate an in-memory computing completed signal after calculating and outputting the intermediate cache data, so as to control the register counter to continue to accumulate and count the second serial number of the read registers. More specifically:

[0072] In the working stage of the computing engine in the embodiments of the present invention, dynamic computing configuration is realized through the following process:

[0073] 1. Distributed configuration bus: Broadcast the control word containing the first serial number of the in-memory computing unit operation block and the channel configuration parameters (input channel parameter C i , output channel parameter C o ) output by the decoder unit to each in-memory computing unit through the distributed configuration bus, determine the enabled in-memory computing unit through the first serial number of the in-memory computing unit operation block, and select to make the enable signal en of the in-memory computing unit to be enabled valid from all the enable signals (en1, en2,..., en x ) of the in-memory computing units, keep the enable signals en of other in-memory computing units invalid, enable the in-memory computing unit, and call the in-memory computing unit for calculation.

[0074] 2. In-memory computing unit configuration: The in-memory computing unit adopts a parameterized configurable structure design, integrates a parameter configuration interface, and the channel configuration parameters (C i1 , C i2 ,..., C ix , and C o1 , C o2 ,..., C ox)After being passed into the in-memory computing unit, the computing mode of the in-memory computing unit can be configured, and then the input data (fea i1 、fea i2 、……、fea ix ) is calculated.

[0075] In the embodiment of the present invention, the configurable cascaded computing engine is as shown in Figure 1 and Figure 6 , and further includes a first multiplexer, a second multiplexer, and a first two-way selector; wherein,

[0076] The first multiplexer is used to select and output the intermediate cache data calculated by the currently determined enabled in-memory computing unit from all in-memory computing units; the second multiplexer is used to select and output the in-memory computing completion signal generated by the currently determined enabled in-memory computing unit from all in-memory computing units; the first two-way selector is used to select to store the intermediate cache data output by the first multiplexer into the second BRAM or the third BRAM. More specifically:

[0077] In the embodiment of the present invention, when the operation of the in-memory computing unit is completed once, the in-memory computing unit issues a work completion signal done. The first multiplexer receives the intermediate cache data (fea o1 、fea o2 、……、fea ox ) output by all in-memory computing units, where x represents the number of in-memory computing units in the configurable cascaded computing engine, selects and outputs the intermediate cache data (fea o ) calculated by the currently determined enabled in-memory computing unit, and stores the intermediate buffer data in the second BRAM (denoted as fea a ) or the third BRAM (denoted as fea b ) according to the storage selection signal for the next in-memory calculation; the second multiplexer receives the work completion signals (done1, done2, ……, done x ) output by all in-memory computing units, selects and outputs the work completion signal output by the currently determined enabled in-memory computing unit as the in-memory computing completion signal r_done, and controls the value of n in the register counter to be incremented by 1 through the in-memory computing completion signal r_done. If the system is still in the computing engine working state, the decoder unit continues to read from the configuration register unit to start a new round of computing tasks. Through this pipeline scheduling strategy, the in-memory computing unit is cyclically called until the pipeline work is completed when the cascade termination signal is detected.

[0078] Further, in an embodiment of the present invention, when the storage selection signal (not marked in the figure) is 1, the intermediate cache data is stored in the second BRAM and the intermediate cache data is read from the third BRAM; when the storage selection signal is 0, the intermediate cache data is read from the second BRAM and the intermediate cache data is stored in the third BRAM; the first two-way selector is further configured to select to store the intermediate cache data output by the first multiplexer in the second BRAM or the third BRAM according to the storage selection signal.

[0079] In the embodiment of the present invention, the on-chip storage unit is as Figure 1 and Figure 7 shown, and further includes a second two-way selector and a third two-way selector; wherein, the second two-way selector is configured to read and output the previously stored intermediate cache data from the second BRAM or the third BRAM according to the in-memory computing completion signal and the storage selection signal; the third two-way selector is configured to read and output the last stored intermediate cache data from the second BRAM or the third BRAM according to the cascade termination signal and the storage selection signal. More specifically:

[0080] The on-chip storage unit in the embodiment of the present invention includes a first BRAM (denoted as Bank I) for storing cascade instructions, and a second BRAM (denoted as Bank A) and a third BRAM (denoted as Bank B) for storing data. The cascade instructions (denoted as inst_in in the figure) stored in Bank I are externally input, and in the example of the present invention, they are input through the AXI (Advanced eXtensible Interface, high-performance extended bus interface) bus; both Bank A and Bank B are used to store the data required by the in-memory computing unit, where Bank A is used to store the intermediate cache data and the initial data (denoted as input in the figure) of the external input through the AXI bus, and Bank B is only used to store the intermediate cache data.

[0081] The on-chip storage unit in the embodiment of the present invention has an on-chip dynamic cache mechanism and a data selection and output mechanism. In the example of the present invention, when the storage selection signal is 1, the intermediate cache data is stored in the second BRAM and the intermediate cache data is read from the third BRAM; when the storage selection signal is 0, the intermediate cache data is read from the second BRAM and the intermediate cache data is stored in the third BRAM, that is: when the cascade termination signal is not detected, according to the storage selection signal, the intermediate cache data is directly stored in Bank A or Bank B integrated in the on-chip storage unit through the ping-pong cache mechanism; after the cascade termination signal is detected, the last stored calculation result is directly read from Bank A or Bank B according to the storage selection signal and the calculation result is output to the outside (denoted as output in the figure).

[0082] 1. Dynamic management mechanism of on-chip cache: The core of this mechanism is to temporarily store intermediate calculation results, i.e., intermediate cache data, when the cascade calculation is not completed, to avoid data overflow or repeated calculation. Two independent storage areas, Bank A and Bank B, are integrated on the chip, and the capacity of each storage area matches the maximum size of the intermediate feature map. In the data writing mode (when the storage selection signal is 1), the intermediate cache data generated by the current calculation is written into Bank A in real time, while the next in-memory calculation unit reads the previously stored intermediate buffer data from Bank B (data reading mode, when the storage selection signal is 1). After an in-memory calculation unit finishes its calculation, the read and write targets are automatically switched through a hardware signal. Bank A switches to the data reading mode, and Bank B switches to the data writing mode to perform similar operations as above. The difference is that the storage selection signal becomes 0, and this signal controls the read and write operations of storage areas Bank A and Bank B.

[0083] 2. Data output process triggered by cascade termination signal: The core of this mechanism is the data output logic after the system finishes working. After the decoder unit parses the cascade process termination operation block, it immediately freezes the pipeline and stops the scheduling of the in-memory calculation unit. At the same time, it selects the currently active storage area (the storage selection signal is 1 and it is in the data writing mode, indicating that the latest storage operation has been performed), Bank A or Bank B, as the final output data. This storage area contains the calculation results written by the last in-memory calculation unit, and outputs this calculation result to the outside.

[0084] After the above design, the configurable cascaded circuit for efficient in-memory processing is designed. The circuit proposed in the present invention is based on the modular hardware design concept, integrating an embedded state machine module, a cascaded instruction distributor unit, a configuration register unit, a configurable cascaded computing engine, and an on-chip storage unit to form a complete in-memory computing processing chain. The operating state of the circuit proposed in the present invention is dynamically regulated by the embedded state machine, and successively executes stages such as system initialization, instruction stream parsing, in-memory computing task scheduling, and data output control. The cascaded instruction distributor unit decomposes the original cascaded instructions into discrete operation blocks through a bit field segmentation and address mapping mechanism and loads them into multiple registers in the configuration register unit; the operation block parsing module and parameter discrimination module designed by the decoder unit synchronously parse the operation code and parameter field in the operation block to generate a control word containing the in-memory computing unit number and channel configuration parameters. The configurable cascaded computing engine distributes the control word to all in-memory computing units through a parameterized distributed configuration bus. After driving them to complete the specified computing tasks, it adopts a dual-BRAM dynamic caching strategy to write the intermediate cached data into the on-chip storage unit in real time. This process controls the in-memory computing units to continuously iterate and compute through hardware-level loops until the cascaded termination signal triggers the output of the final computing result. Finally, the computing result stored in the second BRAM or the third BRAM for the last time is output according to the AXI bus protocol through a high-speed bus, completing the end-to-end computing process.

[0085] To verify the effectiveness of the configurable cascaded circuit for efficient in-memory processing provided by the embodiments of the present invention, the following experiments are carried out for verification.

[0086] The present invention is compared with typical processors such as CDLUT (Jia S.C. Research on Efficient Super-Resolution Reconstruction Technology for Single-Frame Images[D]. Xidian University, 2024), IDLA (Gao P, Huang Z, Ye H, et al. IDLA: An Instruction-based Adaptive CNN Accelerator[C] / / 2020 IEEE 15th International Conference on Solid-State & Integrated Circuit Technology (ICSICT). IEEE, 2020: 1-3.), and TinyNPU-F (Xu K.R. Design and Research of High-Energy-Efficient Deep Neural Network Acceleration Chip[D]. Xidian University, 2022.), as shown in Table 2.

[0087] Table 2 Resource Consumption of Different Processors

[0088]

[0089] Experimental results show that the circuit proposed in the present invention shows significant technical advantages in hardware resource efficiency and system flexibility compared with other processors. The circuit proposed in the present invention has good computing universality and scalability: through the design of combining cascaded instruction allocation units, configuration register units, decoder units and configurable cascade computing engines, it supports the free combination of computing units including but not limited to convolution operators, convolution lookup table operators, etc., to build programmable in-memory computing hardware circuits. This feature effectively breaks through the functional limitations of existing processors on a single operator, and provides a new implementation path for the design of in-memory processing circuits for heterogeneous computing. Experimental data show that the circuit proposed in the present invention has the advantages of high hardware flexibility and low resource requirements, and provides a feasible hardware solution for the deployment of edge smart devices.

[0090] To summarize, the configurable cascade circuit for efficient in-memory processing proposed in an embodiment of the present invention is a configurable cascade in-memory processing circuit that supports configuration and has a high processing speed. It breaks through the limitations of the hardware solidification of traditional in-memory computing circuits through the innovative instruction-driven configurable cascade computing engine design, and realizes that the computing logic and data flow path can be configured according to the task; the four-layer collaborative design of the configurable computing engine, the cascade instruction distributor unit, the configuration register unit, and the on-chip storage unit constructs an integrated "computing-caching-transmission" pipeline, and through the cascade instruction-driven resource configuration mechanism and the deep collaboration with the heterogeneous computing engine, all computing, caching, and transmission tasks are completed directly in the cascade circuit, overcoming the bottlenecks of hardware resource solidification, poor task adaptability, and low multi-stage pipeline collaboration efficiency in existing in-memory computing circuits, eliminating the frequent transfer of intermediate data, and providing a highly adaptable hardware foundation for multi-scenario applications on the edge.

[0091] In the description of the present invention, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.

[0092] Although the present invention is described herein in conjunction with various embodiments, in the process of implementing the claimed invention, those skilled in the art may understand and implement other variations of the disclosed embodiments by viewing the specification and its drawings. In the specification, the word "comprising" does not exclude other components or steps, and "a" or "one" does not exclude multiple situations. Certain measures are recorded in different embodiments, but this does not mean that these measures cannot be combined to produce good results.

[0093] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A configurable cascade circuit for efficient in-memory processing, characterized in that: The configurable cascade circuit comprises: The on-chip storage unit includes a first BRAM, a second BRAM and a third BRAM; the first BRAM is used to store externally input cascade instructions, the cascade instructions include cascade process termination operation blocks and / or several in-memory computing unit operation blocks; the second BRAM and the third BRAM are used to alternately store intermediate cache data; the second BRAM is also used to store externally input initial data; The state machine module is used to make the state machine module transfer from the waiting state to the cascade instruction pre-configuration state according to the start signal input from the outside; A cascade instruction distributor unit is used to read the cascade instruction from the first BRAM in the cascade instruction pre-configuration state, and divide the cascade instruction into multiple sub-instructions through a bit field segmentation mechanism, and distribute each sub-instruction to a corresponding configuration register unit through an address mapping mechanism, and generate an instruction transmission completion signal after all sub-instruction distribution is completed, so that the state machine module is transferred from the cascade instruction pre-configuration state to the computing engine working state; The decoder unit is used for reading the cascade process termination operation block or the in-memory computing unit operation block corresponding to the sub-instruction from each configuration register unit in the computing engine working state respectively: if the reading result is the in-memory computing unit operation block, parsing and outputting the control word including the first sequence number of the corresponding in-memory computing unit operation block and the channel configuration parameter; if the reading result is the cascade process termination operation block, outputting the cascade termination signal so as to make the state machine module transfer from the computing engine working state to the work completion state; A configurable cascade computing engine is used to determine the enabled in-memory computing unit according to the first sequence number in the control word in the working state of the computing engine, and configure the enabled in-memory computing unit according to the channel configuration parameter in the control word, and use the configured in-memory computing unit to calculate the data read from the second BRAM or the third BRAM to output intermediate cache data, so as to store the intermediate cache data in the second BRAM or the third BRAM; wherein the data read from the second BRAM for the first time is the initial data, and the data read from the second BRAM or the third BRAM thereafter is the intermediate cache data.

2. The configurable cascade circuit for efficient in-memory processing according to claim 1, characterized in that: The cascade instruction distributor unit includes an instruction address counter and an instruction distributor; wherein, The instruction address counter is used to accumulate and count the second serial numbers of the registers allocated to the sub-instructions, and to generate an instruction transmission completion signal when the allocated second serial numbers reach a preset allocation serial number threshold; The instruction distributor is used to divide the cascade instruction into multiple sub-instructions according to a preset bit threshold through a bit field segmentation mechanism, and distribute each sub-instruction to a register corresponding to the second serial number through an address mapping mechanism.

3. The configurable cascade circuit for efficient in-memory processing according to claim 1, characterized in that: The configuration register unit includes a register counter and a plurality of registers; wherein, Each register is used to store a cascade process termination operation block or an in-memory computing unit operation block in a corresponding sub-instruction; The register counter is used for accumulating and counting the second serial numbers of the registers read.

4. The configurable cascade circuit for efficient in-memory processing according to claim 3, characterized in that: The decoder unit includes an operation block parsing module and a parameter discrimination module; wherein, The operation block parsing module is used to read the cascade process termination operation block or the in-memory computing unit operation block from the register corresponding to the read second sequence number, and if the read result is the in-memory computing unit operation block, separate the in-memory computing unit operation block into an operation code and a parameter field; The parameter identification module is used to determine the operation block type according to the operation code. If the operation block type is an in-memory computing unit operation block, the corresponding parameter field is parsed to separate the input channel parameters and the output channel parameters, and the output includes the first serial number of the corresponding in-memory computing unit operation block and the control word of the input channel parameters and the output channel parameters. If the operation block type is a cascade process termination operation block, a cascade termination signal is generated.

5. The configurable cascade circuit for efficient in-memory processing according to claim 4, characterized in that: The configurable cascade computing engine includes a distributed configuration bus and a plurality of in-memory computing units; wherein, The distributed configuration bus is used to receive the control word output by the decoder unit and broadcast it to each in-memory computing unit; Each in-memory computing unit is used to determine an enabled in-memory computing unit according to the first serial number in the control word, and configure the enabled in-memory computing unit according to the input channel parameter and the output channel parameter, and use the configured in-memory computing unit to calculate the data read from the second BRAM or the third BRAM to output intermediate cache data, so as to store the intermediate cache data in the second BRAM or the third BRAM.

6. The configurable cascade circuit for efficient in-memory processing according to claim 5, characterized in that: Each in-memory computing unit is further used to generate an in-memory computing completion signal after calculating and outputting the intermediate cache data, so as to control the register counter to continue to accumulate and count the second serial number of the register read.

7. The configurable cascade circuit for efficient in-memory processing according to claim 6, characterized in that: The configurable cascade computing engine further includes a first multiplexer, a second multiplexer, and a first two-way selector; wherein, The first multiplexer is used to select and output the intermediate cache data calculated by the currently enabled in-memory computing unit from all the in-memory computing units; The second multiplexer is used to select and output the in-memory calculation completion signal generated by the in-memory calculation unit currently determined to be enabled from all the in-memory calculation units; The first two-way selector is used to select and store the intermediate cache data output by the first multiplexer into the second BRAM or the third BRAM.

8. The configurable cascade circuit for efficient in-memory processing according to claim 7, characterized in that: When the storage selection signal is 1, the intermediate cache data is stored in the second BRAM and the intermediate cache data is read from the third BRAM; when the storage selection signal is 0, the intermediate cache data is read from the second BRAM and the intermediate cache data is stored in the third BRAM; The first two-way selector is further configured to select, according to the storage selection signal, to store the intermediate cache data output by the first multiplexer into the second BRAM or the third BRAM.

9. The configurable cascade circuit for efficient in-memory processing according to claim 8, characterized in that: The on-chip storage unit further includes a second two-way selector and a third two-way selector; wherein, The second two-way selector is used to read and output the intermediate cache data stored last time from the second BRAM or the third BRAM according to the in-memory calculation completion signal and the storage selection signal; The third two-way selector is used to read and output the intermediate cache data stored for the last time from the second BRAM or the third BRAM according to the cascade termination signal and the storage selection signal.

10. The configurable cascade circuit for efficient in-memory processing according to claim 1, characterized in that: The state machine module is also used to make the state machine module transfer from the work completion state to the waiting state according to the waiting signal.

Citation Information

Cited By

  • Performance counting device and chip

    CN121501622A