In-memory computing processor, data processing method, and computer device
By building a storage unit and a computing integrated processing engine in the processor, and setting up multiple integrated processing units in it, the problems of low computing efficiency and energy efficiency in traditional processors are solved, and efficient data parallel computing is achieved.
Patent Information
- Application Number
- CN202510199885.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-02-24
AI Technical Summary
The "storage wall" problem in the traditional von Neumann architecture has led to the need to improve computing efficiency and energy efficiency of processors when processing large models and artificial intelligence tasks.
A memory and computing integrated processor is designed. By building the memory unit and the memory and computing integrated processing engine into the processor, and setting up multiple memory and computing integrated processing units in the memory and computing integrated processing engine, data parallel computing and computing efficiency are improved.
By integrating computing functions with storage functions, reducing data transmission, improving computing efficiency and energy efficiency, and supporting parallel data computing, the performance of the processor is significantly improved.
Smart Images

Figure CN119670830B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of novel computing, and particularly to a processor-in-memory processor, a data processing method, and a computer device. Background Art
[0002] In recent years, with the rise of artificial intelligence and large model applications, the demand for processor computing power has been continuously increasing. Although most processors still rely on the von Neumann architecture to meet the growing demand for computing power, in order to develop acceleration computing chips with higher computing power and energy efficiency, it is necessary to solve the "memory wall" problem in the traditional von Neumann architecture. Such a problem requires a new architecture system that integrates storage and computing to solve the problem of separated storage and computing.
[0003] Computing-in-Memory (CiM) is an emerging computing architecture paradigm that aims to integrate computing functions and storage functions in the same processing unit or module. In traditional computer architectures, data usually needs to be transferred from the memory to the processor for computing, and this data transfer consumes a large amount of energy and time. Computing-in-Memory avoids this data movement by performing computing tasks on a new type of memory, thereby improving computing efficiency and energy efficiency. However, most current in-memory computing architectures process data in a serial pipeline mode, and the computing efficiency needs to be improved. Summary of the Invention
[0004] Based on this, in view of the above technical problems, it is necessary to provide a processor-in-memory processor, a data processing method, and a computer device that can improve computing efficiency.
[0005] In a first aspect, this application provides a processor-in-memory processor. It includes: a control module and a data module, and the control module is connected to the data module; wherein, the data module includes a storage unit and a processor-in-memory processing engine, and the processor-in-memory processing engine is connected to the storage unit;
[0006] The processor-in-memory processing engine includes N processor-in-memory processing units and a first calculator group, and the processor-in-memory processing unit is connected to the first calculator group; each of the processor-in-memory processing units includes M processor-in-memory array units and a second calculator group, and the processor-in-memory array unit is connected to the second calculator group; both N and M are natural numbers greater than 1; each of the processor-in-memory array units includes a processor-in-memory array, an array control unit, and a readout circuit, and the processor-in-memory array includes a plurality of cross-arranged memory devices;
[0007] The control module is used to control the processor-in-memory processing engine to obtain data from the storage unit and instruct the processor-in-memory processing engine to calculate the obtained data.
[0008] In one embodiment, the first calculator group includes a first accumulator group and / or a first quantizer group; wherein, the first accumulator group includes N accumulators, and the accumulators are connected to the memory - in - computing processing unit one by one; the first quantizer group includes N quantizers, and the quantizers are connected to the memory - in - computing processing unit one by one;
[0009] The second calculator group includes a second accumulator group and / or a second quantizer group; wherein, the second accumulator group includes M accumulators, and the accumulators are connected to the memory - in - computing array unit one by one; the second quantizer group includes M quantizers, and the quantizers are connected to the memory - in - computing array unit one by one.
[0010] In one embodiment, the data module further includes: a data pre - loading unit, a channel selection unit, and a data pre - processing unit; wherein,
[0011] The input end of the data pre - loading unit is connected to the storage unit, the output end of the data pre - loading unit is connected to the input end of the channel selection unit, and the data pre - loading unit is used to pre - load data from the storage unit;
[0012] The input end of the channel selection unit is further connected to the storage unit, the output end of the channel selection unit is connected to the input end of the data pre - processing unit, and the channel selection unit is used to select the data to be calculated from the storage unit and the data pre - loading unit;
[0013] The output end of the data pre - processing unit is connected to the memory - in - computing processing engine, and the data pre - processing unit is used to pre - process the data output by the channel selection unit.
[0014] In one embodiment, the data pre - loading unit includes N register groups, and the register groups are connected to the data pre - processing unit; the data pre - processing unit is used to pre - process the data output by the channel selection unit, including:
[0015] The data pre - processing unit pre - processes the data of each register group and transmits the pre - processing result to the corresponding memory - in - computing processing unit; or,
[0016] The data pre - processing unit aggregates the data of multiple register groups and transmits the aggregation result to one of the preset memory - in - computing processing units.
[0017] In one embodiment, the data module further includes: a data post-processing unit, an input end of the data post-processing unit is connected to the in-memory computing processing engine, and an output end of the data post-processing unit is connected to the storage unit; wherein,
[0018] The data post-processing unit is configured to post-process the calculation result of the in-memory computing processing engine and transmit the post-processed result to the storage unit.
[0019] In one embodiment, the data module further includes: a write verification unit, the write verification unit is respectively connected to the in-memory computing processing engine and the storage unit; wherein,
[0020] The write verification unit is configured to configure the write verification programming mode of the in-memory computing processing engine, and the write verification programming mode includes a first mode and a second mode; wherein,
[0021] In the first mode, write verification programming is performed on the in-memory computing device at a single address through an immediate write instruction per unit.
[0022] In the second mode, after fetching values from the storage unit through an address write instruction, write verification is performed on the in-memory computing device at the same row address.
[0023] In one embodiment, the data module further includes: a data write-back unit, the data write-back unit is respectively connected to the in-memory computing processing engine and the storage unit; wherein,
[0024] The data write-back unit is configured to sequentially return the calculation result of the in-memory computing processing engine to the storage unit according to the order of the write-back data.
[0025] In one embodiment, the data write-back unit includes: N first write-back FIFOs, W second write-back FIFOs, a data path switch, and an arbiter; W is a natural number greater than or equal to 1;
[0026] The first write-back FIFO is configured to temporarily store the write-back data and write-back address generated by the in-memory computing processing engine and send a request to the arbiter based on the write-back address;
[0027] The arbiter is configured to feedback an acknowledgment signal to the request with the highest priority according to the priority of the request;
[0028] The data path switch is configured to select the target write-back data that obtains the acknowledgment signal, so that the target write-back data is transmitted to the second write-back FIFO;
[0029] The second write-back FIFO is used to temporarily store the target write-back data and the corresponding write-back address, and when the storage unit is available, write the target write-back data into the storage unit based on the corresponding write-back address.
[0030] In one embodiment, the control module includes: a control unit, a register unit, and an instruction unit. The control unit is respectively connected to the register unit, the instruction unit, and the data module; wherein,
[0031] The register unit is used to configure the operation mode and operation address of the data module, and record the operation state of the data module;
[0032] The control unit is used to generate control signals and control the register unit and the data module based on the control signals;
[0033] The instruction unit is used to define an instruction set, which includes control instructions for the control unit to execute and calculation instructions for the memory-computation integrated processing engine to execute.
[0034] In one embodiment, the control instructions include at least one of the following: load data instruction, aligned load data instruction, store instruction, register configuration instruction, write to the memory-computation device by immediate value instruction, write to the memory-computation device by data instruction, data send instruction, module control instruction;
[0035] The calculation instructions include at least one of the following: immediate write-back vector matrix multiplication instruction with data preloading, non-write-back vector matrix multiplication instruction with data preloading, immediate write-back vector matrix multiplication instruction with data immediate fetching, non-write-back vector multiplication instruction with data immediate fetching, immediate write-back matrix multiplication instruction with data preloading, non-write-back matrix multiplication instruction with data preloading, immediate write-back matrix multiplication instruction after data preprocessing, non-write-back matrix multiplication instruction after data preprocessing.
[0036] In a second aspect, the present application provides a data processing method. Applied to the memory-computation integrated processor described in the first aspect above, the method includes:
[0037] The control module generates control signals and sends the control signals to the data module;
[0038] The data module performs data memory access operations and data calculation operations based on the control signals.
[0039] In one embodiment, the data module performing data memory access operations and data calculation operations based on the control signals includes:
[0040] The lThe feature vectors of the layer nodes are loaded into the data preloading unit and aggregated through the data preprocessing unit to obtain an aggregation result; l is a positive integer;
[0041] The aggregation result is input into the memory-computation integrated processing engine for vector matrix calculation, and the obtained vector matrix calculation result is calculated through a non-linear activation function by the data post-processing unit to obtain the feature vector of the l +1 layer.
[0042] In one embodiment, the feature vectors of the l layer nodes are loaded into the data preloading unit and aggregated through the data preprocessing unit to obtain an aggregation result, including:
[0043] Taking adjacent n cycles as a group, the data of the current batch in the feature vectors of the target node and its neighboring nodes in the l layer are sequentially loaded into the data preloading unit according to the cycles; n is a natural number greater than 1;
[0044] In the n+1th cycle, the data of the next batch in the feature vectors of the target node of the
[0045] layer is loaded into the data preloading unit, and the data loaded into the data preloading unit in the previous batch is aggregated through the data preprocessing unit to obtain an aggregated point embedding vector;
[0046] In a third aspect, the present application provides a computer device. The above-mentioned memory-computation integrated processor described in the first aspect is deployed in the computer device.
[0047] The above-mentioned memory-computation integrated processor, data processing method, and computer device, by integrating the storage unit and the memory-computation integrated processing engine into the processor, and setting multiple memory-computation integrated processing units in the memory-computation integrated processing engine, thus, on the basis of realizing memory-computation integration, using a single memory-computation integrated processing unit as the minimum operation unit to process data, supporting data parallel computing, and improving the computing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 is a schematic structural diagram of a memory-computation integrated processor in an embodiment;
[0049] Figure 2 is Figure 1 a schematic structural diagram of the memory-computation integrated processing engine in
[0050] Figure 3 is Figure 2 a schematic structural diagram of the memory-computation integrated array unit in
[0051] Figure 4 Schematic diagram of the structure of the in-memory computing processing engine in an embodiment;
[0052] Figure 5 Schematic diagram of the structure of the data module in an embodiment;
[0053] Figure 6 Schematic diagram of the structure of the data preloading unit in an embodiment;
[0054] Figure 7 Schematic diagram of the structure of the data module in an embodiment;
[0055] Figure 8 Schematic diagram of the data flow within the data module in an embodiment;
[0056] Figure 9 Schematic diagram of the structure of the data module in an embodiment;
[0057] Figure 10 Schematic diagram of the state machine of the write verification unit in an embodiment;
[0058] Figure 11 Schematic diagram of the structure of the data module in an embodiment;
[0059] Figure 12 Schematic diagram of the structure of the data write-back unit in an embodiment;
[0060] Figure 13 Schematic diagram of the data flow within the data module in an embodiment;
[0061] Figure 14 Schematic diagram of the structure of the control module in an embodiment;
[0062] Figure 15 Schematic diagram of the control instruction set in an embodiment;
[0063] Figure 16 Schematic diagram of the calculation instruction set in an embodiment;
[0064] Figure 17 Schematic diagram of the structure of the in-memory computing processor in an embodiment;
[0065] Figure 18 Schematic diagram of the process of the data processing method in an embodiment;
[0066] Figure 19 Schematic diagram of the process of graph convolution calculation in an embodiment;
[0067] Figure 20 Schematic diagram of the process of graph convolution calculation in an embodiment;
[0068] Figure 21 It is a schematic flow diagram of data aggregation processing in an embodiment;
[0069] Figure 22 It is a schematic diagram of node feature aggregation in an embodiment;
[0070] Figure 23 It is a schematic diagram of parallel operation of a processing unit with integrated memory and computing in an embodiment;
[0071] Figure 24 It is a schematic structural diagram of a computer device in an embodiment. Specific implementation manners
[0072] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0073] Unless otherwise defined, the technical terms or scientific terms involved in the present application shall have the general meanings understood by those with ordinary skills in the technical field to which the present application belongs. In the present application, words such as "a", "one", "a kind of", "the", "these", etc. do not represent a limitation in quantity, and they can be singular or plural. The terms "including", "comprising", "having" and any variations thereof involved in the present application are intended to cover non-exclusive inclusion; for example, a process, method, system, product or device including a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent in these processes, methods, products or devices. The terms "connection", "coupling", "coupled" and the like involved in the present application are not limited to physical or mechanical connections, but may include electrical connections, whether directly or indirectly. The term "a plurality" involved in the present application means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may represent: A exists alone, A and B exist simultaneously, and B exists alone. Usually, the character " / " indicates that the objects associated before and after are in an "or" relationship. The terms "first", "second", "third", etc. involved in the present application are only used to distinguish similar objects and do not represent a specific sorting for the objects.
[0074] In an embodiment, a processing unit with integrated memory and computing is provided. Figure 1 It is a schematic structural diagram of the processing unit with integrated memory and computing in the embodiment of the present application, as Figure 1As shown in the figure, the memory - in - processing unit includes: a control module and a data module, where the control module is connected to the data module; among them, the data module includes a storage unit and a memory - in - processing engine, and the memory - in - processing engine is connected to the storage unit; the control module is used to control the memory - in - processing engine to obtain data from the storage unit and instruct the memory - in - processing engine to calculate the obtained data.
[0075] In this embodiment, the storage unit and the memory - in - processing engine are integrated in the data module. The data module can be regarded as a whole structure, which supports data access, storage, and calculation. Among them, the storage unit is used to store data, and the memory - in - processing engine is used to perform calculation operations. Under the control of the control module, it obtains data from the storage unit and performs calculations, and the calculation results can be returned to the storage unit for storage. The control module is used to control the data module. In addition to controlling the memory - in - processing engine to obtain data from the storage unit and instructing the memory - in - processing engine to calculate the obtained data, it can also realize the switching of different operation modes of the entire memory - in - processing unit and data flow control.
[0076] Figure 2 and Figure 3 respectively show the schematic structural diagrams of the memory - in - processing engine in this embodiment. As Figure 2 shown, the memory - in - processing engine includes N memory - in - processing units and a first calculator group, and the memory - in - processing units are connected to the first calculator group; each memory - in - processing unit includes M memory - in - array units and a second calculator group, and the memory - in - array units are connected to the second calculator group; both N and M are natural numbers greater than 1; as Figure 3 shown, each memory - in - array unit includes a memory - in - array, an array control unit, and a read - out circuit, and the memory - in - array includes a plurality of cross - arranged memory - in - devices.
[0077] In the memory - in - processing engine, data calculations are performed with a single memory - in - processing unit as the minimum operation unit, and the first calculator group accumulates and / or quantifies the multiple calculation results of each memory - in - processing unit.
[0078] In the memory - in - processing unit, each memory - in - array unit performs data calculations, and the second calculator group accumulates and / or quantifies the multiple calculation results of each memory - in - array unit.
[0079] In the in-memory computing array unit, multiple in-memory computing devices are arranged in a cross pattern to form an array. Assume the array size is A×B, where A is the number of rows and B is the number of columns, and each in-memory computing device corresponds to a different value. The array control unit is used to control the input voltage value, and the readout circuit is used to read the output current value. Voltages can be input on different rows, such as inputting voltage 1 on the first row, voltage 2 on the second row, ……, voltage A on the A-th row. Each in-memory computing device will output a current, and according to Kirchhoff's current law, the total current output in each column is the sum of the currents output by each in-memory computing device in that column, such as outputting current 1 in the first column, current 2 in the second column, ……, and current B in the B-th column. Among them, the in-memory computing device can be any memory, such as FLASH (Flash Memory), RRAM (Resistive Random Access Memory), MRAM (Magnetoresistive Random Access Memory), or FeFET (Ferroelectric Field-Effect Transistor).
[0080] In the in-memory computing processor of this embodiment, by integrating the storage unit and the in-memory computing processing engine into the processor, and by setting multiple in-memory computing processing units in the in-memory computing processing engine, thus, on the basis of realizing in-memory computing, using a single in-memory computing processing unit as the minimum computing unit to process data, supporting data parallel computing, and improving the computing efficiency.
[0081] In one embodiment, Figure 4 A structural schematic diagram of an in-memory computing processing engine is provided, as Figure 4 shown. The first calculator group includes a first accumulator group and / or a first quantizer group. Among them, the first accumulator group includes N accumulators, and the accumulators are connected to the in-memory computing processing units one by one, and are used to accumulate the multiple calculation results of the corresponding in-memory computing processing units. The first quantizer group includes N quantizers, and the quantizers are connected to the in-memory computing processing units one by one, and are used to quantize the multiple calculation results of the corresponding in-memory computing processing units. The second calculator group includes a second accumulator group and / or a second quantizer group. Among them, the second accumulator group includes M accumulators, and the accumulators are connected to the in-memory computing array units one by one, and are used to accumulate the multiple calculation results of the corresponding in-memory computing array units. The second quantizer group includes M quantizers, and the quantizers are connected to the in-memory computing array units one by one, and are used to quantize the multiple calculation results of the corresponding in-memory computing array units.
[0082] In some of these embodiments, the first calculator group may further include a data buffer group, which includes N data buffers. The data buffers are connected to the processing unit for in-memory computing and storage one by one, and are used to temporarily store the data after passing through the quantizer, providing data input for subsequent data processing.
[0083] In one embodiment, Figure 5 A schematic structural diagram of a data module is provided, as Figure 5 shown. The data module further includes: a data preloading unit, a channel selection unit, and a data preprocessing unit. Among them, the input end of the data preloading unit is connected to the storage unit, the output end of the data preloading unit is connected to the input end of the channel selection unit, and the data preloading unit is used to preload data from the storage unit. The input end of the channel selection unit is also connected to the storage unit, the output end of the channel selection unit is connected to the input end of the data preprocessing unit, and the channel selection unit can be switched through different instructions to select the data to be calculated from the storage unit and the data preloading unit. The output end of the data preprocessing unit is connected to the processing engine for in-memory computing and storage, and the data preprocessing unit is used to preprocess the data output by the channel selection unit.
[0084] In this embodiment, if the data preloading unit receives an enable signal, the channel selection unit selects to read data from the data preloading unit to the data preprocessing unit. If the data preloading unit does not receive an enable signal, the channel selection unit selects to directly read data from the storage unit to the data preprocessing unit.
[0085] Figure 6 A schematic structural diagram of a data preloading unit is provided, as Figure 6 shown. The data preloading unit includes N register groups, and the register groups are connected to the data preprocessing unit. Among them, the data bit width of a single register group is the same as the data amount that can be stored in a single address in the storage unit, that is, the maximum data amount input by the data preloading unit from the storage unit at one time. After the data preloading unit executes data preloading, based on the register unit in the control module, the data is pre-input to the data preprocessing unit for data preprocessing operations.
[0086] In some embodiments, the data preprocessing unit will maintain the corresponding processing unit for in-memory computing and storage during the preprocessing process. For example, after the data preprocessing unit performs preprocessing operations such as non-linear activation function operations or quantization operations on the data of each register group, the preprocessing results are transmitted to the corresponding processing unit for in-memory computing and storage.
[0087] In some of these embodiments, the data preprocessing unit changes the corresponding in-memory computing processing unit during preprocessing. For example, after the data preprocessing unit aggregates the data of multiple register groups, it transmits the aggregation result to one of the preset in-memory computing processing units. Among them, data aggregation includes summing, averaging, finding the maximum value, or finding the minimum value of the data. In some of these embodiments, the preprocessing unit may not perform preprocessing and directly output the data to the in-memory computing processing unit.
[0088] In one embodiment, Figure 7 A schematic structural diagram of a data module is provided, as Figure 7 shown, the data module further includes: a data post-processing unit, the input end of the data post-processing unit is connected to the in-memory computing processing engine, and the output end of the data post-processing unit is connected to the storage unit; wherein, the data post-processing unit is used to perform post-processing on the calculation result of the in-memory computing processing engine, such as numerical accumulation, quantization, and non-linear activation function operations, and transmit the post-processing result to the storage unit.
[0089] Figure 8 A schematic data flow diagram within a data module is provided, as Figure 8 shown, during the calculation process, the data flow starts from the storage unit and is pre-loaded into the data preloading unit according to different calculation instruction types. If the data preloading unit receives an enable signal, the channel selection unit selects to read data from the data preloading unit to the data preprocessing unit. If the data preloading unit does not receive an enable signal, the channel selection unit selects to directly read data from the storage unit to the data preprocessing unit. The data preprocessing unit preprocesses the data and transmits the preprocessing result to the in-memory computing processing engine for calculation. In the in-memory computing processing engine, each in-memory computing processing unit calculates the data in parallel, accumulates and quantizes the multiple calculation results, and temporarily stores the calculation results in the data buffer. The data buffer then hands over the calculation results to the data post-processing unit for further processing, and the finally calculated data will be transmitted back to the storage unit.
[0090] In one embodiment, Figure 9 A schematic structural diagram of a data module is provided, as Figure 9 shown, the data module further includes: a write verification unit, the write verification unit is respectively connected to the in-memory computing processing engine and the storage unit; wherein, the write verification unit is used to configure the write verification programming mode of the in-memory computing processing engine, and the write verification programming mode includes a first mode and a second mode; wherein, in the first mode, write verification programming is performed on the in-memory device of a single address through an immediate write instruction per unit; in the second mode, after taking values from the storage unit through an address write instruction, write verification is performed on the in-memory devices of the same row address.
[0091] In this embodiment, the write verification unit is used for the programming and post-programming verification operations of the memory and computing devices in the memory and computing array, including two write verification programming modes. The first mode is the manual (MANUAL) mode, and the second mode is the automatic (AUTO) mode. Parameters related to the operation mode, write pulse time, and write verification times can be set through the register unit. Among them, the operation mode can be the data preprocessing mode, the data postprocessing mode, the memory and computing engine operation mode, the accumulation mode, and the quantization mode.
[0092] Figure 10 A state machine schematic diagram of the write verification unit is provided, as Figure 10 shown. The two write verification programming modes can be controlled by the state machine. Both modes include four stages, namely the idle (IDLE), write (WRITE), verify (VERIFY), and done (DONE) states. Among them, in the idle stage, the write verification unit is in a dormant state and does not perform any operations. In the write stage, the write verification unit performs a write operation on the memory and computing device based on the pulse width configuration of the configuration register. By setting the time parameter, if the operation times out, it is regarded as completed and enters the verification stage. In the verification stage, the current stage mainly reads the data of the memory and computing device and compares it with the target data. For example, if the read data matches the target data, the current write operation is completed; otherwise, it returns to the write stage to restart the write. In the done stage, it is used for the state recovery after the write verification operation ends, outputting an end signal, and preparing for the next write verification.
[0093] Among them, Condition 1 is: (verification success == 1 && verification quantity != Ncell) || (verification success == 0 && number of times exceeding verification == 0) || (verification success == 0 && number of times exceeding verification == 1 && verification quantity != Ncell).
[0094] Condition 2 is: (verification success == 1 || verification success == 0 && number of times exceeding verification == 1) && verification quantity != Ncell.
[0095] Condition 3 is: verification success == 1 || number of times exceeding verification == 1.
[0096] Condition 4 is: verification success == 0 && number of times exceeding == 0.
[0097] It can be switched between the manual and automatic state machines through mode control. The automatic mode adds the automatic jump of the write device address and the loading operation of multiple write data on the basis of the manual mode. In addition to the normal process write operation, the maximum number of attempts can also be set through the register unit. When the number of attempts of the memory and computing device exceeds the maximum value, it will be automatically skipped.
[0098] In one embodiment, Figure 11A schematic structural diagram of a data module is provided, as Figure 11 shown. The data module further includes: a data write-back unit, which is respectively connected to the in-memory computing processing engine and the storage unit; wherein, the data write-back unit is used to return the calculation results of the in-memory computing processing engine to the storage unit in sequence according to the order of the write-back data.
[0099] In one embodiment, Figure 12 A schematic structural diagram of a data write-back unit is provided, as Figure 12 shown. The data write-back unit includes: N first write-back FIFOs, W second write-back FIFOs, a data path switch, and an arbiter; W is a natural number greater than or equal to 1; the first write-back FIFO is used to temporarily store the write-back data and write-back address generated by the in-memory computing processing engine, and send a request to the arbiter based on the write-back address; the arbiter is used to feedback an acknowledgment signal to the highest-priority request according to the priority of the request; the data path switch is used to select the target write-back data that obtains the acknowledgment signal, so that the target write-back data is transmitted to the second write-back FIFO; the second write-back FIFO is used to temporarily store the target write-back data and the corresponding write-back address, and when the storage unit is available, write the target write-back data into the storage unit based on the corresponding write-back address.
[0100] In this embodiment, the first write-back FIFO is used as the write-back FIFO of the input channel, the second write-back FIFO is used as the write-back FIFO of the output channel, N is the number of in-memory computing processing units, and W is consistent with the number of write-back channels supported by the storage unit, and the minimum value requirement of W is 1.
[0101] Figure 13 A schematic data flow diagram within the data module is provided, as Figure 13 shown. During the calculation process, the data flow starts from the storage unit and is pre-loaded into the data preloading unit according to different calculation instruction types. If the data preloading unit receives an enable signal, the channel selection unit selects to read data from the data preloading unit to the data preprocessing unit. If the data preloading unit does not receive an enable signal, the channel selection unit selects to directly read data from the storage unit to the data preprocessing unit. The data preprocessing unit preprocesses the data, and transmits the preprocessing result to the in-memory computing processing engine for calculation, and temporarily stores the calculation result in the data buffer. If it is an immediate write-back instruction type, the data will be immediately output to the data write-back unit after entering the data buffer. Otherwise, the data will be continuously stored in the data buffer and wait for the STORE instruction. The starting position and the ending position of the flowing calculation data are both in the storage unit, and the network weight-related data is stored in the in-memory computing array unit.
[0102] In one embodiment, Figure 14 A schematic structural diagram of a control module is provided, asFigure 14 As shown in the figure, the control module includes: a control unit, a register unit, and an instruction unit. The control unit is respectively connected to the register unit, the instruction unit, and the data module. Among them, the register unit is used to configure the operation mode and operation address of the data module, and record the operation state of the data module. The control unit is used to generate control signals and control the register unit and the data module based on the control signals. The instruction unit is used to define an instruction set, which includes control instructions for the control unit to execute and calculation instructions for the memory-computation integrated processing engine to execute.
[0103] In this embodiment, the control unit is respectively connected to each component in the memory-computation integrated processor and is used to control operations related to instruction and data scheduling and calculation. The control unit may specifically include an instruction control unit, a data control unit, and a data processing control unit. In addition to controlling each component in the memory-computation integrated processor, the control unit is also responsible for controlling the data communication protocol between the memory-computation integrated processor and the external system bus, adapting each communication protocol for the communication between external data and the system bus through an interface module, and controlling the data flow. The register unit may specifically include a status controller and a configuration register, which are used for parameter configuration of the mode and operation address and recording of the operation state, and can be configured through instructions or bus data access. The instruction unit is used to customize various operations related to the reduced instruction set, including a program counter, an instruction memory, instruction decoding, and distribution.
[0104] The control unit and the register unit jointly implement the switching of different operation modes of the memory-computation processor and the control of the data flow. Among them, the interrupt enable, operation state, node enable, clock enable, data base address, data source address, data target address, PC jump base address, data preprocessing mode, data postprocessing mode, memory-computation integrated processing engine operation mode, memory-computation integrated calculation address, accumulation mode, and quantization mode can be pre-configured in the status register and configuration register in the register unit. The control unit can generate corresponding control signals based on the configuration in the register unit.
[0105] The control module of this embodiment has a built-in customized reduced instruction set and configuration register, enabling the memory-computation integrated processor to have reconfigurability and programmability. By providing different combinations of calculation instructions, multiple memory-computation integrated processing units can enter pipelined parallel computing, improving the throughput of the entire memory-computation integrated processor.
[0106] Figure 15 This is a schematic diagram of the control instruction set in this embodiment. Figure 16 This is a schematic diagram of the calculation instruction set in this embodiment. As Figure 15 and Figure 16As shown, this embodiment adopts a 32-bit custom instruction set, including control instructions related to data scheduling and calculation instructions related to in-memory computing. Among them, the opcode is 4 bits, used to represent the operation instruction; buffer index, used to index the specified buffer; byte mask, used to select the value fetched from the storage address and load it into the buffer; start address, used to specify the starting address loaded into the data buffer; storage address, used for direct data access to the storage unit; storage address offset, used to add with the base address in the configuration register to calculate the address for data access; type parameter is 4 bits, used for adjustment of operations for different instructions; unit index, used to specify the in-memory computing unit corresponding to the instruction; in-memory array index, used to specify the in-memory array corresponding to the instruction; in-memory device address, used for indexing the in-memory device; in-memory device write address, used to specify the target starting address of the instruction for writing data to the in-memory device.
[0107] Table 1 shows the 32-bit custom instruction set of this embodiment.
[0108] Table 1
[0109]
[0110] Among them, the control instructions can be used for data access, weight reading and writing, and data stream processing, including data loading (LOAD), data loading by byte (LOADB), data storage (STORE), register configuration (CONFIG), writing to the in-memory device by unit immediate number (WRITEB), writing to the in-memory device by address (WRITED), reading data to the outside (READ), and control-related instructions (CTRL), such as accumulator clearing, cache shifting, PC jumping, and inserting no-op instructions. The following introduces each control instruction.
[0111] LOAD: Load the data at the storage unit address into a preload buffer (i.e., data preloading unit) with a specified index.
[0112] LOADB: Load the data at the storage unit address into a preload buffer with a specified index.
[0113] STORE: Write back the calculation result stored in the data buffer to the direct storage address. The type 4'b0000 means writing back the buffer data, and the type 4'b0001 means writing back the accumulated data.
[0114] CONFIG: Set data to the configuration register through the direct register address. The operation type [2], 1 means the first half of the 64-bit register, and 0 means the second half; the operation type [1:0] respectively represents the byte selection signal, 11 is half word, and 01 and 10 respectively represent the bytes of the first half and the second half.
[0115] WRITEB: Write to a single specific location of the memory - in - computing device. Operation type 0000 is writing 0, equivalent to reset; operation type 0001 is writing 1, equivalent to set.
[0116] WRITED: Write data on a specific half - row of the RAM from the calculated data storage address (storage address offset + address register).
[0117] SEND: Read data from the storage or memory - in - computing array through a direct address and then send it to the bus.
[0118] CTRL: All operations are defined by different operation types. 0000 is to clear the accumulator, 0001 is to shift the data pre - load register; 0010 is for PC jump; 0011 is for instruction skip and insert bubble.
[0119] Among them, the computing instructions can be used for the computing of the memory - in - computing processing engine, including the immediate write - back vector and matrix multiplication operations (vMUL, mvMUL) for data pre - loading, the immediate write - back vector and matrix multiplication operations (vMULI, mvMULI) for immediate data fetching, the non - write - back vector and matrix multiplication operations (vMULA, mvMULA) for immediate data fetching, and the non - write - back vector and matrix multiplication operations (mvaMUL, mvaMULA) after data pre - processing. The following introduces each computing instruction.
[0120] vMUL: Send the data pre - loaded into a single pre - load register group (the register group in the data pre - load unit) to the specified memory - in - computing processing unit for computing multiply - accumulate, and then directly store the result into the storage unit (destination data storage address = storage address offset + address register).
[0121] vMULA: Send the data pre - loaded in a single pre - load register group to the specified memory - in - computing processing unit for computing multiply - accumulate, and multiple rounds of accumulation can be performed without direct write - back.
[0122] vMULI: Send the data stored in the storage unit (source data storage address = storage address offset + address register 0) to the specified memory - in - computing processing unit for computing multiply - accumulate, and then directly store the result into the storage unit (destination data storage address = storage address offset + address register).
[0123] vMULAI: Send the data stored in the storage unit (source data storage address = storage address offset + address register 0) to the specified memory - in - computing processing unit for computing multiply - accumulate, and multiple rounds of accumulation can be performed without direct write - back.
[0124] mvMUL: Simultaneously send the data stored in multiple data buffers to the processing-in-memory unit enabled by a mask for computing multiply-accumulate, and then directly store the result in the storage unit (destination data storage address = storage address offset + address register);
[0125] mvULA: Simultaneously send the data stored in multiple data buffers to the processing-in-memory unit enabled by a mask for computing multiply-accumulate, which can perform multiple rounds of accumulation without direct write-back;
[0126] mvaMUL: Send the preprocessed data to the specified processing-in-memory unit for computing multiply-accumulate, and then directly store the result in the storage unit (destination data storage address = storage address offset + address register);
[0127] mvaMULA: Send the preprocessed data to the specified processing-in-memory unit for computing multiply-accumulate, which can perform multiple rounds of accumulation without direct write-back.
[0128] In one embodiment, Figure 17 a structural schematic diagram of a processing-in-memory processor is provided, as Figure 17 shown. The processing-in-memory processor includes: a control module and a data module; the control module includes a control unit, a register unit, and an instruction unit, and the control unit is respectively connected to the register unit, the instruction unit, and the data module; the data module includes a storage unit, a processing-in-memory processing engine, a data preloading unit, a channel selection unit, a data preprocessing unit, a data postprocessing unit, a write verification unit, and a data write-back unit. Among them, the storage unit, the data preloading unit, the channel selection unit, the data preprocessing unit, the processing-in-memory processing engine, and the data postprocessing unit are connected in sequence. The write verification unit is respectively connected to the processing-in-memory processing engine and the storage unit, and the data write-back unit is respectively connected to the data postprocessing unit and the storage unit.
[0129] In this embodiment, after the instruction unit fetches and decodes the instruction, the data module receives the control signal sent by the control unit, performs data memory access and data processing operations, and sequentially passes through the pipeline to each component. Among them, the storage unit is used for storing initial data and calculation result data, and can transfer the data to the processing-in-memory processing engine via the data preloading unit or send it to the system bus. For the internal data flow calculation of the data module, reference can be made to the above embodiment, and the description of this embodiment will not be repeated.
[0130] The in-memory computing processor provided in this embodiment is suitable for data processing, including neural network acceleration and matrix calculation acceleration. It adopts a custom reduced instruction set (refer to Table 1 above) and a pipeline structure (refer to the connection relationship of each device inside the data module in the above embodiment), and has programmability, high versatility, and flexibility. It can realize the combination of different computing instructions, support the pipelined parallel computing of different in-memory computing processing units, and improve the throughput of the in-memory computing processor. The in-memory computing processor has reconfigurability and programmability. With the combination of programming instructions and configuration registers, it can be flexibly used in different application scenarios and be optimized in different application scenarios of in-memory computing technologies.
[0131] In one embodiment, a data processing method is provided. Figure 18 As a schematic flowchart of this data processing method, taking the application of this method to the in-memory computing processor in any of the above embodiments as an example, it includes the following steps:
[0132] Step S101, the control module generates a control signal and sends the control signal to the data module;
[0133] Step S102, the data module performs data memory access operations and data calculation operations based on the control signal.
[0134] In this embodiment, after the control signal is sent to the data module, the data module performs instruction fetching and decoding on the control signal, executes data memory access and data processing operations, and sequentially transfers the data to each component through the pipeline. Among them, the data module can specifically include a storage unit and an in-memory computing processing engine. The storage unit is used for storing initial data and calculation result data, and the in-memory computing processing engine is used for performing calculations and returning the calculation results to the storage unit. For the internal data flow calculation of the data module, reference can be made to the above embodiment, and no further description will be given in this embodiment. Among them,
[0135] In one embodiment, taking the application of the in-memory computing processor to graph convolution calculation as an example. During the graph convolution calculation process, it is necessary to find neighboring nodes based on the target node for node embedding feature extraction, including point embedding aggregation, activation function, and vector matrix multiplication. Assuming that for the target node v, there are neighboring nodes in the set u, and the current l+1 layer node feature aggregation formula is:
[0136] ;
[0137] where l+1 the embedding feature of the target node v in the l layer needs to be calculated based on the point embedding features of the previous Figure 19 layer. N(v) represents the number of neighboring nodes of the target node v participating in feature extraction, and L is the total number of calculation layers. This is the schematic flowchart of the graph convolution calculation in this embodiment, asFigure 19 As shown, the process includes the following steps:
[0138] Step S201: Load the feature vectors of the nodes in the l-th layer into the data preloading unit, and perform aggregation processing through the data preprocessing unit to obtain an aggregation result; l is a positive integer.
[0139] Figure 20 This is a schematic diagram of the graph convolution calculation process in this embodiment. As Figure 20 shown, the target node v has 2 neighboring nodes , respectively, in l layer has feature vectors . l The feature vectors in the layer will be preloaded into the data preloading unit, and then data aggregation will be performed through the data preprocessing unit.
[0140] Step S202: Input the aggregation result into the in-memory computing processing engine for vector matrix calculation, and perform a non-linear activation function calculation on the obtained vector matrix calculation result through the data post-processing unit to obtain the feature vectors of the (l + 1)-th layer.
[0141] The output result of the data preprocessing unit will be input into the in-memory computing processing engine for vector matrix calculation, and then the calculation result will complete the non-linear activation function calculation through the data post-processing unit to obtain l+1 the feature vectors of the layer
[0142] Figure 21 This is a schematic diagram of the data aggregation processing flow in this implementation. As Figure 21 shown, load the feature vectors of the nodes in the l layer into the data preloading unit, and perform aggregation processing through the data preprocessing unit to obtain an aggregation result, including the following steps:
[0143] Step S301: Take adjacent n cycles as a group, and load the data of the current batch in the feature vectors of the target node and its neighboring nodes in the l layer into the data preloading unit in sequence; n is a natural number greater than 1.
[0144] Figure 22 This is a schematic diagram of node feature aggregation in this embodiment. As Figure 22 shown, the cycle interval n depends on the number of data to be aggregated for each batch of preprocessing. Since each batch needs to aggregate the feature vectors of 1 target node and 2 neighboring nodes, n = 3. At cycle 0, the data from the i-th to the (i + 7)-th in The i-th to (i + 7)-th data in [[]] is pre-loaded into the pre-loading unit; at cycle 2, The i-th to (i + 7)-th data in [[]] is pre-loaded into the pre-loading unit. Wherein, the variable i represents the position of the first data of the current batch in the dot embedding vector, 0 ≤ i ≤ I, and I is the length of the dot embedding vector.
[0145] Step S302, in the (n + 1)-th cycle, load the data of the next batch in the feature vector of the target node of the l-th layer into the data pre-loading unit, and aggregate the data loaded into the data pre-loading unit in the previous batch through the data pre-processing unit to obtain the aggregated dot embedding vector.
[0146] At cycle 3, The (i + 8)-th to (i + 15)-th data in [[]] is pre-loaded into the pre-loading unit. Meanwhile, The i-th to (i + 7)-th data in [[]] completes aggregation.
[0147] Step S303, input the aggregated dot embedding vector into the in-memory computing processing engine for vector matrix calculation.
[0148] Output the aggregated dot embedding vector The (i + 8)-th to (i + 15)-th data in [[]] goes to the in-memory computing processing engine for calculation; continue to pre-load data based on each cycle later, and output the aggregated dot embedding vector every n cycle intervals The data in [[]] goes to the in-memory computing processing engine for calculation. With such settings, data parallel computing is realized, and the computing efficiency is improved.
[0149] Figure 23 This is a schematic diagram of the parallel operation of the in-memory computing processing unit in this embodiment. As Figure 23 shown, when multiple in-memory computing processing units perform parallel processing, where Ndata represents the number of data to be aggregated during pre-processing, and Npe represents the number of in-memory computing processing units used. In the case where the computing operation of the in-memory computing processing unit takes 8 cycles:
[0150] When Ndata = 8 and Npe = 1, in the case of pipeline processing, each task on average takes 8 cycles to complete processing;
[0151] When Ndata = 6 and Npe = 2, in the case of pipeline processing, each task on average takes 6 cycles to complete processing;
[0152] When Ndata = 3 and Npe = 3, in the case of pipeline processing, each task on average takes 3 cycles to complete processing.
[0153] This embodiment can adjust the parallelism of the in-memory computing processing unit to achieve a better throughput rate in different application scenarios by matching different programming schemes.
[0154] In one embodiment, a computer device is provided. Figure 24 It is a schematic structural diagram of the computer device, as Figure 24 shown. The in-memory computing processor of any of the above embodiments is deployed in the computer device. The computer device may further include a memory and a communication interface. The in-memory computing processor is connected to the memory and the communication interface through a system bus. Among them, the in-memory computing processor is used to provide computing and control capabilities. The memory of the computer device may be a non-volatile storage medium and an internal memory. Among them, the non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database is used to store data. The communication interface of the computer device is used to communicate with an external terminal through a network connection.
[0155] It should be noted that specific examples of the in-memory computing processor in this embodiment may refer to the examples described in the above embodiments and optional implementation manners, and will not be repeated in this embodiment.
[0156] Those skilled in the art can understand that Figure 24 the structure shown in is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0157] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties.
[0158] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0159] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0160] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A storage and computing integrated processor, characterized in that: include: A control module and a data module, wherein the control module is connected to the data module; wherein the data module includes a storage unit and a storage-computing integrated processing engine, and the storage-computing integrated processing engine is connected to the storage unit; The storage-computation integrated processing engine includes N storage-computation integrated processing units and a first calculator group, and the storage-computation integrated processing unit is connected to the first calculator group; each of the storage-computation integrated processing units includes M storage-computation integrated array units and a second calculator group, and the storage-computation integrated array unit is connected to the second calculator group; N and M are both natural numbers greater than 1; each of the storage-computation integrated array units includes a storage-computation integrated array, an array control unit and a readout circuit, and the storage-computation integrated array includes a plurality of storage-computation devices arranged crosswise; The control module is used to control the storage-computation integrated processing engine to obtain data from the storage unit, and instruct the storage-computation integrated processing engine to calculate the obtained data; The data module further includes: a data write-back unit, which is connected to the storage-computation integrated processing engine and the storage unit respectively; wherein the data write-back unit is used to return the calculation results of the storage-computation integrated processing engine to the storage unit in sequence according to the order of writing back data; The data write-back unit includes: N first write-back FIFOs, W second write-back FIFOs, a data path switch and an arbiter; W is a natural number greater than or equal to 1; The first write-back FIFO is used to temporarily store the write-back data and write-back address generated by the storage-computation integrated processing engine, and send a request to the arbitrator based on the write-back address; The arbiter is used to feed back a confirmation signal to the highest priority request according to the priority of the request; The data path switch is used to select the target write-back data that obtains the confirmation signal, so that the target write-back data is transmitted to the second write-back FIFO; The second write-back FIFO is used to temporarily store the target write-back data and the corresponding write-back address, and when the storage unit is available, write the target write-back data into the storage unit based on the corresponding write-back address.
2. The integrated storage and computing processor according to claim 1, characterized in that: The first calculator group includes a first accumulator group and / or a first quantizer group; wherein the first accumulator group includes N accumulators, and the accumulators are connected to the storage-computation integrated processing unit one by one; the first quantizer group includes N quantizers, and the quantizers are connected to the storage-computation integrated processing unit one by one; The second calculator group includes a second accumulator group and / or a second quantizer group; wherein, the second accumulator group includes M accumulators, and the accumulators are connected one-to-one with the storage and computing array units; the second quantizer group includes M quantizers, and the quantizers are connected one-to-one with the storage and computing array units.
3. The integrated storage and computing processor according to claim 1, characterized in that: The data module also includes: a data preloading unit, a channel selection unit and a data preprocessing unit; wherein, The input end of the data preloading unit is connected to the storage unit, the output end of the data preloading unit is connected to the input end of the channel selection unit, and the data preloading unit is used to preload data from the storage unit; The input end of the channel selection unit is also connected to the storage unit, the output end of the channel selection unit is connected to the input end of the data preprocessing unit, and the channel selection unit is used to select the data to be calculated from the storage unit and the data preloading unit; The output end of the data preprocessing unit is connected to the storage-computation integrated processing engine, and the data preprocessing unit is used to preprocess the data output by the channel selection unit.
4. The integrated storage and computing processor according to claim 3, characterized in that: The data preloading unit includes N register groups, and the register groups are connected to the data preprocessing unit; the data preprocessing unit is used to preprocess the data output by the channel selection unit, including: The data preprocessing unit preprocesses the data of each register group and transmits the preprocessing result to the corresponding storage-computation integrated processing unit; or The data preprocessing unit aggregates the data of the plurality of register groups and transmits the aggregation result to one of the preset storage-computation integrated processing units.
5. The integrated storage and computing processor according to claim 1, characterized in that: The data module further includes: a data post-processing unit, the input end of the data post-processing unit is connected to the storage and computing integrated processing engine, and the output end of the data post-processing unit is connected to the storage unit; wherein, The data post-processing unit is used to post-process the calculation results of the storage and computing integrated processing engine and transmit the post-processing results to the storage unit.
6. The integrated storage and computing processor according to claim 1, characterized in that: The data module further includes: a write verification unit, which is connected to the storage-computation integrated processing engine and the storage unit respectively; wherein, The write verification unit is used to configure the write verification programming mode of the storage-computation integrated processing engine, and the write verification programming mode includes a first mode and a second mode; wherein, In the first mode, write verification programming is performed on the storage computing device at a single address by writing an immediate unit value instruction; In the second mode, write verification is performed on the storage and computing devices at the same row address after taking values from the storage cells according to address write instructions.
7. The integrated storage and computing processor according to claim 1, characterized in that: The control module includes: a control unit, a register unit and an instruction unit, and the control unit is connected to the register unit, the instruction unit and the data module respectively; wherein, The register unit is used to configure the operation mode and operation address of the data module, and record the operation status of the data module; The control unit is used to generate a control signal and control the register unit and the data module based on the control signal; The instruction unit is used to define an instruction set, which includes control instructions for execution by the control unit and calculation instructions for execution by the storage-computation integrated processing engine.
8. The integrated storage and computing processor according to claim 7, characterized in that: The control instruction includes at least one of the following: a load data instruction, an align load data instruction, a storage instruction, a register configuration instruction, a write storage computing device instruction according to an immediate number, a write storage computing device instruction according to data, a data sending instruction, and a module control instruction; The calculation instructions include at least one of the following: vector-matrix multiplication instructions that are written back immediately with data preloaded, vector-matrix multiplication instructions that are not written back with data preloaded, vector-matrix multiplication instructions that are written back immediately with data immediately fetched, vector-matrix multiplication instructions that are not written back with data immediately fetched, matrix multiplication instructions that are written back immediately with data preloaded, matrix multiplication instructions that are not written back with data preloaded, matrix multiplication instructions that are written back immediately after data preprocessing, and matrix multiplication instructions that are not written back after data preprocessing.
9. A data processing method, characterized in that: Applied to the storage-computing integrated processor according to any one of claims 1 to 8, the method comprising: The control module generates a control signal and sends the control signal to the data module; The data module performs a data access operation and a data calculation operation based on the control signal.
10. The data processing method according to claim 9, characterized in that: The data module performs a data access operation and a data calculation operation based on the control signal, including: The first l The feature vectors of the layer nodes are loaded into the data preloading unit and aggregated by the data preprocessing unit to obtain the aggregated result; l is a positive integer; The aggregation result is input into the storage and computing integrated processing engine for vector matrix calculation, and the obtained vector matrix calculation result is calculated by the data post-processing unit for nonlinear activation function calculation to obtain the first l The feature vector of the +1 layer.
11. The data processing method according to claim 10, characterized in that: The first l The feature vectors of the layer nodes are loaded into the data preloading unit and aggregated by the data preprocessing unit to obtain the aggregation results, including: Take n adjacent periods as a group and l The current batch of data in the feature vectors of the target node and its neighboring nodes in the layer are sequentially loaded into the data preloading unit according to the cycle; n is a natural number greater than 1; In the n+1th cycle, the l The next batch of data in the feature vector of the target node of the layer is loaded into the data preloading unit, and the data preprocessing unit aggregates the previous batch of data loaded into the data preloading unit to obtain an aggregated point embedding vector; The aggregated point embedding vector is input into the storage-computation integrated processing engine to perform the vector matrix calculation.
12. A computer device, characterized in that: The computer device is deployed with a storage-computing integrated processor as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Storage and calculation chip and storage and calculation method supporting convolution operation and sine and cosine function operation
CN117217272A
Data processing device based on storage and calculation integrated unit
CN117519802A