Data processing method, device and equipment based on multifunctional multiplexing pooling layer circuit

By designing a data processing method based on a multifunctional multiplexed pooling layer circuit, and using accumulator and register configurations to achieve efficient pooling operations, the existing pooling layer circuits have been solved, and the effects of high throughput and multifunctional configurations have been achieved.

CN119990208APending Publication Date: 2025-05-13TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411954850.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing pooled layer circuit has low flexibility, low throughput and large area, making it difficult to meet the needs of efficient data processing in areas such as autonomous driving and security monitoring.

Method used

A data processing method based on a multifunctional multiplexed pooling layer circuit is designed, and the interfaces of the transmitting module, collecting module and pooling circuit are connected. The accumulative comparator is used to realize average pooling, maximum pooling and upsampling pooling, and the dual-thread multi-pooling function is realized through register configuration.

Benefits of technology

It improves the flexibility and functional density of the pooled layer circuit, enhances throughput, reduces circuit area, and meets the needs of distributed computing architecture for high throughput and multifunctional configurations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990208A_ABST
    Figure CN119990208A_ABST
Patent Text Reader

Abstract

The invention relates to a data processing method, device and equipment based on a multifunctional multiplexing pooling layer circuit, the multifunctional multiplexing pooling layer circuit comprises a data sending module, a data receiving module and a pooling circuit, the pooling circuit is connected with the data sending module and the data receiving module through interfaces, the pooling circuit comprises an accumulation comparator, and the accumulation comparator is connected with the data receiving module. The method comprises the steps that if data pooling requirements exist, data to be pooled are sent to a pooling circuit through a data sending module, a target pooling thread is determined from the pooling circuit based on task processing requirements, and the data pooling requirements comprise the data to be pooled, the task processing requirements of the data to be pooled and the data sending requirements; and based on the target pooling thread, performing average pooling, maximum pooling and up-sampling pooling on the to-be-pooled data by using an accumulation comparator to obtain a target pooling result, and sending the target pooling result to a data receiving module based on a data sending demand. Therefore, the problems that an existing pooling layer circuit is low in flexibility, low in throughput rate and large in area are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision technology, and in particular to a data processing method, device and equipment based on a multifunctional multiplexed pooling layer circuit. Background Art

[0002] With the rapid development of information technology, society's demand for automated visual processing continues to grow, especially in the fields of autonomous driving, security monitoring, industrial automation, medical diagnosis, etc., and it is expected that computers can accurately identify and understand various target objects in images like humans, so as to achieve more efficient and safer intelligent services. The development of deep learning technology, especially the breakthrough of convolutional neural networks (CNN, Convolutional Neural Networks), has made significant progress in target recognition technology and has become a research hotspot in the field of computer vision, aiming to solve the problems of target detection, classification and positioning in complex scenes.

[0003] In the target recognition network, the implementation circuit of the pooling layer usually uses a specific hardware unit to efficiently perform the pooling operation. A common implementation circuit is a pooling unit based on an application-specific integrated circuit (ASIC). Its working principle is to receive each pixel data into the buffer, compare, accumulate or copy the data according to the needs, and output the result after completing a sliding window of data processing. However, this implementation circuit will become a performance bottleneck in the distributed computing architecture. There are problems such as lack of versatility, excessive storage usage, and slow processing speed. There is a lot of room for optimization and reuse.

[0004] The real-time requirement of target recognition places extremely high demands on chip processing speed and throughput. The current pooling layer circuit design method does not fully utilize pipeline processing and parallel processing, has a low throughput and relatively slow speed, and may interrupt the upstream data flow waiting for pooling processing. In addition, in the current pooling layer implementation circuit, the ASIC used to implement its function is too single. For example, the ASIC circuit that implements maximum value pooling cannot implement average value pooling. Two sets of circuits are required to implement it, which is costly and cannot be flexibly configured. It is necessary to add redundant data handling to implement different pooling functions. At the same time, for common average value pooling and maximum value pooling, the accumulation and comparison process can reuse adders. However, the current circuit implementation method does not include such a design, and the area cost is high. Summary of the invention

[0005] The present application provides a data processing method, device and equipment based on a multifunctional multiplexed pooling layer circuit to solve the problems of low flexibility, low throughput and large area of ​​the current pooling layer circuit.

[0006] The first aspect of the present application provides a data processing method based on a multifunctional multiplexed pooling layer circuit, wherein the multifunctional multiplexed pooling layer circuit includes: a data transmission module, a data receiving module and a pooling circuit, wherein the pooling circuit is respectively connected to the data transmission module and the data receiving module through interfaces, and the pooling circuit includes an accumulation comparator, wherein the method includes the following steps: determining whether there is a pooling data demand, wherein the pooling data demand includes data to be pooled, task processing requirements of the data to be pooled and data sending requirements; if there is a pooling data demand, using the data transmission module to send the data to be pooled to the pooling circuit, and based on the task processing requirements, determining a target pooling thread from the pooling circuit; based on the target pooling thread, using the accumulation comparator to perform average value pooling, maximum value pooling and upsampling pooling on the data to be pooled to obtain a target pooling result, and sending the target pooling result to the data receiving module based on the data sending requirement.

[0007] Optionally, based on the target pooling thread, the accumulator comparator is used to perform average pooling, maximum pooling and upsampling pooling on the data to be pooled to obtain a target pooling result, and sent to one or more designated downstream ends according to the configuration, including: if the target pooling thread is a single-thread, based on the single-thread, the accumulator comparator is used to perform average pooling, maximum pooling or upsampling pooling on the data to be pooled to obtain a single-thread pooling result, and the single-thread pooling result is used as the target pooling result; if the target pooling thread is a multi-thread, based on the multi-thread, the data to be pooled is performed average pooling, maximum pooling and upsampling pooling in parallel to obtain a multi-thread pooling result, and the multi-thread pooling result is used as the target pooling result, and sent to one or more designated downstream ends according to the configuration.

[0008] Optionally, the pooling circuit also includes an SRAM (Static Random Access Memory) cache unit and a configuration register, and the configuration register includes a two-level register. When the pooling circuit is used to receive the pooled data to be downsampled, it includes: judging whether the feature map channel group of the pooled data to be downsampled is greater than or equal to a preset threshold; if the feature map channel group of the pooled data to be downsampled is less than the preset threshold, using the two-level register to temporarily store the column comparison intermediate value and the accumulated intermediate value of the feature map channel; otherwise, after the pooling circuit receives all channel elements of a full convolution kernel, each column of data of each convolution kernel channel is taken in advance in a preset pipeline manner, and the accumulated sum or maximum value of each column of data is stored in the SRAM cache unit, and the accumulated sum or maximum value is compared with the new pooled data to be downsampled transmitted by the sending module through the two-level register delay to obtain the intermediate value, and the intermediate value is written into the storage address of the corresponding column in the SRAM cache unit.

[0009] Optionally, when using the pooling circuit to receive pooled data to be downsampled, it includes: under the downsampling configuration, determining whether there is at least one pixel in the SRAM cache unit; if the at least one pixel exists in the SRAM cache unit, taking it out from the SRAM cache unit according to the preset pipeline method, and comparing / accumulating it with the next pixel value sent by the upstream number sending module, and after the comparison / accumulation of the last pixel value in each column is completed, storing the taken pixel value in the read data register; at this time, the register stores the comparison / accumulation results of each column of all channels of each feature map in turn.

[0010] Optionally, after writing the comparison result to the storage address of the corresponding column in the SRAM cache unit, it includes: under the downsampling pooling configuration, taking out each column of comparison / accumulation data of the corresponding convolution kernel from the SRAM cache unit according to the preset pipeline method, and comparing / accumulating with the next column of data taken out from the SRAM cache unit, and after the comparison / accumulation of the last column of data is completed, storing each column of data taken out into the read data register; when the output valid signal of the read data register is a high level and the sending signal of the sending module is a pull-up signal, it is determined that the downsampling of the pooled data to be downsampled is completed.

[0011] Optionally, when using the pooling circuit to receive pooled data to be upsampled, it includes: under the upsampling configuration, directly storing the pixel points sent from the upstream in the SRAM cache unit, and then taking them out from the SRAM cache unit to the read data register according to a preset pipeline method; when the output valid signal of the read data register is a high level and the sending signal of the sending module is a pull-up signal, it represents that an upsampling output is completed, and when the number of outputs is equal to the pooling kernel size number, it is determined that the upsampling of the pooled data to be upsampled is completed.

[0012] Optionally, based on the target pooling thread, the accumulator comparator is used to perform average pooling and maximum pooling on the data to be pooled to obtain a target pooling result, including: when performing average pooling, the target pooling result is obtained by performing an addition operation; when performing maximum pooling, a subtraction operation is performed to obtain a subtraction result, and based on the subtraction result, two multiplexers are used to output the larger value of the data to be pooled to obtain the target pooling result.

[0013] Optionally, based on the target downsampling pooling thread, the accumulation comparator is used to perform average pooling, maximum pooling and upsampling pooling on the pooling data to be downsampled to obtain the target pooling result, it also includes: if the data type of the pooling data to be downsampled is an INT4 data type, the INT4 data type is converted to an INT9 data type; if the data type of the pooling data to be downsampled is an INT8 type, the INT8 type is converted to an INT18 type; if the data type of the pooling data to be downsampled is other data types, the type is converted to a longer data type in which the mean accumulation will not overflow.

[0014] The second aspect of the present application provides a data processing device based on a multifunctional multiplexed pooling layer circuit, wherein the multifunctional multiplexed pooling layer circuit includes: a data transmission module, a data receiving module and a pooling circuit, wherein the pooling circuit is respectively connected to the data transmission module and the data receiving module through interfaces, and the pooling circuit includes an accumulative comparator, wherein the device includes: a judgment module, used to judge whether there is a pooling data demand, wherein the pooling data demand includes the data to be pooled, the task processing demand of the data to be pooled and the data sending demand; a determination module, used to use the data transmission module to send the data to be pooled to the pooling circuit if the pooling data demand exists, and determine the target pooling thread from the pooling circuit based on the task processing demand; a pooling module, used to perform average value pooling, maximum value pooling and upsampling pooling on the data to be pooled based on the target pooling thread using the accumulative comparator to obtain a target pooling result, and send the target pooling result to the data receiving module based on the data sending demand.

[0015] Optionally, the pooling module is also used to: if the target pooling thread is a single-thread, then based on the single thread, use the accumulator comparator to perform average pooling, maximum pooling or upsampling pooling on the data to be pooled, to obtain a single-thread pooling result, and use the single-thread pooling result as the target pooling result; if the target pooling thread is a multi-thread, then based on the multi-thread, perform average pooling, maximum pooling and upsampling pooling on the data to be downsampled in parallel, to obtain a multi-thread pooling result, and use the multi-thread pooling result as the target pooling result, and send it to one or more designated downstream ends according to the configuration.

[0016] Optionally, the pooling circuit also includes an SRAM cache unit and a configuration register, and the configuration register includes a two-level register. When the pooling circuit receives the pooled data to be downsampled, the pooling module is also used to: determine whether the feature map channel group of the pooled data to be downsampled is greater than or equal to a preset threshold; if the feature map channel group of the pooled data to be downsampled is less than the preset threshold, the two-level register is used to temporarily store the column comparison intermediate value and the cumulative intermediate value of the feature map channel; otherwise, after the pooling circuit receives all channel elements of a full convolution kernel, each column of data of each convolution kernel channel is taken in advance according to a preset pipeline method, and the cumulative sum or maximum value of each column of data is stored in the SRAM cache unit, and the intermediate value is obtained by comparing the cumulative sum or maximum value with the new pooled data to be downsampled transmitted by the sending module through the two-level register delay, and the intermediate value is written into the storage address of the corresponding column in the SRAM cache unit.

[0017] Optionally, when using the pooling circuit to receive pooled data to be downsampled, the pooling module is also used to: under the downsampling configuration, determine whether there is at least one pixel in the SRAM cache unit; if there is at least one pixel in the SRAM cache unit, take it out from the SRAM cache unit according to the preset pipeline method, and compare / accumulate it with the next pixel value sent by the upstream number sending module, and after the comparison / accumulation of the last pixel value in each column is completed, store the taken pixel value in the read data register; at this time, the register stores the comparison / accumulation results of each column of all channels of each feature map in turn.

[0018] Optionally, after writing the comparison result to the storage address of the corresponding column in the SRAM cache unit, the pooling module is also used to: under the downsampling pooling configuration, take out each column of comparison / accumulation data of the corresponding convolution kernel from the SRAM cache unit according to the preset pipeline method, and compare / accumulate it with the next column of data taken out from the SRAM cache unit, and after the comparison / accumulation of the last column of data is completed, store each column of data taken out into the read data register; when the output valid signal of the read data register is a high level and the sending signal of the sending module is a pull-up signal, it is determined that the downsampling of the pooled data to be downsampled is completed.

[0019] Optionally, when the pooling circuit is used to receive pooled data to be upsampled, the pooling module is also used to: under the upsampling configuration, directly store the pixel points sent from the upstream in the SRAM cache unit, and then take them out from the SRAM cache unit to the read data register according to a preset pipeline method; when the output valid signal of the read data register is a high level and the sending signal of the sending module is a pull-up signal, it represents that an upsampling output is completed, and when the number of outputs is equal to the pooling kernel size number, it is determined that the upsampling of the pooled data to be upsampled is completed.

[0020] Optionally, the pooling module is also used to: during average pooling, perform an addition operation through an accumulator comparator to obtain the target pooling result; during maximum pooling, perform a subtraction operation through an accumulator comparator to obtain a subtraction operation result, and based on the subtraction operation result, use two multiplexers to output the larger value of the data to be pooled to obtain the target pooling result.

[0021] Optionally, based on the target downsampling pooling thread, the accumulation comparator is used to perform average pooling or maximum pooling on the pooling data to be downsampled to obtain the target pooling result, the pooling module is also used to: if the data type of the pooling data to be downsampled is an INT4 data type, convert the INT4 data type into an INT9 data type; if the data type of the pooling data to be downsampled is an INT8 type, convert the INT8 type into an INT18 type; if the data type of the pooling data to be downsampled is other data types, convert the type into a longer data type in which the mean accumulation will not overflow.

[0022] The third aspect of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the data processing method based on the multifunctional multiplexing pooling layer circuit as described in the above embodiment.

[0023] The fourth aspect of the present application provides a computer-readable storage medium on which a computer program is stored. The program is executed by a processor to implement the data processing method based on the multifunctional multiplexing pooling layer circuit as described in the above embodiment.

[0024] In the above implementation, if there is a demand for pooled data, the data to be pooled is sent to the pooling circuit using the sending module, and based on the task processing requirements, the target pooling thread is determined from the pooling circuit, wherein the pooling data requirements include the data to be pooled, the task processing requirements of the data to be pooled, and the data sending requirements, based on the target downsampling pooling thread, the accumulator comparator is used to perform average value pooling, maximum value pooling, and upsampling pooling on the data to be pooled, to obtain the target pooling result, and the target pooling result is sent to the receiving module based on the data sending requirements. Thus, the problems of low flexibility, low throughput, and large area of ​​the current pooling layer circuit are solved. The circuit can receive the data to be pooled from the upstream module in each clock cycle, and flexibly control the thread pooling tasks of two completely different tasks through different register configurations, support multiple data types such as Int4, Int8, and FP16, and compress the cache in the pixel receiving cache stage to times the kernel size of the existing solution, and meet the requirements of the distributed computing architecture for high throughput and multi-functional configuration of the pooling layer.

[0025] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0027] Figure 1 A flowchart of a data processing method based on a multifunctional multiplexing pooling layer circuit provided according to an embodiment of the present application;

[0028] Figure 2 is a schematic diagram of the interface between the pooling circuit and the upstream and downstream modules according to one embodiment of the present application;

[0029] Figure 3 is a schematic diagram of the internal structure of a pooling circuit according to an embodiment of the present application;

[0030] Figure 4 is a schematic diagram of a single pooling core processing flow according to an embodiment of the present application;

[0031] Figure 5 A schematic diagram of an up-sampling and down-sampling data flow diagram according to an embodiment of the present application;

[0032] Figure 6A schematic diagram of a column head data receiving pipeline register architecture according to an embodiment of the present application;

[0033] Figure 7 A schematic diagram of a non-column-head data receiving pipeline register architecture according to an embodiment of the present application;

[0034] Figure 8 A schematic diagram of an SRAM address management module structure according to an embodiment of the present application;

[0035] Fig. 9 A schematic diagram of a hardware implementation of a pooling circuit input pipeline according to an embodiment of the present application;

[0036] Fig.10 A timing diagram of a receiving pipeline register according to an embodiment of the present application;

[0037] Fig.11 A schematic diagram of a transmission pipeline register architecture according to an embodiment of the present application;

[0038] Fig.12 A timing diagram of a sending pipeline register according to an embodiment of the present application;

[0039] Fig.13 A schematic diagram of a register configuration architecture according to an embodiment of the present application;

[0040] Fig.14 is a schematic diagram of a buffer in a pooling circuit according to an embodiment of the present application;

[0041] Fig.15 A schematic diagram of a cache Buffer architecture according to an embodiment of the present application;

[0042] Fig.16 is a schematic diagram of a comparison / accumulation unit according to an embodiment of the present application;

[0043] Fig.17 A schematic diagram of an adder multiplexing structure according to an embodiment of the present application;

[0044] Fig.18 A schematic diagram of a data processing device based on a multifunctional multiplexing pooling layer circuit according to an embodiment of the present application;

[0045] Fig.19 Schematic diagram of the structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0046] Embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.

[0047] The following describes the data processing method, device and equipment based on the multifunctional multiplexed pooling layer circuit of the embodiment of the present application with reference to the accompanying drawings. In view of the problems of low flexibility, low throughput and large area of ​​the current pooling layer circuit mentioned in the above background technology, the present application provides a data processing method based on the multifunctional multiplexed pooling layer circuit, in which if there is a pooling data demand, the data to be pooled is sent to the pooling circuit by using the sending module, and based on the task processing demand, the target pooling thread is determined from the pooling circuit, wherein the pooling data demand includes the data to be pooled, the task processing demand of the data to be pooled and the data sending demand, based on the target pooling thread, the accumulator comparator is used to perform average value pooling, maximum value pooling and upsampling pooling on the data to be pooled, and the target pooling result is obtained, and the target pooling result is sent to the receiving module based on the data sending demand. As a result, the current problems of low flexibility, low throughput and large area of ​​the pooling layer circuit are solved, and multiple pooling types, multiple data types and multi-threaded parallel pipeline design computing circuits can be flexibly configured. Through pipeline architecture design and internal storage management, the pooling circuit can continuously receive the pixel data to be processed sent from the upstream without interrupting the upstream data flow of the in-memory computing chip, thereby meeting the high throughput of the distributed computing architecture. The area required for the circuit is reduced by reusing adders and comparators in the average pooling and maximum pooling stages, and the dual-threaded multi-pooling function is realized through register configuration, thereby improving the circuit flexibility and functional density.

[0048] Specifically, Figure 1 A flowchart of a data processing method based on a multifunctional multiplexing pooling layer circuit provided in an embodiment of the present application.

[0049] like Figure 1 As shown, the multifunctional multiplexing pooling layer circuit includes: a data transmission module, a data receiving module and a pooling circuit, the pooling circuit is respectively connected to the data transmission module and the data receiving module through an interface, the pooling circuit includes an accumulating comparator, wherein the data processing method based on the multifunctional multiplexing pooling layer circuit includes the following steps:

[0050] In step S101 , it is determined whether there is a pooling data demand, wherein the pooling data demand includes data to be pooled, task processing demand of the data to be pooled, and data sending demand.

[0051] In step S102, if there is a demand for pooled data, the data to be pooled is sent to the pooling circuit by using the data transmission module, and based on the task processing demand, a target pooling thread is determined from the pooling circuit.

[0052] In step S103, based on the target pooling thread, the cumulative comparator is used to perform average pooling, maximum pooling and upsampling pooling on the pooled data to obtain the target pooling result, and the target pooling result is sent to the data receiving module based on the data sending requirement.

[0053] Optionally, in some embodiments, based on the target pooling thread, an accumulator comparator is used to perform average pooling, maximum pooling and upsampling pooling on the pooled data to obtain a target pooling result, and the result is sent to one or more designated downstream ends according to the configuration, including: if the target pooling thread is a single-threaded thread, based on the single-thread, an accumulator comparator is used to perform average pooling, maximum pooling or upsampling pooling on the pooled data to obtain a single-threaded pooling result, and the single-threaded pooling result is used as the target pooling result; if the target pooling thread is a multi-threaded thread, based on the multi-threaded thread, the average pooling, maximum pooling and upsampling pooling are performed on the pooled data in parallel to obtain a multi-threaded pooling result, and the multi-threaded pooling result is used as the target pooling result, and the result is sent to one or more designated downstream ends according to the configuration.

[0054] Specifically, the architecture of the dual-channel multi-function multiplexing pooling layer circuit (hereinafter referred to as MFOP, Multi-Function OfPooling) and the upstream and downstream modules of the chip is as follows Figure 2 shown.

[0055] The clock signal, reset signal and memory reset signal of MFOP are connected to the corresponding signals of the top layer. The signal that controls the MFOP pooling function is sent by the register configuration module (Rwreg, Read Write reg), which contains info and done signals. The info signal is the specific content of the configuration, and the done signal is the valid flag of the info signal. To avoid configuration conflicts, the flag should be pulled low for one cycle and then pulled high when switching configurations.

[0056] MFOP obtains a data packet containing pixel data and associated information from an upstream transmission module (i.e., the transmission module in the embodiment of the present application). In each cycle, only one of the multiple transmission modules connected to MFOP is valid, that is, only one associated valid signal is pulled high, and the data packet transmitted in the corresponding direction is valid.

[0057] After completing the pooling of one kernel data configured by Rwreg, MFOP will send the corresponding data to the downstream receiving module (i.e., the receiving module in the embodiment of the present application) in the corresponding direction according to the configuration requirements, that is, pull up the valid signal to send, and load the sending data information into the corresponding direction sending data packet. When the receiving module in the corresponding direction completes data reception and notifies MFOP by pulling up the rdy signal in the corresponding direction, the sending of one data is completed, and MFOP continues to send the subsequent data to the receiving module in this form. At the same time, the cache space released by MFOP after the data transmission is completed will notify the scheduler through a series of cache release signals to configure the sending of subsequent data.

[0058] Among them, the internal structure of MFOP is as follows Figure 3 As shown in the figure, the MFOP is composed of a configuration register (configreg), a comparison / accumulation unit (addcmp, add and compare) with the number of threads supported, an SRAM cache unit (Buffer), and logic control. The function of the configuration register (configreg) is to receive the pooling configuration signal from Rwreg and save it locally in the MFOP. When the upstream configuration information changes, the information saved in the configuration register will be updated by refreshing the cfg_done signal. The comparison / accumulation unit (addcmp) with the number of threads supported is the core unit for completing the pooling operation. The comparator and the adder are integrated by multiplexing the adder to complete the maximum value pooling comparison and average value pooling accumulation and averaging process of inputs of different data types. The cache unit is used to save the intermediate data of the pooling, including the pixel value of the upsampling module and the processing intermediate value of the downsampling module. The logic control module in the MFOP configures other small modules according to the configreg information, sends the corresponding data, address and signal to the corresponding position, and then receives the processed data and sends it to the downstream.

[0059] Optionally, in some embodiments, the pooling circuit also includes an SRAM cache unit and a configuration register, the configuration register includes a two-level register, and when the pooling circuit is used to receive the pooled data to be downsampled, it includes: determining whether the feature map channel group of the pooled data to be downsampled is greater than or equal to a preset threshold; if the feature map channel group of the pooled data to be downsampled is less than the preset threshold, using the two-level register to temporarily store the column comparison intermediate value and the accumulated intermediate value of the feature map channel; otherwise, after the pooling circuit receives all channel elements of a full convolution kernel, each column of data of each convolution kernel channel is taken in advance according to a preset pipeline method, and the accumulated sum or maximum value of each column of data is stored in the SRAM cache unit, and the accumulated sum or maximum value is compared with the new pooled data to be downsampled transmitted by the sending module through the two-level register delay to obtain the intermediate value, and the intermediate value is written into the storage address of the corresponding column in the SRAM cache unit.

[0060] Optionally, in some embodiments, when a pooling circuit is used to receive pooled data to be downsampled, it includes: under the downsampling configuration, determining whether there is at least one pixel in the SRAM cache unit; if there is at least one pixel in the SRAM cache unit, taking it out from the SRAM cache unit according to a preset pipeline method, and comparing / accumulating it with the next pixel value sent by the upstream number sending module, and after the comparison / accumulation of the last pixel value in each column is completed, storing the taken pixel value in the read data register; at this time, the register stores the comparison / accumulation results of each column of all channels of each feature map in turn.

[0061] Optionally, in some embodiments, after writing the comparison result to the storage address of the corresponding column in the SRAM cache unit, it includes: under the downsampling pooling configuration, taking out each column of comparison / accumulation data of the corresponding convolution kernel from the SRAM cache unit in a preset pipeline manner, and comparing / accumulating with the next column of data taken out from the SRAM cache unit, and after the comparison / accumulation of the last column of data is completed, storing each column of data taken out into the read data register; when the output valid signal of the read data register is a high level and the sending signal of the sending module is a pull-up signal, it is determined that the downsampling of the pooled data to be downsampled is completed.

[0062] Optionally, in some embodiments, when a pooling circuit is used to receive pooled data to be upsampled, it includes: under the upsampling configuration, directly storing pixel points sent from upstream in an SRAM cache unit, and then taking them out from the SRAM cache unit to a read data register according to a preset pipeline method; when the output valid signal of the read data register is a high level and the sending signal of the sending module is a pull-up signal, it represents that an upsampling output is completed, and when the number of outputs is equal to the pooling kernel size number, it is determined that the upsampling of the pooled data to be upsampled is completed.

[0063] Under the demand of storage and computing chips, the pooling module needs to receive data from the data transmission module in every cycle, which puts great demands on the throughput of the circuit. Therefore, the embodiment of the present application adopts the method of pipeline design to achieve the purpose of improving the module throughput. The process is as follows: Figure 4 As shown in Figure 1, the data transmission in the hardware is as follows Figure 5 shown.

[0064] For the downsampling module, when the pooling circuit receives the data pipeline

[0065] When the processing feature map channel group channel group is greater than or equal to the preset threshold 3, the pooling circuit receives data from the corresponding direction according to the input data valid signal, quantizes the data so that each flit (256 bits) of data is expanded to 576 bits, and stores the data in SRAM.

[0066] After receiving a full kernel channel, the data of the kernel channel column is taken out in advance in a pipeline manner, compared with the new number transmitted by the number sending module, and the result is overwritten to the same address of SRAM. The design here adjusts the receiving data pipeline, and the key is to realize 1) SRAM stores the cumulative sum or maximum value of each column, not a single pixel, 2) realize the corresponding comparison of data by adding register delay, such as when calculating the mean pooling, when (2 rows, 2 columns, 1 channel) data arrives, the data in the SRAM address taken out is the cumulative sum of (0 rows, 2 columns, 1 channel) and (1 row, 2 columns, 1 channel), that is, the sum of (0 rows, 2 columns, 1 channel), (1 row, 2 columns, 1 channel), (3 rows, 2 columns, 1 channel), 3) after completing the accumulation or comparison, write the result to the storage address of the corresponding column in SRAM.

[0067] The received data pipeline stores the data of each channel and each column in the SRAM in the order of (channel0, column 0), (channel1, column 0)...(channel n, column 0), (channel0, column 1).

[0068] In MFOP, the configuration register specifies the size of the SRAM space occupied by each thread by giving a base address base and a top address end. In addition, in the pooled application of the embodiment of the present application, address reading is sequential reading and writing without jumping. Therefore, it can be determined that the initial address of each thread is base, and each time a read / write operation is completed, the read / write address is +1, and when the read / write address reaches end, it jumps back to base to continue.

[0069] The receiving data pipeline architecture is divided into two types according to whether it is the first element of the column, such as Figure 6 Shown and Figure 7 shown.

[0070] After the received data arrives, it will be selected by the input valid signal and delayed by two registers, register 0 and register 1. Figure 8 The data cache address management and input data selection are implemented by hardware. SRAM_ab is the SRAM write address register, and SRAM_aa is the SRAM read address register. The signals SRAM_base_X and SRAM_end_X correspond to the base address signal and the top address signal of the X thread respectively.

[0071] When the processing feature graph channel group is less than the preset threshold value 3, it is a special case. In this case, the upstream sends data too fast, resulting in a delay in the process of storing and retrieving the corresponding data in SRAM, and a read-after-write (WAR) conflict will occur. The data read out is not the accumulated value or maximum value of the corresponding column. The way to solve this problem in the embodiment of the present application is to use 2 (threads) * 2 (channel groups) * 576bit = 2Kbit registers to temporarily store column comparison / accumulated intermediate data in this special case, and then write to SRAM to avoid the delay caused by directly reading and writing SRAM. The rest of the data processing flow is the same. Registers are not used to store intermediate variables unless there are special circumstances, because the corresponding intermediate variable space requirement = thread * channel group number is too large, the register occupies too much area, and this part of the register resources is not easy to reuse, so SRAM is used to store this part of the data while meeting the pipeline length.

[0072] exist Fig. 9 In the figure, two directions of dual-thread input and two directions of output in four directions are taken as examples. The input_sel signal is the valid selection signal in the four directions of input, register 0 (receive_reg0) and register 1 (receive_reg1) are two-stage input data delay registers, the accumulator comparator output (out_addcmp1) signal is the output result of the accumulation / comparison unit, the if_col_1st signal is the column head element signal, and SRAM_db is the SRAM write data register. In addition, whether the signal of the first pixel in the column selects whether the result of addcmp or the result of receive_reg1 is stored in the SRAM. The reason for the delay through the two-stage registers is to allow the data to be accumulated / compared to complete the process of accumulation / comparison-writing to SRAM-reading from SRAM, such as the timing Fig.10 As shown in the figure, data0 needs to be compared with data3. If there is no register delay, data3 will arrive before data0 is taken out of SRAM, causing a conflict and failing to form a pipeline. After adding registers, although the pipeline length becomes longer, MFOP can receive, compare and read different data in each cycle, increasing the system throughput.

[0073] The counter monitors the receive data pipeline. When a channel completes the kernel reception and writes it into SRAM, the transmit data pipeline is started.

[0074] When the data transmission pipeline is started, there is at least one complete kernel of a channel in the SRAM. Each column of data corresponding to the kernel is taken out from the SRAM in a pipeline manner and stored in the temporary register. After the next column of data is taken out, it is compared and the result is overwritten to the temporary register. It should be noted that the average value pooling starts to calculate the average value after the last column comparison.

[0075] like Fig.11 As shown, when the data transmission pipeline is started, the data in the SRAM is taken out in sequence to the output register. Every time the vld-rdy signal handshake is successful, the output counter is increased by 1. When the output counter reaches the upsampling number configured by the register, the SRAM fetch address is changed to take out the next data to the output register and continue to output until the upsampling process of all data in the SRAM is completed.

[0076] like Fig.12 As shown, after the accumulation or comparison of the last column is completed, the 576-bit inter-column comparison result is dequantized into 1flit data, and the output register is assigned, and the output valid signal is valid. The output end and the downstream module shake hands through the vld-rdy signal. When the output valid signal is high and the rdy of the transmission module is invalid, the output register value remains unchanged and the previous pipeline is paused; when the output valid signal is high. And the rdy of the transmission module is pulled high, that is, after the reception is completed, the pipeline continues to output the next data until the downsampling process of all data in the SRAM is completed.

[0077] It should be noted that the processing of the upsampling module is basically the same as the downsampling process. The difference is that each upsampled data needs to occupy an SRAM space, instead of each feature map column data occupying a SRAM space.

[0078] For receive data pipeline: in upsampling mode, each receive data is stored in SRAM; when at least one pixel is stored in SRAM, the transmit data pipeline is started.

[0079] For the transmit data pipeline: The transmit data pipeline architecture is the same as the upsampling process.

[0080] The current target recognition neural network has multiple pooling layers, and the pooling types and pooling parameters of different pooling layers are different. In addition, the computational parallelism requirements of the pooling layer are high, and the circuit needs to be optimized. The embodiment of the present application realizes flexible control of the dual-thread pooling function circuit through different register configurations, supports different pooling layers and different pooling parameters, such as Fig.13 shown.

[0081] This circuit implements the 64-bit wide configuration signal to pre-configure the processing function before the data arrives. For different pooling layers of the network, this configuration can be used to directly convert the pooling layer circuit function.

[0082] The pooling layer circuit supports multiple data transmission and reception modes: 1 to 1, 2 to 1, and 2 to 2. For single-threaded pooling tasks, the circuit supports receiving data from any of the four upstream modules and sending data to any of the four downstream modules. For pooling tasks that require parallel processing, the top-level controller will be mapped to two threads supported by the pooling layer circuit, sending data from two of the four upstream modules, and sending data to any one or two of the four downstream modules as needed after processing.

[0083] When the configuration thread is completed, the top-level controller modifies the 64-bit wide configuration signal to the corresponding pooling task configuration for the next stage.

[0084] Optionally, in some embodiments, based on the target downsampling pooling thread, the downsampling pooling data is subjected to average pooling, maximum pooling and upsampling pooling using an accumulator comparator to obtain a target pooling result, and further includes: if the data type of the downsampling pooling data is an INT4 data type, the INT4 data type is converted to an INT9 data type; if the data type of the downsampling pooling data is an INT8 type, the INT8 type is converted to an INT18 type; if the data type of the downsampling pooling data is other data types, the type is converted to a longer data type in which the mean accumulation will not overflow.

[0085] Specifically, the pooling process requires a large amount of data cache to store the pooling calculation results of the incoming data. SRAM storage is used to achieve high-density storage. The cache used by MFOP is as follows: Fig.14 shown.

[0086] For the INT4 / INT8 multiplexing case, the maximum downsampling kernel is a pooling circuit with a 5x5 task. The 64-bit INT4 arriving in each clock cycle needs to be saved with the INT9 data type to achieve the accumulation of 25 kernel numbers without data overflow. Therefore, the input 256b data is expanded to 576b for storage, that is, each INT4 is converted to INT9, and INT8 is converted to INT18, to ensure that there is no precision loss in the accumulated data before the final division of the average value pooling.

[0087] For example, Fig.15The buffer implemented by the circuit consists of four 64x128b SRAMs, which are stored in the output order. For example, if 64 INT4 type data are input, each INT4 data is expanded into INT9 and stored in one address space of the buffer; if 32 INT8 type data are input, each INT8 data is expanded into INT18 and stored in one address space of the buffer.

[0088] Optionally, in some embodiments, based on the target pooling thread, an accumulator comparator is used to perform average pooling, maximum pooling and upsampling pooling on the data to be pooled to obtain a target pooling result, including: during average pooling, an addition operation is performed by the accumulator comparator to obtain a target pooling result; during maximum pooling, a subtraction operation is performed by the accumulator comparator to obtain a subtraction operation result, and based on the subtraction operation result, two multiplexers are used to output the larger value of the data to be pooled to obtain the target pooling result.

[0089] Specifically, the circuit uses an adder to call a module within DesignWare. The comparator used for maximum pooling can be an adder for average pooling. When performing maximum pooling, the DW module is adjusted to be configured as subtraction, and the comparison result is determined according to the result sign bit.

[0090] like Fig.16 As shown, the multiplexing unit receives two 576-bit data A and B to be compared or accumulated, and sends the processing result to the out signal. The module addcmp needs to configure its processing mode. The relevant signals include: 1) Data type signal datatype. Determines how addcmp should divide A and B; 2) Calculation mode signal op. Determines whether addcmp performs accumulation operation on the data or comparison-output larger value operation; 3) Mean division div_enable and div. Determine whether addcmp performs average division on the accumulated value, and the divisor value.

[0091] In the embodiment of the present application, the DesignWare adder and two multiplexers are used to implement the functions of the adder and comparator modules. The specific circuit diagram is as follows: Fig.17 As shown. The circuit function is controlled by the op signal. When the function is average pooling, op = 0, the addcmp module selects Z = A + B and outputs the sum of the two numbers; when the function is maximum pooling, op = 1, the first multiplexer selects the maximum value between A and B according to the sign bit Z[-1] of the adder output result, and the second multiplexer selects the value of op = 1 and outputs the maximum value between A and B.

[0092] In summary, the advantages of the embodiments of the present application are as follows:

[0093] (1) In the current pooling computing circuit, after receiving a kernel data and completing the pooling processing, the next kernel data is received. The data throughput is not fixed, which will cause the upstream data pipeline to be interrupted during the period between two kernels; in addition, different network pooling functions require different computing units, which takes up a large amount of unnecessary area and increases the overall circuit cost and power consumption. In order to meet the requirements of the pooling layer circuit implemented by the in-memory computing chip, there is an urgent need for a flexible, multi-functional, high-throughput, and high-density computing unit to implement the function. Therefore, the embodiment of the present application proposes a pipeline design computing circuit that uses register configuration and can flexibly configure multiple pooling types, multiple data types, and multi-threaded parallelism.

[0094] (2) Through pipeline architecture design and internal storage management, the pooling circuit can continuously receive the pixel data to be processed sent from the upstream without interrupting the upstream data flow of the in-memory computing chip, thereby meeting the high throughput of the distributed computing architecture; by reusing adders and comparators in the average pooling and maximum pooling stages, the required circuit area is reduced; and by configuring registers to realize dual-thread multi-pooling functions, the circuit flexibility and functional density are improved.

[0095] It should be noted that other alternative solutions can also be used to implement the technical solutions in the embodiments of the present application:

[0096] (1) Using storage similar to SRAM;

[0097] (2) Multifunctional pooling circuits for different data types;

[0098] (3) Different numbers of input and output combinations form multifunctional pooling circuits.

[0099] According to the data processing method based on the multifunctional multiplexed pooling layer circuit proposed in the embodiment of the present application, if there is a demand for pooled data, the data to be pooled is sent to the pooling circuit by using the data sending module, and based on the task processing demand, the target pooling thread is determined from the pooling circuit, wherein the pooling data demand includes the data to be pooled, the task processing demand of the data to be pooled, and the data sending demand, based on the target pooling thread, the accumulator comparator is used to perform average pooling, maximum pooling, and upsampling pooling on the data to be pooled to obtain the target pooling result, and the target pooling result is sent to the data receiving module based on the data sending demand. As a result, the current problems of low flexibility, low throughput and large area of ​​the pooling layer circuit are solved, and multiple pooling types, multiple data types and multi-threaded parallel pipeline design computing circuits can be flexibly configured. Through pipeline architecture design and internal storage management, the pooling circuit can continuously receive the pixel data to be processed sent from the upstream without interrupting the upstream data flow of the in-memory computing chip, thereby meeting the high throughput of the distributed computing architecture. The area required for the circuit is reduced by reusing adders and comparators in the average pooling and maximum pooling stages, and the dual-threaded multi-pooling function is realized through register configuration, thereby improving the circuit flexibility and functional density.

[0100] Next, a data processing device based on a multifunctional multiplexing pooling layer circuit proposed according to an embodiment of the present application is described with reference to the accompanying drawings.

[0101] Fig.18 It is a block diagram of a data processing device based on a multifunctional multiplexing pooling layer circuit according to an embodiment of the present application.

[0102] Among them, the multifunctional multiplexing pooling layer circuit includes: a data transmission module, a data receiving module and a pooling circuit. The pooling circuit is connected to the data transmission module and the data receiving module through interfaces respectively. The pooling circuit includes an accumulator comparator. The data processing device based on the multifunctional multiplexing pooling layer circuit includes: a judgment module 100, a determination module 200 and a pooling module 300.

[0103] Among them, the judgment module 100 is used to determine whether there is a pooling data demand, and the pooling data demand includes the data to be pooled, the task processing demand of the data to be pooled and the data sending demand; the determination module 200 is used to use the sending module to send the data to be pooled to the pooling circuit if there is a pooling data demand, and determine the target pooling thread from the pooling circuit based on the task processing demand; the pooling module 300 is used to use the accumulator comparator to perform average pooling, maximum pooling and upsampling pooling on the data to be pooled based on the target pooling thread, obtain the target pooling result, and send the target pooling result to the receiving module based on the data sending demand.

[0104] Optionally, in some embodiments, the pooling module 300 is also used for: if the target pooling thread is a single-thread, then based on the single thread, using an accumulator comparator to perform average pooling, maximum pooling or upsampling pooling on the pooled data to obtain a single-thread pooling result, and using the single-thread pooling result as the target pooling result; if the target pooling thread is a multi-thread, then based on the multi-thread, performing average pooling, maximum pooling and upsampling pooling on the pooled data in parallel to obtain a multi-thread pooling result, and using the multi-thread pooling result as the target pooling result, and sending it to one or more designated downstream ends according to the configuration.

[0105] Optionally, in some embodiments, the pooling circuit also includes an SRAM cache unit and a configuration register, the configuration register includes a two-level register, and when the pooling circuit receives the pooling data to be downsampled, the pooling module 300 is also used to: determine whether the feature map channel group of the pooling data to be downsampled is greater than or equal to a preset threshold; if the feature map channel group of the pooling data to be downsampled is less than the preset threshold, the two-level register is used to temporarily store the column comparison intermediate value and the cumulative intermediate value of the feature map channel; otherwise, after the pooling circuit receives all channel elements of a full convolution kernel, each column of data of each convolution kernel channel is taken in advance according to a preset pipeline method, and the cumulative sum or maximum value of each column of data is stored in the SRAM cache unit, and the cumulative sum or maximum value is compared with the new pooling data to be downsampled transmitted by the sending module through the two-level register delay to obtain the intermediate value, and the intermediate value is written into the storage address of the corresponding column in the SRAM cache unit.

[0106] Optionally, in some embodiments, when a pooling circuit is used to receive pooled data to be downsampled, the pooling module is also used for 300: in the downsampling configuration, determining whether there is at least one pixel in the SRAM cache unit; if there is at least one pixel in the SRAM cache unit, taking it out from the SRAM cache unit according to a preset pipeline method, and comparing / accumulating it with the next pixel value sent by the upstream number sending module, after the comparison / accumulation of the last pixel value in each column is completed, the taken-out pixel value is stored in the read data register; at this time, the comparison / accumulation results of each column of all channels of each feature map are stored in the register in sequence.

[0107] Optionally, in some embodiments, after writing the comparison result to the storage address of the corresponding column in the SRAM cache unit, the pooling module 300 is also used to: under the downsampling pooling configuration, take out each column of comparison / accumulation data of the corresponding convolution kernel from the SRAM cache unit in a preset pipeline manner, and compare / accumulate it with the next column of data taken out from the SRAM cache unit, and after the comparison / accumulation of the last column of data is completed, store each column of data taken out into the read data register; when the output valid signal of the read data register is high and the sending signal of the sending module is a pull-up signal, it is determined that the downsampling of the pooled data to be downsampled is completed.

[0108] Optionally, in some embodiments, when a pooling circuit is used to receive pooled data to be upsampled, the pooling module 300 is also used to: under the upsampling configuration, directly store the pixel points sent from the upstream in the SRAM cache unit, and then take them out from the SRAM cache unit to the read data register according to a preset pipeline method; when the output valid signal of the read data register is a high level and the sending signal of the sending module is a pull-up signal, it represents that an upsampling output is completed, and when the number of outputs is equal to the pooling kernel size number, it is determined that the upsampling of the pooled data to be upsampled is completed.

[0109] Optionally, in some embodiments, the pooling module 300 is also used to: during average pooling, perform an addition operation through the accumulator comparator to obtain a target pooling result; during maximum pooling, perform a subtraction operation through the accumulator comparator to obtain a subtraction operation result, and based on the subtraction operation result, use two multiplexers to output the larger value in the data to be pooled to obtain the target pooling result.

[0110] Optionally, in some embodiments, when the target pooling result is obtained by using an accumulator comparator to perform average pooling or maximum pooling on the downsampled pooling data based on the target downsampled pooling thread, the pooling module 300 is also used to: if the data type of the downsampled pooling data is an INT4 data type, convert the INT4 data type to an INT9 data type; if the data type of the downsampled pooling data is an INT8 type, convert the INT8 type to an INT18 type; if the data type of the downsampled pooling data is other data types, convert the type to a longer data type in which the mean accumulation will not overflow.

[0111] It should be noted that the aforementioned explanation of the data processing method embodiment based on the multifunctional multiplexing pooling layer circuit is also applicable to the data processing device based on the multifunctional multiplexing pooling layer circuit of this embodiment, and will not be repeated here.

[0112] According to the data processing device based on the multifunctional multiplexed pooling layer circuit proposed in the embodiment of the present application, if there is a demand for pooled data, the data to be pooled is sent to the pooling circuit by using the data sending module, and based on the task processing demand, the target pooling thread is determined from the pooling circuit, wherein the pooling data demand includes the data to be pooled, the task processing demand of the data to be pooled, and the data sending demand, and based on the target pooling thread, the accumulator comparator is used to perform average pooling, maximum pooling, and upsampling pooling on the data to be pooled to obtain the target pooling result, and the target pooling result is sent to the data receiving module based on the data sending demand. As a result, the current problems of low flexibility, low throughput and large area of ​​the pooling layer circuit are solved, and multiple pooling types, multiple data types and multi-threaded parallel pipeline design computing circuits can be flexibly configured. Through pipeline architecture design and internal storage management, the pooling circuit can continuously receive the pixel data to be processed sent from the upstream without interrupting the upstream data flow of the in-memory computing chip, thereby meeting the high throughput of the distributed computing architecture. The area required for the circuit is reduced by reusing adders and comparators in the average pooling and maximum pooling stages, and the dual-threaded multi-pooling function is realized through register configuration, thereby improving the circuit flexibility and functional density.

[0113] Fig.19 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device may include:

[0114] Memory 1901 , processor 1902 , and a computer program stored in the memory 1901 and executable on the processor 1902 .

[0115] When the processor 1902 executes the program, the data processing method based on the multifunctional multiplexing pooling layer circuit provided in the above embodiment is implemented.

[0116] Furthermore, the electronic device further comprises:

[0117] The communication interface 1903 is used for communication between the memory 1901 and the processor 1902 .

[0118] The memory 1901 is used to store computer programs that can be executed on the processor 1902 .

[0119] The memory 1901 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.

[0120] If the memory 1901, the processor 1902 and the communication interface 1903 are implemented independently, the communication interface 1903, the memory 1901 and the processor 1902 can be connected to each other through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig.19 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0121] Optionally, in a specific implementation, if the memory 1901, the processor 1902 and the communication interface 1903 are integrated on a chip, the memory 1901, the processor 1902 and the communication interface 1903 can communicate with each other through an internal interface.

[0122] The processor 1902 may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.

[0123] An embodiment of the present application also provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the data processing method based on the multifunctional multiplexing pooling layer circuit is implemented as above.

[0124] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.

[0125] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0126] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.

[0127] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable storage medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable storage medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples (non-exhaustive list) of computer-readable storage media include the following: an electrical connection with one or N wirings (electronic device), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable storage medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways if necessary, and then stored in a computer memory.

[0128] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiment, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0129] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.

[0130] In addition, each functional unit in each embodiment of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0131] The computer-readable storage medium mentioned above may be a read-only memory, a disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limiting the present application. A person of ordinary skill in the art may change, modify, replace and modify the above embodiments within the scope of the present application.

Claims

1. A data processing method based on a multifunctional multiplexing pooling layer circuit, characterized in that: The multifunctional multiplexing pooling layer circuit includes: a data transmission module, a data receiving module and a pooling circuit, wherein the pooling circuit is respectively connected to the data transmission module and the data receiving module through interfaces, and the pooling circuit includes an accumulating comparator, wherein the method includes: Determine whether there is a pooling data demand, wherein the pooling data demand includes the data to be pooled, the task processing demand of the data to be pooled, and the data sending demand; If there is a demand for the pooled data, the data to be pooled is sent to the pooling circuit by using the data transmission module, and a target pooling thread is determined from the pooling circuit based on the task processing demand; Based on the target pooling thread, the accumulator comparator is used to perform average pooling, maximum pooling and upsampling pooling on the data to be pooled to obtain a target pooling result, and the target pooling result is sent to the data receiving module based on the data sending requirement.

2. The method according to claim 1, characterized in that Based on the target pooling thread, the cumulative comparator is used to perform average pooling, maximum pooling and upsampling pooling on the pooled data to obtain the target pooling result, and sends it to one or more designated downstream ends according to the configuration, including: If the target pooling thread is a single thread, then based on the single thread, the accumulator comparator is used to perform average pooling, maximum pooling, or upsampling pooling on the data to be pooled to obtain a single-thread pooling result, and the single-thread pooling result is used as the target pooling result; If the target pooling thread is multi-threaded, then based on the multi-thread, average pooling, maximum pooling and upsampling pooling are performed on the data to be pooled in parallel to obtain a multi-threaded pooling result, and the multi-threaded pooling result is used as the target pooling result, and is sent to one or more designated downstream ends according to the configuration.

3. The method according to claim 1, characterized in that The pooling circuit further includes an SRAM cache unit and a configuration register, wherein the configuration register includes two levels of registers. When the pooling circuit receives the pooling data to be downsampled, it includes: Determine whether the feature map channel group of the to-be-downsampled pooled data is greater than or equal to a preset threshold; If the feature map channel group of the to-be-downsampled pooled data is smaller than the preset threshold, the two-level register is used to temporarily store the column comparison intermediate value and the accumulated intermediate value of the feature map channel; otherwise, after the pooling circuit receives all channel elements of a full convolution kernel, each column of data of each convolution kernel channel is taken in advance in a preset pipeline manner, and the accumulated sum or the maximum value of each column of data is stored in the SRAM cache unit, and the accumulated sum or the maximum value is compared with the new to-be-downsampled pooled data transmitted by the data transmission module through the two-level register delay to obtain the intermediate value, and the intermediate value is written into the storage address of the corresponding column in the SRAM cache unit.

4. The method according to claim 3, characterized in that: When the pooling circuit is used to receive the pooling data to be downsampled, the method includes: In the downsampling configuration, determining whether there is at least one pixel in the SRAM cache unit; If at least one pixel point exists in the SRAM cache unit, it is taken out from the SRAM cache unit according to the preset pipeline method, and compared / accumulated with the next pixel point value sent by the upstream number sending module. After the comparison / accumulation of the last pixel point value in each column is completed, the taken pixel point value is stored in the read data register; at this time, the comparison / accumulation results of each column of all channels of each feature map are stored in the register in turn.

5. The method according to claim 3, characterized in that: After writing the comparison result into the storage address of the corresponding column in the SRAM cache unit, comprising: Under the downsampling pooling configuration, each column of comparison / accumulation data corresponding to the convolution kernel is taken out from the SRAM cache unit according to the preset pipeline mode, and compared / accumulated with the next column of data taken out from the SRAM cache unit, and after the comparison / accumulation of the last column of data is completed, each column of data taken out is stored in the read data register; When the output valid signal of the read data register is at a high level and the sending signal of the sending module is a pull-up signal, it is determined that the downsampling of the to-be-downsampled pooled data is completed.

6. The method according to claim 5, characterized in that When the pooling circuit is used to receive the pooling data to be upsampled, the method includes: In the upsampling configuration, the pixels sent from the upstream are directly stored in the SRAM cache unit, and then taken out from the SRAM cache unit to the read data register according to a preset pipeline method; When the output valid signal of the read data register is at a high level and the sending signal of the sending module is a pull-up signal, it represents that an upsampling output is completed. When the number of outputs is equal to the pooling kernel size number, it is determined that the upsampling of the pooled data to be upsampled is completed.

7. The method according to claim 1, characterized in that The target pooling thread, using the accumulator comparator to perform average pooling, maximum pooling and upsampling pooling on the data to be pooled, to obtain a target pooling result, includes: During average value pooling, an addition operation is performed by the accumulator comparator to obtain the target pooling result; During maximum pooling, a subtraction operation is performed through an accumulator comparator to obtain a subtraction operation result, and based on the subtraction operation result, two multiplexers are used to output a larger value in the data to be pooled to obtain the target pooling result.

8. The method according to claim 3, characterized in that When the target down-sampling pooling thread is based on the target down-sampling pooling thread, and the accumulator comparator is used to perform average pooling, maximum pooling and up-sampling pooling on the to-be-down-sampled pooling data, and the target pooling result is obtained, the method further includes: If the data type of the pooled data to be downsampled is the INT4 data type, the INT4 data type is converted to the INT9 data type; if the data type of the pooled data to be downsampled is the INT8 type, the INT8 type is converted to the INT18 type; if the data type of the pooled data to be downsampled is other data types, the type is converted to a longer data type in which the mean accumulation does not overflow.

9. A data processing device based on a multifunctional multiplexing pooling layer circuit, characterized in that: The multifunctional multiplexing pooling layer circuit includes: a data transmission module, a data receiving module and a pooling circuit, wherein the pooling circuit is respectively connected to the data transmission module and the data receiving module through an interface, and the pooling circuit includes an accumulation comparator, wherein the device includes: A judgment module, used to judge whether there is a pooling data demand, wherein the pooling data demand includes the data to be pooled, the task processing demand of the data to be pooled, and the data sending demand; A determination module, configured to use the sending module to send the data to be pooled to the pooling circuit if there is a demand for the pooled data, and determine a target pooling thread from the pooling circuit based on the task processing demand; A pooling module is used to perform average pooling, maximum pooling and upsampling pooling on the data to be pooled based on the target pooling thread and using the accumulator comparator to obtain a target pooling result, and send the target pooling result to the data receiving module based on the data sending requirement.

10. An electronic device, characterized in that: It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement a data processing method based on a multifunctional multiplexing pooling layer circuit as described in any one of claims 1 to 8.