Image serializer and electronic device
By designing a dedicated image serializer hardware module, the feature matrix data is automatically generated and output directly to the processor, which solves the problem of duplicate feature matrix data during image serialization, and achieves the effect of reducing memory and transmission overhead and improving execution efficiency.
Patent Information
- Application Number
- CN202510433888.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-04-08
AI Technical Summary
In the process of image serialization, the generated feature matrix data contains a large amount of duplicate data due to the special behavior of convolution during image serialization, which increases memory and data transmission overhead, and thus increases communication delay.
A dedicated hardware module, namely an image serializer, is designed to read original feature data from off-chip memory, and automatically generate feature matrix data, which is directly output to the multiplication calculation processor, reducing the storage and transmission of feature matrix data.
It effectively reduces memory and data transmission overhead, reduces communication delay, improves the execution efficiency of artificial intelligence tasks, and reduces the required memory resources and computing resources.
Smart Images

Figure CN119963402B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and more particularly to an image serializer and an electronic device. Background Art
[0002] Im2col (image to column) slides a template by column, converts the data within each window into a column vector, and arranges the columns into a new matrix, thereby converting the input data into matrix-form data suitable for convolution operations.
[0003] In related technologies, due to the special behavior mode of convolution, the feature matrix data obtained after the input data is processed by im2col contains a large amount of duplicate data, and the data volume far exceeds the original data volume. During the image serialization process, it is necessary to cache and transmit the feature matrix data, which not only requires a large amount of memory resources, but also increases the data transmission overhead and the computational resource overhead, resulting in communication delays. Summary of the Invention
[0004] The present invention provides an image serializer and an electronic device, which can effectively reduce the memory and data transmission overhead, effectively reduce the communication delay, improve the execution efficiency of artificial intelligence tasks, and further effectively reduce the memory resources and computational resources required for artificial intelligence tasks.
[0005] To solve the above technical problems, the present invention provides the following technical solutions:
[0006] On the one hand, the present invention provides an image serializer, including an input end, a read-write control module, a data storage module, and an output end; the input end receives read-write control parameters and original feature data; the read-write control module determines a data writing method and a data reading method based on the read-write control parameters, writes the original feature data into the data storage module according to the data writing method; under the control of the data reading method, reads corresponding data from the data storage module according to the convolution operation behavior mode to generate feature matrix data; the output end outputs the feature matrix data.
[0007] On the other hand, the present invention provides an electronic device, including an off-chip memory, a processor, and an image serializer as described in the embodiments of the present invention; wherein, the off-chip memory stores original feature data; the image serializer reads the original feature data from the off-chip memory and outputs corresponding feature matrix data; the processor performs matrix multiplication and addition calculations on the feature matrix data and the corresponding feature matrix data to obtain a result matrix.
[0008] The advantages of the technical solution provided by the present invention are as follows. The im2col operation is designed as a dedicated hardware module, namely an image columnarizer, which can be directly integrated into a neural network accelerator without relying on an external central processing unit. The communication and interaction operations between the neural network accelerator and the central processing unit are reduced, effectively reducing the communication latency and facilitating the improvement of the execution efficiency of artificial intelligence tasks. During the task execution process, the image columnarizer reads the original feature data from the off-chip memory and automatically generates the feature matrix data. Only the original feature data with a small amount of data needs to be directly cached, reducing the storage overhead of directly saving the feature matrix data in the cache, reducing the occupation of memory resources, and effectively reducing the memory overhead. In the process of converting the original feature data into the feature matrix data, only the transmission of the original feature data with a small amount of data is involved, and there is no need to transmit the feature matrix data, effectively reducing the data transmission overhead from the off-chip memory to the on-chip memory, and thus effectively reducing the memory resources and computing resources required for artificial intelligence tasks. In addition, the electronic device has corresponding advantages. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the technical solutions of the present invention or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0010] Figure 1 It is a schematic diagram of an image columnarization process in an exemplary scenario;
[0011] Figure 2 It is a schematic diagram of convolution with a dilation rate of 1 in an exemplary scenario;
[0012] Figure 3 It is a schematic diagram of convolution with a dilation rate of 2 in an exemplary scenario;
[0013] Figure 4 It is a schematic diagram of convolution with a dilation rate of 3 in an exemplary scenario;
[0014] Figure 5 It is a schematic diagram of the framework of a neural network accelerator in an exemplary application scenario provided by the present invention;
[0015] Figure 6 It is a schematic diagram of a feature matrix and a weight matrix in an exemplary application scenario provided by the present invention;
[0016] Figure 7 It is a structural framework diagram of an image columnarizer in an exemplary embodiment provided by the present invention;
[0017] Figure 8Schematic diagram of the storage method under an exemplary embodiment of the original feature data provided by the present invention;
[0018] Figure 9 Schematic diagram of the storage method under another exemplary embodiment of the original feature data provided by the present invention;
[0019] Figure 10 Structural framework diagram of another exemplary embodiment of the data storage module provided by the present invention;
[0020] Figure 11 Structural framework diagram of another exemplary embodiment of the image serializer provided by the present invention;
[0021] Figure 12 Structural framework diagram of an exemplary embodiment of the write control module provided by the present invention;
[0022] Figure 13 Schematic diagram of the write control process provided by the present invention;
[0023] Figure 14 Structural framework diagram of an exemplary embodiment of the read control module provided by the present invention;
[0024] Figure 15 Schematic diagram of the data filling method provided by the present invention;
[0025] Figure 16 Schematic diagram of the principle of the data coordinate generation circuit provided by the present invention;
[0026] Figure 17 Schematic diagram of the boundary range provided by the present invention;
[0027] Figure 18 Schematic diagram of the state transition conditions of each counter provided by the present invention;
[0028] Figure 19 Schematic diagram of the data piecing process provided by the present invention;
[0029] Figure 20 Schematic diagram of the state transition conditions of the read counter provided by the present invention;
[0030] Figure 21 Structural diagram of an exemplary embodiment of the electronic device provided by the present invention. Detailed implementation mode
[0031] In order to enable those skilled in the art to better understand the technical solution of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. Among them, the terms "first", "second", "third", "fourth", etc. in the specification and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "including" and "having" and any variations of the two are intended to cover non-exclusive inclusions. The term "exemplary" means "used as an example, embodiment or illustrative". Any embodiment described here as "exemplary" is not necessarily interpreted as being superior or better than other embodiments.
[0032] The convolution operation is to slide each convolution kernel on the feature data with a certain step size and expansion rate, and then multiply the feature data and the convolution kernel data accordingly and then accumulate them. It is a common data processing method in the field of machine learning and deep learning. Figure 1 For example, the input feature data is three-dimensional data, that is, it includes three dimensions: H (high), W (width), and C (channel). Figure 1 Use H I , W I and C I Filters (convolution filter) includes Co convolution kernels with three dimensions H, W, C. The number of channels of each convolution kernel is the same as the number of channels of the feature data, which can be represented by H F , W F and C I For example, when the step size is 1 and the dilation rate is 1, the data corresponding to the first slide of the convolution kernel on the feature data is the first row of the feature matrix. This row is multiplied and accumulated with the first column element of the weight matrix (i.e., the first convolution kernel) to obtain the first element in the upper left corner of the result matrix. F =W F =2, H I =W I =3, Co=3 as an example, the convolution kernel can slide 4 times on the feature data. Correspondingly, the number of rows in the feature matrix is 4, and the number of elements in each row is H F *W F *C I , the number of columns of the weight matrix is 3, and the number of elements in each column is H F *W F *C I The step length in the horizontal or vertical direction is usually the same. When the step length is not 1, the distance moved by each slide is no longer 1. When the expansion rate is not 1, the feature data corresponding to each element in the convolution kernel will also change accordingly, with H F =W F =3 convolution kernel in H I =WI Taking the example of sliding on the feature data with a dilation rate of = 5, Figure 2 , Figure 3 and Figure 4 During the convolution process with dilation rates of 1, 2, and 3 respectively, the black dots represent the elements of the convolution kernel. With a dilation rate of 2, the feature data elements corresponding to two adjacent elements of the convolution kernel are separated by 1 feature element. With a dilation rate of 3, the feature data elements corresponding to two adjacent elements of the convolution kernel are separated by 2 feature elements.
[0033] To improve the data processing efficiency, im2col can be used to convert the input data, such as Figure 1 the feature data, into a matrix form suitable for convolution operations, such as Figure 1 the feature matrix in. If the feature data input is image data, the image data can be converted into matrix features through im2col. When performing various artificial intelligence tasks involving convolution operations, the original convolution operation can be converted into a matrix multiplication operation through im2col, and the matrix multiplication unit of various neural network accelerators can be used for accelerated calculation. Neural network accelerators include GPU (Graphics Processing Unit), DPU (Data Processing Unit), and FPGA (Field-Programmable Gate Array). Since im2col can change the process of discontinuously reading input feature data according to the sliding window of the convolution kernel, that is, the data read is discontinuous in memory, into a process of continuously reading the feature matrix in space, it can also accelerate the calculation of the convolution operation, thereby improving the data processing efficiency and the execution efficiency of artificial intelligence tasks.
[0034] The process of the related technology performing the above-mentioned im2col-based convolution calculation operation through a neural network accelerator includes:
[0035] First, a software program is written and executed on a general-purpose processor such as a CPU (Central Processing Unit). The im2col operation is performed on the input feature data to generate a feature matrix, and the convolutional kernel data is rearranged by flattening to obtain a weight matrix. Then, the feature matrix and the weight matrix are stored in a large-capacity slow off-chip memory, such as DDR (Double Data Rate) or HBM (High Bandwidth Memory). Then, any neural network accelerator reads the feature matrix and the weight matrix from the off-chip memory into a high-speed on-chip memory with limited capacity, such as SRAM (Static Random Access Memory). Finally, the neural network accelerator sends the data in the on-chip memory to the matrix multiplication unit to perform matrix multiplication calculations. The matrix multiplication unit can be, for example, a systolic array or a multiply-accumulate tree to complete the convolution operation.
[0036] For the above process, the weight matrix is a rearrangement of the convolutional kernel data without duplicate data. Due to the special behavior of convolution, the feature matrix obtained through the im2col operation contains a large amount of duplicate data. That is to say, the number of elements contained in the feature matrix far exceeds the size of the original feature data. For example, the original feature data includes 18 elements, and the generated feature matrix contains 32 elements. Although a general-purpose processor can flexibly implement any arrangement of the input feature data or convolutional kernel data, since the feature matrix contains a large amount of duplicate data, this will lead to an increase in storage overhead and data transmission overhead. For a slow and large-capacity off-chip memory, the memory occupancy resources required for the feature matrix will increase. It is more unacceptable to store this feature matrix in the expensive but limited-capacity on-chip memory of the neural network accelerator, seriously occupying the high-speed on-chip memory resources of the accelerator. After the im2col operation is completed on a general-purpose processor such as a CPU, the feature matrix with an increased data volume needs to be transferred from the off-chip memory to the high-speed on-chip memory of the neural network accelerator. The transfer of the data volume will increase the data transmission overhead and reduce the overall computing efficiency. In addition, a neural network model usually contains multiple convolution operations. If the neural network accelerator frequently relies on an external general-purpose processor such as a CPU to complete the im2col operation, more communication delays will be introduced, causing the matrix multiply-accumulate calculation unit in the neural network accelerator to be idle and increasing the calculation delay.
[0037] In view of this, the present invention designs a dedicated hardware module for the im2col operation, that is, an image columnarizer. The image columnarizer reads the original feature data from the off-chip memory, can automatically generate feature matrix data through the internal circuit structure, and then sends the feature matrix data to the processor for multiplication and addition calculation, effectively reducing the memory and data transmission overhead, effectively reducing the communication delay, improving the execution efficiency of artificial intelligence tasks, and further effectively reducing the memory resources and computing resources required for artificial intelligence tasks. Based on the above technical solution of the present invention, combined with Figure 5 Some possible application scenarios related to the technical solution of the present invention are introduced by way of example, which may include the following content:
[0038] Integrate the image columnarizer 502 into the neural network accelerator 50. For example, the neural network accelerator can be integrated in a GPU or FPGA, and then the neural network accelerator is deployed on the server 5. For example, the FPGA can be inserted into the server through a PCIe (peripheral component interconnect express, high-speed serial computer expansion bus) slot. The neural network accelerator includes a first off-chip memory 501 and a matrix multiplication and addition calculation unit 503. The first off-chip memory 501 stores the original feature data. The data storage module of the image columnarizer 502 uses a storage array, receives the read-write control parameters and the original feature data through the input end, and uses the read-write control module to determine the data writing method and data reading method based on the read-write control parameters, and writes the original feature data into the storage array according to the data writing method. Under the control of the data reading method, the corresponding data is read from the storage array according to the convolution operation behavior mode to generate Figure 6 The feature matrix data shown, and the feature matrix data is input into the matrix multiplication and addition calculation unit 503 through the output end. The matrix multiplication and addition calculation unit 503 performs Figure 6 Multiplication and addition operations on the feature matrix data and the corresponding weight features to obtain the final result matrix. It can be seen that the storage array of the image columnarizer 502 only needs to cache the original feature data, automatically reads the original feature data stored in the storage array according to the convolution behavior mode through the internal structure, and then assembles and combines it into feature matrix data and sends it, avoiding directly storing the huge data after the im2col expansion, effectively reducing the memory and data transmission overhead. The whole process also requires the neural network accelerator to interact and communicate with the CPU of the server, effectively reducing the communication delay, realizing the accelerated training or accelerated inference of the neural network model, and being able to improve the execution efficiency of artificial intelligence tasks based on the neural network model.
[0039] It should be noted that the above application scenarios are only shown for the convenience of understanding the ideas and principles of the present invention, and the embodiments of the present invention are not limited in this regard. On the contrary, the embodiments of the present invention can be applied to any applicable scenario. After introducing the technical solutions of the present invention, the various non-limiting embodiments of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0040] First, please refer to Figure 7 , Figure 7 which is a structural framework diagram of an exemplary embodiment of the image serializer provided in this embodiment. This embodiment may include the following content:
[0041] The image serializer 502 may include an input end 703, a read-write control module 701, a data storage module 702, and an output end 704. Among them, the input end 703 may be, for example, a data interface or a user port. Through the input end 703, the read-write control parameters defined by the user for generating the corresponding hardware circuit can be received. By the custom read-write control parameters, it can be ensured that the finally generated hardware circuit is a hardware circuit that meets the user's requirements. Since the image serializer 502 is for implementing the convolution operation, the read-write control parameters should at least include the convolution operation parameters, and the convolution operation parameters at least include the dilation rate, the convolution kernel size, and the sliding step. When writing a program, the dilation rate can be expressed as params_conv_dil, the convolution kernel size can be expressed as params_conv_ksize. The size of the convolution kernel, that is, the sizes of Hf and Wf, are generally the same and can be represented by one parameter; the sliding step can be expressed as params_conv_stride, and the sliding steps in the horizontal and vertical directions are generally the same and are also represented by one parameter. As for other parameter types of the read-write control parameters, they can be determined according to the read-write control method to be implemented by the read-write control module 701 and the physical parameters of the storage components of the data storage module 702. The present invention does not limit this. In addition, the original feature data of the present invention refers to the original data that needs to be image serialized. The original feature data can be pre-stored in a specified storage hardware, which is outside the image serializer 502. For the convenience of description, it can be defined as an off-chip memory. When the image serializer 502 needs to perform an image serialization operation on the original feature data, the image serializer 502 can obtain the original feature data stored in the off-chip memory through the input end 703 and cache it in the data storage module 702. In other words, the data storage module 702 stores the original feature data input from the off-chip memory according to certain rules under the control of the read-write control module 701. Its essence is a storage hardware for caching the received original feature data so that the read-write control module 701 can read and process it. It is a storage hardware inside the image serializer 502.
[0042] In this embodiment, the read-write control module 701 determines the data writing method and the data reading method based on the read-write control parameters. The data writing method refers to the rule according to which the original feature data is stored in the data storage module 702, that is, the original feature data is written into the data storage module 702 according to the data writing method. The data reading method controls how to read data from the data storage module 702 and generate the corresponding feature matrix data based on the convolutional behavior pattern or the behavior pattern of the image serialization operation. That is, under the control of the data reading method, the corresponding data is automatically read from the data storage module 702 according to the convolutional operation behavior mode, and the feature matrix data is automatically generated. The output end 704 sends out the feature matrix data output by the read-write control module 701, for example, sends it to the PE (Processing Element) array for multiplication and addition calculation, or the multiplication and addition tree.
[0043] Among them, the image serializer 502 can be implemented using RTL (register transfer language), such as VHDL (Very High Speed Integrated Circuit Hardware Description Language) and Verilog (hardware description language). That is, the hardware code for implementing the above functions can be written using RTL language, and these hardware codes can be processed using tools to generate the corresponding physical circuit hardware. As for how to write the hardware code and which tool to use to generate the circuit, as long as the image serializer 502 can implement the functions to be implemented by the functional modules of the image serializer 502 described in the present invention, those skilled in the art can handle it flexibly, which does not affect the implementation of the present invention.
[0044] In the technical solution provided in this embodiment, the im2col operation is designed as a dedicated hardware module, that is, the image serializer 502, which can be directly integrated into the neural network accelerator without relying on an external central processing unit. The communication and interaction operations between the neural network accelerator and the central processing unit are reduced, effectively reducing the communication delay and facilitating the improvement of the execution efficiency of artificial intelligence tasks. During the task execution process, the image serializer 502 reads the original feature data from the off-chip memory and automatically generates the feature matrix data. Only the original feature data with a small amount of data needs to be directly cached, reducing the storage overhead of directly storing the feature matrix data in the cache, reducing the occupation of memory resources, and effectively reducing the memory overhead; during the process of converting the original feature data into the feature matrix data, only the transmission of the original feature data with a small amount of data is involved, and there is no need to transmit the feature matrix data, effectively reducing the data transmission overhead from the off-chip memory to the on-chip memory, and thus effectively reducing the memory resources and computing resources required for artificial intelligence tasks.
[0045] In order to further improve the processing efficiency of the image serializer 502 for the original feature data, based on the above embodiments, the present invention also optimizes the storage method of storing the original feature data in the off-chip memory, which may include the following contents:
[0046] It can be understood that the reading and writing of the original feature data by the image serializer 502 of the present invention are both realized under the control of read-write control parameters. The storage method of the original feature data determines the reading and writing of the original feature data by the image serializer 502. Therefore, when making certain agreements and changes to the storage format of the original feature data in the off-chip memory, it is also necessary to be based on the read-write control parameters. The read-write control parameters of this embodiment at least include the number of elements included in a single read-write operation. The dimension of the original feature data includes the number of channels. Taking the data storage module 702 including multiple memories as an example, combined with the image serialization behavior mode of processing data row by row, each memory of the corresponding data storage module 702 is essentially a row buffer. Correspondingly, the number of elements included in a single read-write operation can be the number of elements included in a read operation or a write operation of a single LineBuffer (row buffer), that is, the number of elements included in a single read-write operation is the number of elements included in a single read-write operation of the row buffer, which can be defined as NUM_CH. NUM_CH can adopt a default value of 32, for example. The original feature data is three-dimensional data, and the three dimensions are length Hi, width Wi, and the number of channels Ci respectively. As Figure 8 shown, in this embodiment, (Hi, Wi) is used as the coordinate, and the Ci dimension is conditionally regarded as a whole. The data segmentation parameter is determined according to the number of channels and the number of elements, and the original feature data is cut into multiple data blocks with the number of channels as a whole according to the data segmentation parameter, and then stored at the target storage location. The target storage location can be the off-chip memory, or can be refined to the storage address of the off-chip memory. Exemplarily, after splitting the data blocks, they can be stored in the order from left to right and from top to bottom. Of course, they can also be stored in other orders, which do not affect the implementation of the present invention.
[0047] In this embodiment, in order to facilitate the reading and writing of the original feature data, the number of elements included in a single read-write operation needs to match the amount of elements of the original feature data stored in one address. Based on this, the present invention is divided into two cases for storage according to the size relationship between the number of channels of the original feature data and the number of elements included in a single read-write operation:
[0048] The first storage case: The value of the number of channels of the original feature data is greater than the number of elements included in a single read / write operation. According to the number of elements, the data with the same width dimension and height dimension in the original feature data is cut into multiple data blocks, and each data block is respectively stored in each address space of the target storage location. In this embodiment, ideally, the number of elements included in each data block is equal to NUM_CH. The number of elements is the data corresponding to a coordinate (Hi, Wi), which is defined as the first type of data block. Inevitably, there will be data blocks with the number of elements less than NUM_CH, which is defined as the second data block. That is, the data length value of the first type of data block is the same as the value of the number of elements, and the data length value of the second type of data block is less than the value of the number of elements. After cutting the data with the same width dimension and height dimension in the original feature data into the first type of data block and the second type of data block according to NUM_CH, in the way that one first type of data block is stored in one address space, each first type of data block is sequentially stored in each first type of address space of the target storage location; after all the first type of data blocks are filled, in the way that each second type of data block is filled into one address space until there is no remaining storage space, each second type of data block is sequentially stored in each second type of address space of the target storage location; among them, the start and end addresses of each first type of address space are connected, and the end address of the last first type of address space is connected to the start address of the first second type of address space. As Figure 8 shown, the dimensions of the original feature data are Hi = Wi = 224. The original feature data takes (Hi, Wi) as coordinates, and the Ci dimension is conditionally regarded as a whole and stored in the order from left to right and from top to bottom. The first type of address space is address 0, address 1, etc., and the second type of address space is the address for storing the remaining data. When Ci > NUM_CH, for each position index, the data of the first NUM_CH channels is saved first. After saving the data of the first NUM_CH channels for all position indexes, the data of the next NUM_CH channels for each subsequent position index is saved, and so on until all the original feature data is stored.
[0049] The second storage case: The number of channels of the original feature data is less than or equal to the number of elements included in a single read / write operation. The length value of the padding data is determined according to the difference between the number of channels and the number of elements, and the padding data type is determined. For example, it can be padded by Padding (filling) 0. The data with the same width dimension and height dimension in the original feature data and the padding data are respectively stored in each address space of the target storage location. As Figure 9As shown, the dimensions of the original feature data are Hi = Wi = 224. The original feature data uses (Hi, Wi) as coordinates, and the Ci dimension is conditionally regarded as a whole and stored in the order from left to right and from top to bottom. When Ci ≤ NUM_CH, Padding0 is performed, and each data with the same width dimension and height dimension and the padding data together occupy 1 address space.
[0050] Based on the storage method of the original feature data recorded in the above embodiment, when reading the original feature data, what is actually read is the first NUM_CH data at the (Hi, Wi) coordinates. For example, when Ci = 60, when reading the data at address 0 from the off-chip memory, what is actually read is the first NUM_CH = 32 data at the coordinates Hi = 1, Wi = 1. When reading address 1, what is read is the first NUM_CH data at the coordinates Hi = 1, Wi = 2; when reading address 224 * 224, what is read is the last 60 - NUM_CH = 28 data at the coordinates Hi = 1, Wi = 1, and so on.
[0051] As can be seen from the above, in this embodiment, the channel number dimension of the original feature data is segmented according to the granularity of NUM_CH, and then the segmented data blocks are stored in sequence. The data storage method matches the data reading method, which is beneficial to improving the image serialization operation efficiency of the image serializer 502 for the original feature data.
[0052] In order to further improve the image serialization operation efficiency of the image serializer 502 for the original feature data, the off-chip memory can be set to actively send data to the image serializer 502. Correspondingly, the image serializer 502 includes, for example, a data acquisition port provided at the input end 703, or a port of a module that implements write control in the read-write control module 701. The off-chip memory sends the original feature data to this data acquisition port. Similarly, the data sending should match the data reading. In this embodiment, two parameters, namely the single data sending parameter, the data validity parameter, and the element bit width of a single read-write operation, can also be predefined. The element bit width of a single read-write operation is further defined based on the number of elements included in a single read-write operation to limit the bit width (bit width) of a single element. Taking the data storage module 702 including multiple memories as an example, combined with the image serialization behavior mode of processing data by rows, each memory of the corresponding data storage module 702 is essentially a row buffer. Correspondingly, the element bit width of a single read-write operation is the bit width of a single element in a single read / write operation of a LineBuffer, which can be defined as LB_ELEWIDTH. LB_ELEWIDTH can, for example, adopt a default value of 16. The single data sending parameter refers to how many elements with what width are sent each time for the original feature data from the off-chip memory. When actually writing a program, for example, this parameter can be represented by lb_data_in. The data validity parameter is used to indicate when the data sent according to the single data sending parameter is valid. When actually writing a program, for example, this parameter can be represented by lb_data_in_valid. After determining NUM_CH and LB_ELEWIDTH, combined with the image serialization behavior mode, lb_data_in can be defined as the data from the off-chip memory, sending NUM_CH elements with a width of LB_ELEWIDTH each time. Correspondingly, lb_data_in_valid indicates when the data of lb_data_in is valid. For example, when the value of lb_data_in_valid is 1, it means valid. Based on the above embodiment, when the off-chip memory stores the original feature data, when Ci is less than NUM_CH or Ci cannot be divided evenly by NUM_CH, there will be padding data. Therefore, when reading the original feature data from the off-chip memory, padding data will be read, that is, data that is not the real original feature data will be read. Based on this, the read-write control parameter at least further includes the number of valid data, which is used to represent the number of valid data in each read data when reading the original feature data from the off-chip memory. Taking the example of reading NUM_CH data each time, the number of valid data represents how many of the NUM_CH data read each time are valid data. When actually writing a program, for example, this parameter can be represented by params_conv_ci.Based on this, the data acquisition port sends a first quantity of original feature data with the width value of the data element being the element bit width of a single read / write operation to the input end 703 from the original feature data in the target storage location. The original feature data block is the original feature data sent in the current round of data transmission and is part of the original feature data stored in the off-chip memory. When it is determined that the original feature data is valid based on the data valid parameter, the data valid parameter is sent to the input end 703; the input end 703 determines the number of valid data included in the original feature data based on the number of valid data.
[0053] As can be seen from the above, in this embodiment, a port for obtaining the original feature data stored according to the agreed specification from the off-chip memory is provided. The off-chip memory sends data through this port in accordance with the agreed manner, that is, the provisions of the read / write control parameter, which facilitates the processing of the original feature data by the image serializer 502 and improves the image serialization operation efficiency of the image serializer 502 for the original feature data.
[0054] In the above embodiments, the structure of the data storage module 702 is not limited in any way. The data storage module 702 in this embodiment may be composed of a group of memories, that is, the data storage module 702 is a storage array. The memories may adopt SRAM (Static Random Access Memory). The data storage module 702 caches the original feature data from off-chip memories through an array composed of multiple independent and high-speed on-chip memories (SRAM). After the read / write control parameters define the convolution operation parameters, the number of elements NUM_CH included in a single read / write operation, and the element bit width LB_ELEWIDTH of a single read / write operation, it is also necessary to further define the storage parameters. The storage parameters at least include the total number of memories, the maximum address supported by a single memory for reading and writing, and the data volume of a single read / write operation. During actual programming, LB_CH can be used to represent the total number of memories, LB_DEPTH can be used to represent the maximum address supported by a single memory for reading and writing, and LB_WIDTH can be used to represent the data volume of a single read / write operation. LB_CH represents how many memories there are in the data storage module 702. For example, the default value 16 can be adopted; this parameter determines that the maximum values of the two dimensions (Hf, Wf) (the length and width of the convolution kernel) of the supported convolution are 16. That is to say, LB_CH memories represent that the maximum convolution kernel size of the supported convolution operation is LB_CH. LB_DEPTH represents the depth of each memory in the data storage module 702, that is, the maximum address supported by each memory for reading and writing + 1. For example, the default value 256 can be adopted. LB_WIDTH is the width of a single memory, that is, the data volume operated by a single read / write operation of the memory. Its value is LB_ELEWIDTH * NUM_CH, and the default value can be 32 * 16 = 512 bit. Considering that the convolution operation or the process of graph serialization is carried out in units of rows and is used to cache the original feature data, each memory is essentially a row buffer, and LineBuffer can be used to represent the memory. In other words, the number of memories included in the storage array is the same as the total number of memories LB_CH in the storage parameters. The total number of memories is determined according to the maximum convolution kernel size of the convolution operation; each memory is independent of each other, and the read / write data width of each memory is determined according to the number of elements and the element bit width. The total storage data volume is determined according to the maximum address supported by a single memory in the storage parameters and the corresponding read / write data width. For example, please refer to Figure 10, the data storage module 702 includes LB_CH independent SRAMs. The read / write data width of each SRAM is NUM_CH * LB_ELEWIDTH, and each SRAM can store at most LB_DEPTH data with a width of NUM_CH * LB_ELEWIDTH, that is, the maximum address range is 0 to LB_DEPTH - 1. Since there are LB_CH SRAMs, the lb_wens (write enable signal), lb_waddrs (write address signal), lb_wdatas (write data signal), lb_rens (read enable signal), lb_raddrs (read address signal), and lb_rdatas (read data signal) of the data storage module 702 all have LB_CH copies, corresponding to different SRAMs.
[0055] As can be seen from the above, the present invention uses a high-speed on-chip memory to form a storage array to store the original feature data, and the storage parameters of the storage array match the read / write control parameters, which can improve the read / write efficiency of the original feature data.
[0056] In the above embodiment, the structure of the read / write control module 701 is not limited in any way. The present invention also gives an exemplary structural manner of the read / write control module 701. In this embodiment, as Figure 11 shown, the read / write control module 701 is composed of a read control module and a write control module. The read control module interacts with both the write control module and the data storage module 702, and the write control module has data interaction with the data storage module 702. The read control module and the write control module will be introduced separately below.
[0057] In this embodiment, the write control module is connected to the input terminal 703, and the parameters input to the write control module are defined as write control parameters, that is, the read / write control parameters include the write control parameters. The write control module at least includes a counter, a write parameter input port, and a write signal output port. The write parameter input port receives the write control parameters. Among them, for the convenience of description, the counter inside the write control module is defined as the write counter. Based on the write control parameters, according to the write counter to count the sending situation of the original feature data and the storage parameters of the data storage module 702, determine the storage location information of the currently input sub-original feature data in the data storage module 702, and generate a write enable signal, a write address signal, and a write data signal according to the storage location information and the sub-original feature data; through the write signal output port, send the write enable signal, the write address signal, and the write data signal to the data storage module 702.
[0058] In this embodiment, it can be understood that the original feature data is sent to the image serializer 502 in batches. Or rather, the image serializer 502 reads the original feature data from the off-chip memory in batches or in multiple rounds. For the convenience of description, the part of the original feature data processed at the current moment in the current round is defined as the sub-original feature data, and the sub-original feature data is part of the original feature data received at the current moment. Among them, the write control parameters of this embodiment are used to control the amount of data written at a time and the total amount of data written to each storage location of the data storage module 702, including but not limited to the write start parameter, the single-data transmission parameter lb_data_in, the data valid parameter lb_data_in_valid, and the write position switching parameter. Among them, the write start parameter indicates that the writing of the LineBuffer is about to start and is used as a reset signal, which can be represented by params_write_start. The write position switching parameter indicates the relationship between the number of data writes and the storage location, and can be represented by params_lb_wswitch_cnt. Taking the data storage module 702 as a storage array as an example, the write position switching parameter refers to how many times the data storage module 702 writes data before switching to the next LineBuffer when writing the original feature data into LB_CH LineBuffers. The data valid parameter indicates whether the sub-original feature data sent according to the single-data transmission parameter is valid. After the write control parameters determine the write start parameter, the single-data transmission parameter, the data valid parameter, and the write position switching parameter, when the write control module receives the write start parameter, it resets the timer; whenever it receives the data valid parameter, it determines the storage location information of the currently input sub-original feature data in the data storage module 702 according to the write position switching parameter and the count value counted by the timer.
[0059] Exemplarily, such as Figure 12As shown, in this embodiment, the data storage module 702 includes multiple memories. The write counter may include a first write counter and a second write counter. The first counter may be represented by wline_cnt, and the second counter may be represented by lb_write_index. The single-data transmission parameter is that the data volume of the sub-raw feature data transmitted each time is the same as the read / write data width of a single memory. The write position switching parameter represents the number of valid data written to a single memory, that is, params_lb_wswitch_cnt represents how many times data is written before switching to the next LineBuffer. The Hi dimension of the raw feature data in the off-chip memory corresponds to the total number of memories included in the data storage module 702, and the Wi dimension corresponds to the depth of a single SRAM memory, that is, the LB_CH LineBuffers can store at most LB_CH rows and LB_DEPTH data per row in the raw feature data. In this way, lb_write_index can reflect the index of the SRAM memory currently being written, and wline_cnt can represent the address of the SRAM memory currently being written. When a data valid parameter is received through the write parameter input port, a write enable signal that enables the write of the target memory and disables the write enables of other memories in the data storage module 702 is generated according to the second count value, a write address signal is generated according to the first count value, and a write data signal is generated according to the sub-raw feature data.
[0060] Based on the above write control module, the working processes of wline_cnt and lb_write_index are as follows Figure 13As shown, before loading the original feature data from off-chip memory in each round, the value of the write start signal can be set to 1, that is, params_write_start == 1, indicating that the write start signal is valid and lasts for one clock cycle. At this time, the first counter and the second counter are reset to 0. Whenever a data valid parameter is received through the write parameter input port, if the first count value of the first write counter is not the write position switching parameter - 1, the value of the current first count value is increased by the first preset value; if the first count value is the write position switching parameter - 1, the first write counter is reset, and the value of the current second count value of the second write counter is increased by the second preset value, so as to represent the address of the target memory currently written by the first write counter, and the second write counter represents the index of the target memory currently written. Among them, the first preset value and the second preset value can be set according to the actual situation. For example, they can both be 1. That is, whenever the original feature data arrives, that is, whenever the data valid parameter is valid, that is, lb_data_in_valid == 1, it is judged whether the first count value of wline_cnt is equal to params_lb_wswitch_cnt - 1. If wline_cnt == params_lb_wswitch_cnt - 1, the first count value of wline_cnt is set to 0 for the next accumulation starting from 0, and at the same time, the second count value of lb_write_index is incremented by 1, that is, lb_write_index++; if wline_cnt is different from params_lb_wswitch_cnt - 1, the first count value of wline_cnt is incremented by 1, that is, wline_cnt++.
[0061] Taking the storage array as multiple SRAMs as an example, when lb_write_index is 0 and lb_data_in_valid == 1, the write enable of the first SRAM in the corresponding lb_wens is valid, and the write enables of the remaining SRAMs are invalid; the value of the first write address in lb_waddrs is assigned the value of wline_cnt, and the value of the first write data in lb_wdatas is assigned the value of lb_data_in; as the sub-original feature data is continuously acquired, wline_cnt gradually increases from 0 to params_lb_wswitch_cnt - 1. Correspondingly, data is written to the addresses 0 to params_lb_wswitch_cnt - 1 of the first SRAM. Figure 8 and Figure 9Taking the storage method shown as an example, under the write parameter control, the original feature data is read from the off-chip memory. For example, if params_lb_wswitch_cnt is set to 10, the data with coordinates from <1, 1> to <1, 10> will be read and stored in the first SRAM; then the data with coordinates from <2, 1> to <2, 10> will be read and stored in the second SRAM. At this time, lb_write_index becomes 1 and wline_cnt becomes 0, and the subsequent params_lb_wswitch_cnt data will be written into the second SRAM, and so on until the write operation is completed. When the write control operation is completed, lb_write_index indicates how many rows of original feature data have been written into the storage array in this round of write operation, and params_lb_wswitch_cnt indicates how many valid data there are in each row.
[0062] As can be seen from the above, in this embodiment, the write control parameter is used to specify how many rows of original feature data are written in each round and how many data are written in each row, and the counter is used to count how many rows of original feature data have been written and how many valid data there are in each row, so as to efficiently and accurately cache the original feature data to the image serializer 502.
[0063] In this embodiment, as Figure 11As shown, the read / write control module 701 further includes a read control module. The parameters input to the read control module are defined as read control parameters. That is to say, the read / write control parameters include the read control parameters. Here, it should be noted that the read / write control parameters include all the control parameters that need to be input to the image serializer 502. Different embodiments focus on describing the content of the current embodiment. To avoid description, the control parameters used in the read process are defined as read control parameters, and the control parameters used in the write process are defined as write control parameters. The same control parameters will be used in the read / write process. That is to say, there will be the same control parameters in the read control parameters and the write control parameters. This does not mean that the read / write control parameters will include duplicate parameters, but after the read / write control parameters are input to the image serializer 502, different functional modules will use them according to their own needs. Among them, the read control module at least includes a read parameter input port, a read data port, and a read data control circuit. The read data port is connected to the data storage module 702, and the read data control circuit is connected to the output terminal 704. The read control parameters can be received through the read parameter input port. The read data control circuit is a circuit converted from the hardware code that reads the original feature data in the data storage module 702 according to the behavior mode of convolution to generate the final feature matrix data. The function it needs to achieve is: determine the target position information of the target original feature data covered by the convolution kernel during the sliding process according to the read control parameters, determine the physical position read information corresponding in the data storage module 702 based on the target position information, and read the target original feature data from the data storage module 702 through the read data port according to the physical position read information, and splice the original feature data according to the storage method of the original feature data and the target original feature data to generate the feature matrix data. Finally, the feature matrix data is output to a specified position, such as a systolic array, through the output terminal 704. Among them, the storage method of the original feature data refers to the storage method of the original feature data in the off-chip memory and / or the storage method of the original feature data on the data storage module 702. The actual data reading position and the identification of padding data are determined according to the storage method. The actual data is the original feature data. The target position information is the Hi coordinate and Wi coordinate of the original feature data covered by the convolution kernel during the sliding process. The physical position read information refers to which memory in the data storage module 702 and which address of the memory to read the data from.
[0064] In order to use the read control circuit to determine the target position information of the target original feature data in the convolution modes with different convolution kernels, different dilation rates, and different padding sizes, the read data control circuit of this embodiment at least includes a data coordinate generation circuit, such as Figure 14As shown, the data coordinate generation circuit is responsible for sequentially generating the Hi coordinate and Wi coordinate of the original feature data covered by the convolutional kernel during the sliding process for the original feature data cached in the data storage module 702. When determining the coordinate position, since the number of channels dimension of the convolutional kernel is always the same as that of the original feature data, the Hi coordinate can be ignored, and finally the target position information is output. In this embodiment, after the write control module writes the original feature data completely into the data storage module 702, the original feature data with Hi = lb_write_index and Wi = params_lb_wswitch_cnt is cached in the data storage module 702. Since the convolution process is the process of sliding and multiplying the convolutional kernel on the original feature data, the convolutional kernel cannot exceed the right boundary and lower boundary of the valid feature data. And during the convolution process, it may be necessary to selectively pad the upper, lower, left, and right of the original feature data, so the right boundary and lower boundary of the valid feature data are variable. Based on this, the target position information of the present invention includes the right boundary position and lower boundary position of the target original feature data, where the right boundary position is the x coordinate value and the lower boundary position is the y coordinate value. The read parameter input port receives the write position switching parameter and the data write position parameter. Among them, the data write position parameter represents the data storage amount of the original feature data in the data storage module 702. Taking the storage array as an example of the data storage module, the data write position parameter refers to how many rows of original feature data are written in the storage array in this round of write operation. When the write control module sets the second counter, the data write position parameter can be represented by lb_write_index, that is, the write control module sends lb_write_index to the read control module, and the read control module uses it as the data write position parameter. Determine the right boundary position of the target original feature data according to the write position switching parameter, and determine the lower boundary position of the target original feature data according to the data write position parameter.
[0065] Exemplarily, in combination with whether to read the original feature data in a padding mode, this embodiment also gives an exemplary determination process of the target position information. The read parameter input port receives the write position switching parameter, the data write position parameter, and the padding mode read parameter. Among them, the padding mode read parameter represents reading data from the data storage module 702 in a padding mode, the padding position, and the padding length. Since padding can be performed in all directions of up, down, left, and right, and the padding is the same for up and down, and the padding is the same for left and right, and the padding lengths in the up-down direction and the left-right direction may be different, the padding length can include a first padding length and a second padding length. The first padding length is the padding length in the horizontal direction, and the second padding length is the padding length in the vertical direction when the horizontal direction padding length is the same as the vertical direction padding length. When the first padding length and the second padding length are the same, the padding length at this time can be defined as the fixed padding length. In the actual application process, the padding mode read parameter can include params_penable, params_pleft_en, parmas_pright_en, parmas_pup_en, parmas_pdown_en, params_conv_phsize, and params_conv_pvsize. Among them, params_penable indicates whether to enable the Padding mode when reading data from the data storage module 702 or the LineBuffer. params_pleft_en indicates whether to insert Padding on the left side of the data when reading data from the data storage module 702; parmas_pright_en indicates whether to insert Padding on the right side of the data when reading data from the data storage module 702; parmas_pup_en indicates whether to insert Padding on the upper side of the data when reading data from the data storage module 702; parmas_pdown_en indicates whether to insert Padding on the lower side of the data when reading data from the data storage module 702. params_conv_phsize represents the size of the unilateral Padding when padding in the horizontal direction (i.e., left or right); params_conv_pvsize represents the size of the unilateral Padding when padding in the vertical direction (i.e., upper or lower). Perform a Padding (fill with 0) operation on the original feature data based on the above padding mode read parameter, such as Figure 15As shown, the two parameters params_conv_phsize and params_conv_pvsize indicate how many zeros are padded in the Hi and Wi dimensions of the actual data; params_penable represents the overall switch for whether to perform Padding, and params_pleft_en, parmas_pright_en, parmas_pup_en, and parmas_pdown_en are the sub-switches for the four directions. When the corresponding switch is valid, that is, when the values of params_pleft_en, parmas_pright_en, parmas_pup_en, and parmas_pdown_en are 1, the corresponding positions will be padded with the corresponding number of zeros according to the values set by params_conv_phsize and params_conv_pvsize.
[0066] Based on the above read control parameters, when the filling position is left filling and right filling, the right boundary position of the target original feature data is determined according to the write position switching parameter and the double first filling length; when the filling position is left filling or right filling, the right boundary position of the target original feature data is determined according to the write position switching parameter and the first filling length; when the filling position is not left filling and not right filling, the right boundary position of the target original feature data is determined according to the write position switching parameter. That is, when params_penable is invalid, the right boundary is params_lb_wswitch_cnt, and the lower boundary is lb_write_index. When params_penable is valid and both params_pleft_en and parmas_pright_en are valid, the right boundary is params_lb_wswitch_cnt + 2 * params_conv_phsize. When only one of params_pleft_en and parmas_pright_en is valid, the right boundary is params_lb_wswitch_cnt + params_conv_phsize. When both params_pleft_en and parmas_pright_en are invalid, the right boundary is params_lb_wswitch_cnt. When the filling position is upper filling and lower filling, the lower boundary position of the target original feature data is determined according to the data write position parameter and the double second filling length; when the filling position is upper filling or lower filling, the lower boundary position of the target original feature data is determined according to the data write position parameter and the second filling length; when the filling position is not upper filling and not lower filling, the lower boundary position of the target original feature data is determined according to the data write position parameter. That is, when both parmas_pup_en and parmas_pdown_en are valid, the lower boundary is lb_write_index + 2 * params_conv_pvsize. When only one of parmas_pup_en and parmas_pdown_en is valid, the lower boundary is lb_write_index + params_conv_pvsize. When both parmas_pup_en and parmas_pdown_en are invalid, the lower boundary is lb_write_index.
[0067] When the above embodiments determine the lower boundary and the right boundary according to different filling methods, the target position information, that is, the corresponding xy coordinates, can be obtained based on the currently received write position switching parameter, data write position parameter, and filling method read parameter.
[0068] As can be seen from the above, the data coordinate generation circuit in this embodiment generates corresponding xy data coordinates within the lower and right boundaries according to the received read control parameters. To facilitate the understanding of the principle of the data coordinate generation circuit of the present invention by those skilled in the art, the following content may be included:
[0069] The data coordinate generation circuit may set a first read counter, a second read counter, a third read counter, a fourth read counter, a fifth read counter, and a sixth read counter. The x coordinate of the original feature data is determined according to the sum of the values of the third read counter and the fifth counter, and the y coordinate of the original feature data is determined according to the sum of the values of the fourth read counter and the sixth counter. Among them, the counting interval of the third read counter and the fourth counter is +params_conv_dil each time, and the counting interval of other counters is +1 each time.
[0070] In this embodiment, the first read counter is used to represent the number of data index times in the x direction inside the convolution kernel during the current sliding, which can be represented by x_index_cnt1. The second read counter represents the number of data index times in the y direction inside the convolution kernel during the current sliding, which can be represented by y_index_cnt1. The third read counter is used to represent the data coordinate in the x direction inside the convolution kernel, which can be represented by x_index_cnt1_with_dil. The fourth read counter is used to represent the data coordinate in the y direction inside the convolution kernel, which can be represented by y_index_cnt1_with_dil. The fifth read counter is used to represent the overall x coordinate position of the current sliding window in the padded feature data, which can be represented by x_index_cnt2. The sixth read counter is used to represent the overall y coordinate position of the current sliding window in the padded feature data, which can be represented by y_index_cnt2. The two together represent the coordinates of the upper left corner element of the convolution kernel in the padded feature data. A seventh counter may also be set. The seventh counter is used to generate a working signal, which can be represented by index_en. When its value is 1, it means starting to work. When the read start signal is detected, this signal indicates that reading the LineBuffer is about to start and is used as a reset signal. That is, when params_lb_read_start == 1, index_en = 1. When y_index_cnt2_finish, then index_en changes from 1 to 0. The change process of these seven counters is as Figure 18 shown Figure 18The "!" in it represents the logical NOT operator. Among them, params_lb_read_start resets all counters to 0 and sets index_en to 1; right_bound is the right boundary, right_bound_nw is the right boundary of the next sliding window, down_bound is the lower boundary, and down_bound_nw is the lower boundary of the next sliding window; the counter starts to work when it detects that index_en is 1.
[0071] Please refer to Figure 16 and Figure 17 , for example, lb_write_index = 6, params_lb_wswitch_cnt = 6, with the size of the convolutional kernel params_conv_ksize being 3*3, the stride params_conv_stride and the dilation rate params_conv_dil both being 1. Through the write operation, the first 6 rows of the original feature data are saved in the first 6 memories of the LB_CH memories in the storage array. Each memory includes 6 data, and it is set to perform Padding on the top, bottom, left, and right of the original feature data, with the size being 2, that is, the fixed padding length params_pvsize = params_phsize = 2. In each sliding, such as Figure 16As shown, the data inside the convolution kernel is indexed in a zigzag pattern. The x - direction coordinate (i.e., x_index_cnt1) needs to be counted params_conv_ksize times in each round, and a total of params_conv_ksize rounds are counted. At the end of each round of counting, an x_index_cnt1_finish (the first counter end signal) signal is generated, that is, x_index_cntl_finish=(index_en&&x_index_cntl==(params_conv_ksize - 1)), and at the same time, x_index_cnt1 is reset to 0. When the x - direction coordinate finishes one round of counting, that is, when an x_index_cnt1_finish signal is generated, the y - direction coordinate (i.e., y_index_cnt1) is counted once, and a total of params_conv_ksize times are counted. When the y - direction coordinate is counted params_conv_ksize times and x_index_cnt1_finish is generated, a y_index_cnt1_finish (the second counter end signal) signal is generated, and at the same time, y_index_cnt1 is reset to 0, that is, y_index_cntl_finish=(x_index_cntl_finish&&y_index_cntl==(params_conv_ksize - 1)). In this way, the counters x_index_cnt1 and y_index_cnt1 can record the number of times of xy data indexing inside the convolution kernel during this sliding, and complete the change of data indexing inside the convolution window. In each sliding, when the dilation rate is not 1, the difference between the x - direction coordinate and the y - direction coordinate of the adjacent internal data of the convolution kernel is params_conv_dil. Therefore, the counters x_index_cnt1 and y_index_cnt1 represent the number of times of data indexing in the x and y directions inside the convolution kernel during this sliding. In this embodiment, the counters x_index_cnt1_with_dil and y_index_cnt1_with_dil are used to record the data coordinates in the x and y directions inside the convolution kernel during this sliding. For each sliding between each other: after determining the index changes in the x and y directions inside the convolution kernel, this embodiment uses the counters x_index_cnt2 and y_index_cnt2 to record the overall x - coordinate position and y - coordinate position in the feature data after Padding of the current sliding window. The two together constitute the coordinates of the upper - left - corner element of the convolution kernel in the feature data after Padding.The initial values of x_index_cnt2 and y_index_cnt2 are both 0. Whenever a y_index_cnt1_finish signal is generated, that is, every time a round of indexing of the feature elements within the convolution kernel window is completed and the next sliding window does not exceed the right boundary, x_index_cnt2 is incremented by params_conv_stride based on its current value. When the y_index_cnt1_finish signal is generated but the next sliding window will exceed the right boundary, x_index_cnt2 is reset to 0. When the x_index_cnt2_finish (the fifth counter end signal) signal is generated, x_index_cnt2_finish = (y_index_cntl_finish && right_bound_nw > right_bound), it indicates that the sliding window of the convolution kernel needs to move downward and return to the left boundary. Whenever the x_index_cnt2_finish signal is generated once, it means that the sliding window has moved to the position allowed by the right boundary in the X direction and the next sliding window does not exceed the lower boundary, and y_index_cnt2 is incremented by params_conv_stride based on its current value. When the x_index_cnt2_finish signal is generated and the next sliding window will exceed the lower boundary, the y_index_cnt2_finish (the sixth counter end signal) signal is generated, y_index_cnt2_finish = (x_index_cnt2_finish && down_bound_nw > down_bound), indicating that the im2col operation for the original feature data written to the data storage module 702 is completely finished. The finally output x coordinate for reading the original feature data is: x_index = x_index_cnt1_with_dil + x_index_cnt2, and the y coordinate of the original feature data is: y_index = y_index_cnt1_with_dil + y_index_cnt2.
[0072] As can be seen from the above, in this embodiment, multiple counters are used to count the sliding situation during the convolution operation, so as to generate corresponding xy data coordinates within the range of the lower boundary and the right boundary according to the received read control parameters.
[0073] In order to use the read control circuit to determine the physical location read information of the target original feature data in convolution modes with different padding sizes, so as to efficiently and accurately read the original feature data cached, the read data control circuit of this embodiment at least includes a read address generation circuit. When the target location is valid, the read address generation circuit determines the physical address according to the target location information of the target original feature data and the storage method of the original feature data; determines the physical location read information of the data storage module 702 according to the physical address and the target location information, and generates a data read signal according to the physical location read information; the data read signal is sent to the read data port. In this embodiment, the read address generation circuit can receive x_index, y_index, and index_valid in the data coordinate generation circuit, and generate a signal for reading the original feature data in the data storage module. Among them, index_valid is used to indicate that the target location is valid.
[0074] Exemplarily, in combination with whether to read the original feature data in a padding manner, this embodiment also gives an exemplary determination process of the physical address. When reading data from the data storage module 702 in a padding manner and the padding position is left padding, the x coordinate in the physical address is the difference between the x coordinate of the target position information and the first padding length; when the padding position is not left padding, the x coordinate in the physical address is the x coordinate of the target position information; when the padding position is upper padding, the y coordinate in the physical address is the difference between the y coordinate of the target position information and the second padding length; when the padding position is not upper padding, the y coordinate in the physical address is the y coordinate of the target position information; when no padding element is added during the storage of the original feature data, that is, when reading data from the data storage module 702 without using the padding method, the x coordinate in the physical address is the x coordinate of the target position information, and the y coordinate in the physical address is the y coordinate of the target position information. In this embodiment, according to params_penable, params_pleft_en, and parmas_pup_en, the x and y coordinates of the physical address are calculated, which can be expressed as real_x_index and real_y_index. When params_penable = 1 and params_pleft_en = 1, then real_x_index = x_index - params_conv_phsize; otherwise, real_x_index = x_index. For example, when the left padding length is 2, actually when x_index is 0 or 1, what it refers to is not the feature data stored in the data storage module 702, but the 0 data of Padding; at this time, there is no need to read the data storage module 702. When params_penable = 1 and params_pup_en = 1, then real_y_index = y_index - params_conv_pvsize; otherwise, real_y_index = y_index.
[0075] After the physical address is determined through the above embodiments, considering the existence of padding data, it is necessary to further determine whether the generated physical address is within the range of real and valid feature data. If it is within the range of real and valid feature data, it is necessary to read the data storage module 702. If it is not within the range of real and valid feature data, it is not necessary to read the data storage module 702. For example, it can be represented by x_index_select and y_index_select. x_index_select indicates being within the range of valid feature data in the x-axis direction, and y_index_select indicates being within the range of valid feature data in the y-axis direction. The data coordinate generation circuit continuously generates x_index and y_index, and the read address generation circuit continuously updates real_x_index, real_y_index, x_index_select, and y_index_select according to x_index and y_index. The read address generation circuit reads the original feature data stored in the data storage module 702 according to the generated series of data coordinates in accordance with the convolutional behavior pattern.
[0076] If the left padding is not enabled, when the x - coordinate value of the physical address is less than the write position switching parameter, it is within the range of valid feature data in the x - axis direction; when the left padding is enabled, when the x - coordinate of the target position information is greater than or equal to the first padding length, and the x - coordinate value of the physical address is less than the write position switching parameter, it is within the range of valid feature data in the x - axis direction; if the upper padding is not enabled, when the y - coordinate value of the physical address is less than the data write position parameter, it is within the range of valid feature data in the y - axis direction; when the upper padding is enabled, when the y - coordinate of the target position information is greater than or equal to the second padding length, and the x - coordinate value of the physical address is less than the data write position parameter, it is within the range of valid feature data in the y - axis direction. That is to say, this embodiment can determine the lower - boundary detection method and the upper - boundary detection method. For the lower - boundary detection method in the x - direction, it is that the LeftPadding is not enabled or the LeftPadding is enabled AND x_index>=params_conv_phsize; for the lower - boundary detection method in the y - direction, it is that the UpPadding is not enabled or the UpPadding is enabled AND y_index>=params_conv_pvsize. For the upper - boundary detection method in the x - direction, it is real_x_index<params_lb_wswitch_cnt; for the upper - boundary detection method in the y - direction, it is real_y_index<lb_writing_index. For x_index_select, during the lower - boundary detection process, if there is LeftPadding, skip the LeftPadding part; during the upper - boundary detection process, if there is RightPadding, do not need to read the data storage module. For y_index_select, during the lower - boundary detection process, if there is UpPadding, skip the UpPadding part; during the upper - boundary detection process, if there is DownPadding, skip the DownPadding part. That is to say, when it is satisfied that it is within the range of valid feature data in both the x - axis direction and the y - axis direction, a data read signal is generated according to the storage position of the target original feature data in the data storage module 702 and the corresponding read address as the physical position read information. In the actual application process, the following relational expression can represent the whole process:
[0077] x_index_select = ((LeftPadding is not enabled) OR (LeftPadding is enabled AND x_index >= params_conv_phsize)) AND (real_x_index < params_lb_wswitch_cnt); y_index_select = ((UpPadding is not enabled) OR (UpPadding is enabled AND y_index >= params_conv_pvsize)) AND (real_y_index < lb_writing_index).
[0078] After the physical address is determined, the storage location of the target original feature data in the data storage module 702 can be determined according to the y-coordinate value of the physical address; the read address of the target original feature data in the storage location can be determined according to the x-coordinate value of the physical address; the storage location of the target original feature data in the data storage module 702 and the corresponding read address are used as the physical location read information. Taking the data storage module 702 as a storage array, and the storage array includes multiple SRAMs as an example, when both x_index_select and y_index_select are valid, the SRAM in the data storage module 702 needs to be read. Among the LB_CH SRAMs, which SRAM to read is determined according to real_y_index. During the application process, the lb_rens (read enable) of the corresponding SRAM can be set to 1 to achieve this. When the SRAM to be read is determined, the data at the real_x_index address is read. During the actual application process, the corresponding lb_raddrs (read address) can be set to real_x_index.
[0079] In order to enable the read control circuit to support convolution modes with different convolution kernels, different dilation rates, and different padding sizes, the read data control circuit of this embodiment at least includes a padding data generation circuit. The padding data generation circuit generates padding data according to the received target position information, physical address, padding mode parameter, data write position parameter, write position switching parameter, and whether it is within the valid feature data range signal, and outputs a valid padding signal and a padding data signal according to the padding data.
[0080] In this embodiment, the read address generation circuit sends the real_x_index, real_y_index, and x_index_select signals to the padding data generation circuit, and the data coordinate generation circuit sends the x_index and y_index to the padding data generation circuit. The padding data generation circuit generates padding data, such as params_penable, params_pleft_en, params_pright_en, parmas_pup_en, params_pdown_en, such as generating 0 data for Padding, according to the continuously updated x_index, y_index, real_x_index, real_y_index, x_index_select, and the set padding mode read parameters. The generation process of the padding data is as follows: If left padding is enabled, when the x coordinate value of the target position information is less than the fixed padding length, the corresponding padding data is generated. For example, when LeftPadding is enabled (i.e., both params_penable and params_pleft_en are 1), and when x_index is less than params_phsize, 0 data for the corresponding position of LeftPadding is generated. If right padding is enabled, when the x coordinate of the physical address is greater than or equal to the write position switching parameter, the corresponding padding data is generated; for example, when RightPadding is enabled (i.e., both params_penable and params_pright_en are 1), and when real_x_index >= params_lb_wswitch_cnt, 0 data for the position filled by RightPadding is generated. If up padding is enabled, when it is within the range of valid feature data in the x-axis direction and the y coordinate value of the physical address is less than the second padding length, the corresponding padding data is generated; for example, when UpPadding is enabled (i.e., both params_penable and params_pup_en are 1), and when x_index_select is valid and y_index < params_conv_pvsize, 0 data for the corresponding position of UpPadding is generated. If down padding is enabled, when it is within the range of valid feature data in the x-axis direction and the y coordinate value of the physical address is greater than or equal to the data write position parameter, the corresponding padding data is generated. For example, when DownPadding is enabled (i.e., both params_penable and params_pdown_en are 1), and when x_index_select is valid and real_y_index >= lb_writing_index, 0 data for the corresponding position of DownPadding is generated.
[0081] In order to realize that the read control circuit supports convolution modes with different convolution kernels, different expansion rates, and different padding sizes, the read data control circuit of this embodiment includes at least a matrix data generation circuit. The above-mentioned padding data generation circuit will output padding_valid and padding_data signals, padding_valid indicates a valid signal of the padding method currently adopted, and padding_data is the padding data generated according to different padding conditions. The padding data generation circuit will send the generated padding_valid and padding_data signals to the matrix data generation circuit, and the matrix data generation circuit also receives the real feature data from the data storage module 702 read by the read address generation circuit, collects these data, and then sends them in a specified manner, such as the dimension size of NUM_CH. Since the padding data generation circuit and the read data storage module 702 will not be performed at the same time, the matrix data generation circuit will perform simple signal selection based on the respective valid signals for the real feature data from the data storage module 702 read by the read address generation circuit and the padding_data of the padding data generation circuit, and obtain the collected valid signal (lb_data_valid, read data valid signal) and data signal (lb_data, here refers to the read data signal).
[0082] In this embodiment, the read control parameters include at least the number of valid data and the number of elements included in a single read and write operation. Whenever the target original feature data is read, a corresponding number of valid original feature data is selected from the target original feature data according to the number of valid data and stored in the shift register; whenever the amount of valid original feature data in the shift register at the current moment is equal to the number of elements, the valid original feature data of the shift register at the current moment is sent as the feature matrix data at the current moment, and a read data valid signal is output at the same time.
[0083] In order to avoid overflow, the storage space of the shift register is greater than or equal to the preset overflow value. The preset overflow value is determined according to the number of elements. For example, if the number of elements is NUM_CH, the preset overflow value can be NUM_CH*2. Correspondingly, the memory space of the shift register can be NUM_CH*2. After NUM_CH data are gathered in the shift register, they will be sent once, so the shift register will not overflow. For example, if Figure 19As shown, each time lb_data contains NUM_CH data, and only params_conv_ci data out of the NUM_CH data are valid. Therefore, splicing is required, that is, NUM_CH valid data need to be collected before sending. Whenever lb_data_valid is valid, that is, when there is new lb_data, params_conv_ci valid data are obtained from lb_data. For example, they can be sequentially stored in a shift register with a size of NUM_CH * 2 in the order from right to left. At the same time, the old data in the shift register at this time needs to be shifted to the right by params_conv_ci data lengths.
[0084] To accurately control the sending of feature matrix data, a counter can be used to record the number of valid data stored in the shift register. For ease of description, the counter used by the read control module is defined as the read counter, which can be represented by write_cnt. Correspondingly, the shift register is connected to the read counter, and the read counter records the number of valid original feature data in the shift register at the current moment. When a read start signal is received, the read counter is reset. When the read data valid signal is valid and the amount of valid original feature data in the shift register at the current moment is not equal to the number of elements, the third count value of the read counter is adjusted to the sum of the current third count value and the number of valid data; when the read data valid signal is valid, the amount of valid original feature data in the shift register at the current moment is equal to the number of elements, and a round of indexing of the feature elements within the convolution kernel window has not been completed, the third count value of the read counter is adjusted to the difference between the number of valid data and the number of elements; when the read data valid signal is valid, the amount of valid original feature data in the shift register at the current moment is equal to the number of elements, and a round of indexing of the feature elements within the convolution kernel window has been completed, the third count value of the read counter is adjusted based on the current third count value, and the increased value is the difference between the number of valid data and the remaining number of elements. When the amount of valid original feature data in the shift register at the current moment is equal to the number of elements, if the actual amount of valid original feature data in the shift register at the current moment is less than the number of elements, the number of filled sending data is determined according to the difference between the number of elements and the remaining number of elements. Corresponding filled sending data are generated according to the number of filled sending data, and the valid original feature data in the shift register at the current moment and the filled sending data are used as the feature matrix data at the current moment for sending.
[0085] In this embodiment, output_once_en is defined to indicate whether the number of data in the shift register has been assembled enough to NUM_CH for one transmission, that is, when output_once_en is valid, it indicates whether the number of data in the shift register has been assembled enough to NUM_CH for one transmission. y_index_cnt1_finish (the second counter end signal) is used to indicate whether the current sliding window is the last successful assembly. When y_index_cnt1_finish is valid, the current sliding window has been the last time to assemble enough NUM_CH valid data. Define params_conv_ci_last (the remaining number of elements) to represent that when the original feature data of a convolution window is assembled in units of NUM_CH data, the last assembly is less than NUM_CH, that is, the remaining number of elements represents the number of valid original feature data remaining when the original feature data of a convolution window is assembled in units of the number of elements and the last assembly is less than the number of elements. For example, when the convolution kernel is 3*3 and Ci is 10, the data covered by one sliding is 3*3*10 = 90. After assembling twice with NUM_CH = 32 data, there are only 26 data in the last time, that is, params_conv_ci_last = 26. The logic of write_cnt recording the number of valid data in the shift register is as Figure 20 shown, which realizes real-time reflection of the available data volume in the shift register: when params_lb_read_start (read start signal) is valid, write_cnt is reset to 0; when lb_data_valid is valid and output_once_en is invalid, write_cnt is incremented by params_conv_ci; when lb_data_valid is valid, output_once_en is valid, and y_index_cnt1_finish is invalid, write_cnt += (params_conv_ci - NUM_CH), when lb_data_valid is valid, output_once_en is valid, and y_index_cnt1_finish is valid, write_cnt += (params_conv_ci - params_conv_ci_last).
[0086] As can be seen from the above, the read control module of this embodiment includes a data coordinate generation circuit, a read address generation circuit, a padding data generation circuit, and a matrix data generation circuit. The data coordinate generation circuit realizes generating the coordinates of the original feature data in the im2col mode through a custom logic. The coordinates continuously generated by the data coordinate generation circuit, combined with the read control parameters, are sent to the read address generation circuit and the padding data generation circuit, realizing the support for convolution modes with different convolution kernels, different dilation rates, and different Padding sizes. Finally, the matrix data generation circuit combines the real feature data read from the on-chip cache and the padding data, forms the feature data, and sends it outwards.
[0087] Finally, the present invention also provides an electronic device, which can be used as a neural network model accelerator to efficiently complete the relevant training tasks and inference tasks of the neural network model. As Figure 21 shown, the electronic device may include an off-chip memory 211, a processor 212, and an imager 502. Among them, the off-chip memory 211 stores the original feature data; the imager 502 reads the original feature data from the off-chip memory 211 and outputs the corresponding feature matrix data; the processor 212 performs matrix multiplication and addition calculations on the feature matrix data and the corresponding feature matrix data to obtain a result matrix. The processor 212 may be, for example, a systolic array or a multiply-accumulate tree, or other computing devices capable of performing multiply-accumulate operations, which do not affect the implementation of the present invention. In addition, in this embodiment, the processor 212 and the imager 502 may be deployed in the neural network model accelerator, such as an FPGA or a GPU. The off-chip memory 211 may be a memory outside the neural network model accelerator. That is to say, the off-chip memory 211 and the neural network model accelerator together constitute the electronic device. Of course, the off-chip memory 211, the processor 212, and the imager 502 may all be deployed in the neural network model accelerator, such as an FPGA or a GPU. That is to say, the off-chip memory 211 may be a memory inside the neural network model accelerator. That is to say, the neural network model accelerator alone can constitute the electronic device.
[0088] As can be seen from the above, the imager of this embodiment is integrated into the electronic device as a dedicated hardware module. The memory of the electronic device only needs to store a small amount of original feature data, and there is no need to cache the entire feature matrix data, avoiding directly storing the huge data after im2col expansion, having low requirements for its memory resources, and also being able to reduce the data transmission overhead from the off-chip memory to the on-chip memory. During the imager operation process of the electronic device, it does not need to rely on an external central processor, reducing the communication and interaction operations between the neural network accelerator and the central processor, effectively reducing the communication delay, being beneficial to improving the execution efficiency of artificial intelligence tasks, and further effectively reducing the memory resources and computing resources required for artificial intelligence tasks.
[0089] The above has introduced in detail an image serializer and an electronic device provided by the present invention. Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. Whether the units and algorithm steps of each example described in the disclosed embodiments are executed in the form of electronic hardware or computer software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, and such implementation should not be considered as exceeding the scope of the present invention. Without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the present invention.
Claims
1. An image arrayer, characterized in that: It includes an input terminal, a read-write control module, a data storage module and an output terminal; The input end receives read-write control parameters and original feature data; The read-write control module determines a data writing mode and a data reading mode based on the read-write control parameters, and writes the original feature data into the data storage module according to the data writing mode; Under the control of the data reading mode, corresponding data is read from the data storage module according to the convolution operation behavior mode to generate feature matrix data; The output terminal outputs the characteristic matrix data; Wherein, the write parameter input port of the read-write control module receives a single data transmission parameter, and the read-write control module includes a first write counter and a second write counter; whenever a data valid parameter is received through the write parameter input port, if the first count value of the first write counter is not the write position switching parameter -1, the value of the current first count value is increased by a first preset value; if the first count value is the write position switching parameter -1, the first write counter is reset, and the value of the current second count value of the second write counter is increased by a second preset value, so that the first write counter indicates the address of the target memory currently being written, and the second write counter indicates the index of the target memory currently being written; Among them, the data storage module includes multiple memories, the single data sending parameter is that the data amount of the sub-original feature data sent each time is the same as the read and write data width of a single memory, the data validity parameter indicates whether the sub-original feature data sent according to the single data sending parameter is valid; the write position switching parameter indicates the relationship between the number of data write times and the storage position.
2. The image arrayer according to claim 1, characterized in that: The read / write control parameter includes at least the number of elements included in a single read / write operation, and the dimension of the original feature data includes the number of channels; Determine a data segmentation parameter according to the number of channels and the number of elements; According to the data segmentation parameters, the original feature data is stored in a target storage location.
3. The image arrayer according to claim 2, characterized in that: The number of channels is less than or equal to the number of elements, and each data and padding data having the same width dimension and height dimension in the original feature data are stored in each address space of the target storage location respectively; the length value of the padding data is the difference between the number of channels and the number of elements; The number of channels is greater than the number of elements. According to the number of elements, the data with the same width dimension and height dimension in the original feature data is cut into multiple data blocks, and each data block is stored in each address space of the target storage location respectively.
4. The image arrayer according to claim 1, characterized in that: The read / write control parameters include at least the number of elements included in a single read / write operation, the element width and storage parameters of a single read / write operation; The data storage module is a storage array, the number of memories included in the storage array is the same as the total number of memories in the storage parameters, and the total number of memories is determined according to the maximum convolution kernel size of the convolution operation; Each memory is independent of each other, and the read and write data width of each memory is determined according to the number of elements and the element bit width, and the total amount of stored data is determined according to the maximum address supported by a single memory in the storage parameters for reading and writing, and the corresponding read and write data width.
5. The image arrayer according to any one of claims 1 to 4, characterized in that: The read-write control module includes a write control module connected to the input end, the read-write control parameters include write control parameters, and the write control module at least includes a write counter, a write parameter input port and a write signal output port; The write parameter input port receives the write control parameter, and the write control parameter controls the amount of data written at a single time and the total amount of data written to each storage location of the data storage module; Based on the write control parameter, according to the write counter statistics the sending status of the original feature data and the storage parameter of the data storage module, determine the storage location information of the currently input sub-original feature data in the data storage module, and generate a write enable signal, a write address signal and a write data signal according to the storage location information and the sub-original feature data; sending the write enable signal, the write address signal and the write data signal to the data storage module through the write signal output port; The sub-original feature data is part of the original feature data received at the current moment.
6. The image arrayer according to any one of claims 1 to 4, characterized in that: The read-write control module includes a read control module, and the read-write control parameter includes a read control parameter; the read control module at least includes a read parameter input port, a read data port and a read data control circuit, the read data port is connected to the data storage module, and the read data control circuit is connected to the output port; The read parameter input port receives the read control parameter; The read data control circuit determines the target position information of the target original feature data covered by the convolution kernel during the sliding process according to the read control parameters, determines the physical position read information corresponding to the data storage module based on the target position information, and reads the target original feature data from the data storage module through the read data port according to the physical position read information, and generates feature matrix data according to the storage method of the original feature data and the target original feature data.
7. The image arrayer according to claim 6, characterized in that: The data reading control circuit at least includes a data coordinate generating circuit for outputting target position information; the target position information includes the right boundary position and the lower boundary position of the target original feature data, wherein the right boundary position is an x coordinate value and the lower boundary position is a y coordinate value; When the read control parameter does not include a fill mode read parameter, the read parameter input port receives a write position switching parameter and a data write position parameter; the write position switching parameter indicates the relationship between the number of data writes and the storage position; the data write position parameter indicates the data storage amount of the original feature data in the data storage module; the right boundary position of the target original feature data is determined according to the write position switching parameter, and the lower boundary position of the target original feature data is determined according to the data write position parameter; When the read control parameter includes a fill mode read parameter, the read parameter input port receives a write position switch parameter, a data write position parameter and the fill mode read parameter; the fill mode read parameter indicates that data is read from the data storage module in a fill mode, including a fill position and a fill length; the fill length includes a first fill length and a second fill length; When the filling position is left filling and right filling, the right boundary position of the target original feature data is determined according to the write position switching parameter and the double first filling length; when the filling position is left filling or right filling, the right boundary position of the target original feature data is determined according to the write position switching parameter and the first filling length; when the filling position is neither left filling nor right filling, the right boundary position of the target original feature data is determined according to the write position switching parameter; When the filling position is upper filling and lower filling, the lower boundary position of the target original feature data is determined according to the data write position parameter and the double second filling length; when the filling position is upper filling or lower filling, the lower boundary position of the target original feature data is determined according to the data write position parameter and the second filling length; when the filling position is neither upper filling nor lower filling, the lower boundary position of the target original feature data is determined according to the data write position parameter.
8. The image arrayer according to claim 6, characterized in that: The read data control circuit at least includes a read address generation circuit; The read address generation circuit determines a physical address according to target position information of the target original feature data and a storage method of the original feature data when the target position is valid; determines physical position read information of the data storage module according to the physical address and the target position information, and generates a data read signal according to the physical position read information; and the data read signal is sent to the read data port.
9. The image arrayer according to claim 8, characterized in that: Determining a physical address according to target location information of the target original feature data and a storage method of the original feature data includes: When data is read from the data storage module in a filling manner, and the filling position is left filling, the x coordinate in the physical address is the difference between the x coordinate of the target position information and the first filling length; when the filling position is not left filling, the x coordinate in the physical address is the x coordinate of the target position information; when the filling position is upper filling, the y coordinate in the physical address is the difference between the y coordinate of the target position information and the second filling length; when the filling position is not upper filling, the y coordinate in the physical address is the y coordinate of the target position information; When the original feature data is stored without adding padding elements, the x coordinate in the physical address is the x coordinate of the target position information, and the y coordinate in the physical address is the y coordinate of the target position information.
10. The image arrayer according to claim 9, characterized in that: Generating a data read signal according to the physical position read information includes: If the left padding is not enabled, when the x-coordinate value of the physical address is less than the write position switching parameter, it is located within the valid feature data range in the x-axis direction; if the left padding is enabled, when the x-coordinate of the target position information is greater than or equal to the first padding length, and the x-coordinate value of the physical address is less than the write position switching parameter, it is located within the valid feature data range in the x-axis direction; If the upper side padding is not enabled, when the y coordinate value of the physical address is less than the data write position parameter, it is located within the valid feature data range in the y-axis direction; if the upper side padding is enabled, when the y coordinate of the target position information is greater than or equal to the second padding length, and the x coordinate value of the physical address is less than the data write position parameter, it is located within the valid feature data range in the y-axis direction; When both the x-axis direction and the y-axis direction are within the valid feature data range, a data read signal is generated according to the storage position of the target original feature data in the data storage module and the corresponding read address as physical position read information.
11. The image arrayer according to claim 6, characterized in that: The read data control circuit includes at least a fill data generation circuit; The filling data generating circuit generates filling data according to the received target position information, physical address, filling mode parameter, data write position parameter, write position switching parameter, and whether it is in the valid characteristic data range signal, and outputs a valid filling signal and a filling data signal according to the filling data; Among them, the generation process of the filling data is as follows: if the left side filling is enabled, when the x-coordinate value of the target position information is less than the fixed filling length, the corresponding filling data is generated; if the right side filling is enabled, when the x-coordinate value of the physical address is greater than or equal to the write position switching parameter, the corresponding filling data is generated; if the upper side filling is enabled, when it is within the valid feature data range in the x-axis direction and the y-coordinate value of the physical address is less than the second filling length, the corresponding filling data is generated; if the lower side filling is enabled, when it is within the valid feature data range in the x-axis direction and the y-coordinate value of the physical address is greater than or equal to the data write position parameter, the corresponding filling data is generated.
12. The image arrayer according to claim 6, characterized in that: The read data control circuit at least includes a matrix data generation circuit; the read control parameters at least include the number of valid data and the number of elements included in a single read and write operation; Whenever the target original feature data is read, a corresponding number of valid original feature data is selected from the target original feature data according to the number of valid data, and stored in a shift register; the storage space of the shift register is greater than or equal to a preset overflow value, and the preset overflow value is determined according to the number of elements; Whenever the amount of valid original feature data of the shift register at the current moment is equal to the number of elements, the valid original feature data of the shift register at the current moment is sent as the feature matrix data at the current moment, and a read data valid signal is outputted at the same time.
13. The image arrayer according to claim 12, characterized in that: The shift register is connected to a read counter, and the read counter records the number of valid original feature data of the shift register at the current moment; When a read start signal is received, the read counter is reset; when a read data valid signal is valid and the amount of valid original feature data of the shift register at the current moment is not equal to the number of elements, the third count value of the read counter is adjusted to be the sum of the current third count value and the number of valid data; When the read data valid signal is valid, the amount of valid original feature data of the shift register at the current moment is equal to the number of elements, and a round of indexing of feature elements in the convolution kernel window has not been completed, the third count value of the read counter is adjusted to be the numerical difference between the number of valid data and the number of elements; When the read data valid signal is valid, the amount of valid original feature data of the shift register at the current moment is equal to the number of elements, and a round of indexing of feature elements in the convolution kernel window is completed, the third count value of the read counter is adjusted based on the current third count value, and the increased value is the difference between the number of valid data and the number of remaining elements; The remaining number of elements refers to the number of valid original feature data remaining when the original feature data of a convolution window is pieced together in units of the number of elements and when the number of elements cannot be pieced together for the last time.
14. An electronic device, characterized in that: comprising an off-chip memory, a processor and an image columnizer as claimed in any one of claims 1 to 13; Wherein, the off-chip memory stores original feature data; The image columnizer reads the original feature data from the off-chip memory and outputs corresponding feature matrix data; the processor performs matrix multiplication and addition calculations on the feature matrix data and corresponding feature matrix data to obtain a result matrix.
Citation Information
Patent Citations
Image processor and its method
JP1999167627A