Image serializer and electronic device
By integrating the image serializer in the neural network accelerator, the original feature data is directly read from off-chip memory and generated feature matrix data, the memory and transmission overhead problems caused by the duplication of feature matrix data during image serialization are solved, and more efficient artificial intelligence task execution is achieved.
Patent Information
- Application Number
- CN202510433888.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-04-08
AI Technical Summary
In the process of image serialization, the generated feature matrix data contains a large amount of duplicate data due to the special behavior of convolution during image serialization, which increases memory and data transmission overhead, and thus increases communication delay.
Design a dedicated hardware module, namely the image serializer, which is directly integrated into the neural network accelerator. Through the read and write control module and the data storage module, the original feature data is read from off-chip memory, and the feature matrix data is automatically generated to reduce the amount of feature matrix data stored and transmitted.
It effectively reduces memory and data transmission overhead, reduces communication delay, improves the execution efficiency of artificial intelligence tasks, and reduces the required memory resources and computing resources.
Smart Images

Figure CN119963402A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an image arrayer and electronic equipment. Background Art
[0002] im2col (image to column) slides the template by columns, converts the data in each window into column vectors, and then arranges them into a new matrix by columns, thereby converting the input data into matrix form data suitable for convolution operations.
[0003] In related technologies, due to the special behavior of convolution, the feature matrix data obtained after the input data is processed by im2col will contain a large amount of repeated data, and the amount of data is far greater than the original data. During the image columnization process, the feature matrix data needs to be cached and transmitted, which not only requires a large amount of memory resources, but also increases the data transmission overhead and computing resource overhead, resulting in communication delays. Summary of the invention
[0004] The present invention provides an image serializer and electronic device, which effectively reduces memory and data transmission overhead, effectively reduces communication delay, improves the execution efficiency of artificial intelligence tasks, and thus effectively reduces the memory resources and computing resources required for artificial intelligence tasks.
[0005] In order to solve the above technical problems, the present invention provides the following technical solutions: On the one hand, the present invention provides an image serializer, including an input end, a read-write control module, a data storage module and an output end; the input end receives read-write control parameters and original feature data; the read-write control module determines a data writing method and a data reading method based on the read-write control parameters, and writes the original feature data to the data storage module according to the data writing method; under the control of the data reading method, reads corresponding data from the data storage module according to the convolution operation behavior mode to generate feature matrix data; the output end outputs the feature matrix data.
[0006] On the other hand, the present invention provides an electronic device, including an off-chip memory, a processor and an image columnizer as described in an embodiment of the present invention; wherein the off-chip memory stores original feature data; the image columnizer reads the original feature data from the off-chip memory and outputs corresponding feature matrix data; the processor performs matrix multiplication and addition calculations on the feature matrix data and the corresponding feature matrix data to obtain a result matrix.
[0007] The advantage of the technical solution provided by the present invention is that the im2col operation is designed as a dedicated hardware module, namely, an image columnizer, which can be directly integrated into the neural network accelerator without the help of an external central processing unit. The communication and interactive operations between the neural network accelerator and the central processing unit are reduced, which effectively reduces the communication delay and is conducive to improving the execution efficiency of artificial intelligence tasks. In the process of executing the task, the image columnizer reads the original feature data from the off-chip memory and automatically generates feature matrix data. It only needs to directly store the original feature data with a small amount of data, which reduces the storage overhead of directly storing the feature matrix data in the cache, reduces the occupation of memory resources, and effectively reduces the memory overhead; in the process of converting the original feature data into feature matrix data, it only involves the transmission of the original feature data with a small amount of data, and there is no need to transmit the feature matrix data, which effectively reduces the data transmission overhead from the off-chip memory to the on-chip memory, thereby effectively reducing the memory resources and computing resources required for artificial intelligence tasks. In addition, the electronic device has corresponding advantages. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] In order to more clearly illustrate the technical solutions of the present invention or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0009] Figure 1 The following is a schematic diagram of the image columnization process in an exemplary scenario; Figure 2 A schematic diagram of convolution with a dilation rate of 1 in an exemplary scenario; Figure 3 A schematic diagram of convolution with a dilation rate of 2 in an exemplary scenario; Figure 4 A schematic diagram of convolution with a dilation rate of 3 in an exemplary scenario; Figure 5 A schematic diagram of a framework of a neural network accelerator for an exemplary application scenario provided by the present invention; Figure 6 A schematic diagram of a feature matrix and a weight matrix for an exemplary application scenario provided by the present invention; Figure 7 A structural framework diagram of an exemplary embodiment of an image serializer provided by the present invention; Figure 8 A schematic diagram of a storage method of original feature data provided by the present invention under an exemplary embodiment; Fig. 9A schematic diagram of a storage method of original feature data provided by the present invention under another exemplary embodiment; Fig.10 A structural framework diagram of another exemplary embodiment of the data storage module provided by the present invention; Fig.11 A structural framework diagram of another exemplary embodiment of the image serializer provided by the present invention; Fig.12 A structural framework diagram of an exemplary implementation of a write control module provided by the present invention; Fig.13 A schematic diagram of a write control flow provided by the present invention; Fig.14 A structural framework diagram of an exemplary implementation of a read control module provided by the present invention; Fig.15 A schematic diagram of a data filling method provided by the present invention; Fig.16 A schematic diagram of the principle of the data coordinate generation circuit provided by the present invention; Fig.17 A schematic diagram of the boundary range provided by the present invention; Fig.18 A schematic diagram of the state transition conditions of each counter provided by the present invention; Fig.19 A schematic diagram of the data piecing together process provided by the present invention; Fig. 20 A schematic diagram of the state transition conditions of the read counter provided by the present invention; Fig.21 This is a structural diagram of an exemplary embodiment of an electronic device provided by the present invention. DETAILED DESCRIPTION
[0010] In order to enable those skilled in the art to better understand the technical solution of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. Among them, the terms "first", "second", "third", "fourth", etc. in the specification and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "including" and "having" and any variations of the two are intended to cover non-exclusive inclusions. The term "exemplary" means "used as an example, embodiment or illustrative". Any embodiment described here as "exemplary" is not necessarily interpreted as being superior or better than other embodiments.
[0011] The convolution operation is to slide each convolution kernel on the feature data with a certain step size and expansion rate, and then multiply the feature data and the convolution kernel data accordingly and then accumulate them. It is a common data processing method in the field of machine learning and deep learning. Figure 1For example, the input feature data is three-dimensional data, that is, it includes three dimensions: H (high), W (width), and C (channel). Figure 1 Use H I , W I and C I Filters (convolution filter) includes Co convolution kernels with three dimensions H, W, C. The number of channels of each convolution kernel is the same as the number of channels of the feature data, which can be represented by H F , W F and C I For example, when the step size is 1 and the dilation rate is 1, the data corresponding to the first slide of the convolution kernel on the feature data is the first row of the feature matrix. This row is multiplied and accumulated with the first column element of the weight matrix (i.e., the first convolution kernel) to obtain the first element in the upper left corner of the result matrix. F =W F =2, H I =W I =3, Co=3 as an example, the convolution kernel can slide 4 times on the feature data. Correspondingly, the number of rows in the feature matrix is 4, and the number of elements in each row is H F *W F *C I , the number of columns of the weight matrix is 3, and the number of elements in each column is H F *W F *C I The step length in the horizontal or vertical direction is usually the same. When the step length is not 1, the distance moved by each slide is no longer 1. When the expansion rate is not 1, the feature data corresponding to each element in the convolution kernel will also change accordingly, with H F =W F =3 convolution kernel in H I =W I =5 feature data sliding up as an example, Figure 2 , Figure 3 and Figure 4 The dilation rates of the convolution are 1, 2, and 3 respectively. The black dots represent the elements of the convolution kernel. When the dilation rate is 2, the feature data elements corresponding to two adjacent elements of the convolution kernel are separated by 1 feature element. When the dilation rate is 3, the feature data elements corresponding to two adjacent elements of the convolution kernel are separated by 2 feature elements.
[0012] In order to speed up data processing efficiency, you can use im2col to convert the input data, such as Figure 1 The feature data is converted into a matrix form suitable for convolution operations, such as Figure 1If the feature data input is image data, the image data can be converted into matrix features through im2col. When executing various artificial intelligence tasks involving convolution operations, the original convolution operation can be converted into a matrix multiplication operation through im2col, and the matrix multiplication units of various neural network accelerators can be used for accelerated calculation, such as GPU (Graphics Processing Unit), DPU (Data Processing Unit), and FPGA (Field-Programmable Gate Array). Since im2col can convert the process of discontinuously reading input feature data according to the sliding window of the convolution kernel, that is, the data read is discontinuous in the memory, into the process of spatially continuous reading of the feature matrix, it can also accelerate the calculation of the convolution operation, thereby improving the data processing efficiency and the execution efficiency of artificial intelligence tasks.
[0013] Related technology The process of performing the above-mentioned convolution calculation operation based on im2col through a neural network accelerator includes: First, a software program is written and executed on a general-purpose processor such as a CPU (Central Processing Unit), and the im2col operation is performed on the input feature data to generate a feature matrix. The convolution kernel data is rearranged by flattening to obtain a weight matrix. Then, the feature matrix and the weight matrix are stored in a large-capacity slow off-chip memory, such as DDR (Double Data Rate) and HBM (High Bandwidth Memory). Then, any neural network accelerator is used to read the feature matrix and the weight matrix from the off-chip memory to a high-speed on-chip memory with limited capacity, such as SRAM (Static Random Access Memory). Finally, the neural network accelerator sends the data in the on-chip memory to the matrix multiplication unit to perform matrix multiplication calculations. The matrix multiplication unit can be, for example, a systolic array or a multiply-add tree to complete the convolution operation.
[0014] For the above process, the weight matrix is a rearrangement of the convolution kernel data, and there is no duplicate data. However, due to the special behavior of convolution, the feature matrix obtained by the im2col operation will contain a large amount of duplicate data, that is, the number of elements contained in the feature matrix far exceeds the size of the original feature data. For example, the original feature data includes 18 elements, and the generated feature matrix contains 32 elements. Although general-purpose processors can flexibly implement any arrangement of input feature data or convolution kernel data, since the feature matrix contains a large amount of duplicate data, this will lead to increased storage overhead and data transmission overhead. For slow, large-capacity off-chip memory, the memory resources required for the feature matrix will increase. It is even more unacceptable to store the feature matrix in the expensive but limited-capacity on-chip memory of the neural network accelerator, which seriously occupies the high-speed on-chip memory resources of the accelerator. After completing the im2col operation on a general-purpose processor such as the CPU, the feature matrix with increased data volume needs to be transferred from the off-chip memory to the high-speed on-chip memory of the neural network accelerator. The data volume transfer will increase the data transmission overhead and reduce the overall computing efficiency. In addition, neural network models usually contain multiple convolution operations, and neural network accelerators frequently use external CPUs and other general-purpose processors to complete im2col operations, which will introduce more communication delays, causing the matrix multiplication and addition computing units in the neural network accelerator to be idle and the computing delay to increase.
[0015] In view of this, the present invention designs a dedicated hardware module for im2col operation, namely, an image columnizer, which reads the original feature data from the off-chip memory, automatically generates feature matrix data through the internal circuit structure, and then sends the feature matrix data to the processor for multiplication and addition calculation, effectively reducing memory and data transmission overhead, effectively reducing communication delay, improving the execution efficiency of artificial intelligence tasks, and thus effectively reducing the memory resources and computing resources required for artificial intelligence tasks. Based on the above technical solution of the present invention, combined with Figure 5 Some possible application scenarios involved in the technical solution of the present invention are introduced by way of example, which may include the following contents: The image columnizer 502 is integrated into the neural network accelerator 50. For example, the neural network accelerator can be integrated into a GPU or an FPGA. Then, the neural network accelerator is deployed on a server 5. For example, the FPGA can be inserted into the server through a PCIe (peripheral component interconnect express) card slot. The neural network accelerator includes a first off-chip memory 501 and a matrix multiplication and addition calculation unit 503. The first off-chip memory 501 stores the original feature data. The data storage module of the image columnizer 502 uses a storage array, receives read-write control parameters and original feature data through an input end, and uses a read-write control module to determine a data writing method and a data reading method based on the read-write control parameters, and writes the original feature data into the storage array according to the data writing method. Under the control of the data reading method, the corresponding data is read from the storage array according to the convolution operation behavior to generate the following Figure 6 The characteristic matrix data shown in FIG. 5 is input to the matrix multiplication and addition calculation unit 503 through the output terminal. The matrix multiplication and addition calculation unit 503 performs Figure 6 The feature matrix data and the corresponding weight features are multiplied and added to obtain the final result matrix. It can be seen that the storage array of the image columnizer 502 only needs to cache the original feature data, automatically read the original feature data stored in the storage array according to the convolution behavior mode through the internal structure, and then assemble and combine it into feature matrix data before sending, avoiding the direct storage of the huge data after the im2col expansion, effectively reducing the memory and data transmission overhead, and the whole process also requires the neural network accelerator to interact and communicate with the server's CPU, effectively reducing communication delays, realizing accelerated training or accelerated reasoning of the neural network model, and improving the execution efficiency of artificial intelligence tasks based on the neural network model.
[0016] It should be noted that the above application scenarios are only shown to facilitate understanding of the ideas and principles of the present invention, and the embodiments of the present invention are not limited in this respect. On the contrary, the embodiments of the present invention can be applied to any applicable scenario. After introducing the technical solution of the present invention, various non-limiting embodiments of the present invention are described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0017] First see Figure 7 , Figure 7 This is a structural framework diagram of an exemplary implementation of the image serializer provided in this embodiment. This embodiment may include the following contents: The image serializer 502 may include an input terminal 703, a read / write control module 701, a data storage module 702, and an output terminal 704. The input terminal 703 may be, for example, a data interface or a user port, through which the read / write control parameters defined by the user to generate the corresponding hardware circuit may be received, and the user-defined read / write control parameters may ensure that the hardware circuit finally generated is a hardware circuit that meets the user's requirements. Since the image serializer 502 is to implement the convolution operation, the read / write control parameters must at least include the convolution operation parameters, and the convolution operation parameters at least include the expansion rate, the convolution kernel size, and the sliding step size. When programming, the expansion rate may be represented as params_conv_dil, the convolution kernel size may be represented as params_conv_ksize, the size of the convolution kernel, i.e., the size of Hf and Wf, these two values are generally the same and may be represented by one parameter; the sliding step size may be represented as params_conv_stride, and the sliding step sizes in the horizontal and vertical directions are generally consistent and may also be represented by one parameter. As for other parameter types of the read-write control parameters, they can be determined according to the read-write control mode to be implemented by the read-write control module 701 and the physical parameters of the storage components of the data storage module 702, and the present invention does not limit this. In addition, the original feature data of the present invention refers to the original data that needs to be image-columnized, and the original feature data can be pre-stored in the specified storage hardware, which is outside the image columnizer 502. For the convenience of description, it can be defined as an off-chip memory. When the image columnizer 502 is required to perform an image columnization operation on the original feature data, the image columnizer 502 can obtain the original feature data stored in the off-chip memory through the input terminal 703 and cache it in the data storage module 702. In other words, the data storage module 702 stores the original feature data input from the off-chip memory according to certain rules under the control of the read-write control module 701. Its essence is a storage hardware for caching the received original feature data so that the read-write control module 701 can read and process it. It is a storage hardware inside the image columnizer 502.
[0018] In this embodiment, the read-write control module 701 determines the data writing mode and the data reading mode based on the read-write control parameters. The data writing mode refers to the rules according to which the original feature data is stored in the data storage module 702, that is, the original feature data is written into the data storage module 702 according to the data writing mode. The data reading mode is based on the convolution behavior mode or the behavior mode of the image columnization operation to control how to read data from the data storage module 702 and generate corresponding feature matrix data, that is, under the control of the data reading mode, the corresponding data is automatically read from the data storage module 702 according to the convolution operation behavior mode, and the feature matrix data is automatically generated. The output terminal 704 sends the feature matrix data output by the read-write control module 701 to the outside, for example, to the PE (Processing Element) array or the multiplication-addition tree for multiplication-addition calculation.
[0019] The image serializer 502 can be implemented using RTL (register transfer language) languages, such as VHDL (Very High Speed Integrated Circuit Hardware Description Language) and Verilog (hardware description language), that is, the RTL language can be used to write hardware codes that implement the above functions, and these hardware codes can be processed using tools to generate corresponding physical circuit hardware. As for how to write hardware codes and what tools to use to generate circuits, as long as the image serializer 502 can implement the functions to be implemented by the functional modules of the image serializer 502 recorded in the present invention, those skilled in the art can handle them flexibly, which will not affect the implementation of the present invention.
[0020] In the technical solution provided in this embodiment, the im2col operation is designed as a dedicated hardware module, namely, the image columnizer 502, which can be directly integrated into the neural network accelerator without the need for an external central processing unit. The communication and interactive operations between the neural network accelerator and the central processing unit are reduced, which effectively reduces the communication delay and is conducive to improving the efficiency of artificial intelligence task execution. In the process of executing the task, the image columnizer 502 reads the original feature data from the off-chip memory and automatically generates feature matrix data. It only needs to directly cache the original feature data with a small amount of data, which reduces the storage overhead of directly storing the feature matrix data in the cache, reduces the occupation of memory resources, and effectively reduces the memory overhead; in the process of converting the original feature data into the feature matrix data, it only involves the transmission of the original feature data with a small amount of data, and does not need to transmit the feature matrix data, which effectively reduces the data transmission overhead from the off-chip memory to the on-chip memory, thereby effectively reducing the memory resources and computing resources required for the artificial intelligence task.
[0021] In order to further improve the processing efficiency of the image columnizer 502 on the original feature data, based on the above embodiment, the present invention further optimizes the storage method of the original feature data stored in the off-chip memory, which may include the following contents: It can be understood that the reading and writing of the original feature data by the image columnizer 502 of the present invention are both implemented under the control of the read-write control parameters. The storage method of the original feature data determines the reading and writing of the original feature data by the image columnizer 502. Therefore, when making certain agreements and changes to the storage format of the original feature data in the off-chip memory, it is also necessary to be based on the read-write control parameters. The read-write control parameters of this embodiment at least include the number of elements included in a single read-write operation, and the dimension of the original feature data includes the number of channels. Taking the data storage module 702 including multiple memories as an example, combined with the image columnization behavior mode of processing data by row, the corresponding memories of the data storage module 702 are essentially a line cache, and accordingly, the number of elements included in a single read-write operation can be the number of elements included in a read operation or write operation of a LineBuffer (line cache), that is, the number of elements included in a single read-write operation is the number of elements included in a single read-write operation of the line cache, which can be defined as NUM_CH, and NUM_CH can use a default value of 32, for example. The original feature data is three-dimensional data, and the three dimensions are length Hi, width Wi and number of channels Ci, such as Figure 8 As shown, this embodiment uses (Hi, Wi) as coordinates, and conditionally regards the Ci dimension as a whole, determines the data segmentation parameters according to the number of channels and the number of elements, and cuts the original feature data into multiple data blocks with the number of channels as a whole according to the data segmentation parameters, and then stores them in the target storage location. The target storage location can be an off-chip memory, or it can be refined to the storage address of the off-chip memory. Exemplarily, after the data blocks are divided, they can be stored in order from left to right and from top to bottom. Of course, they can also be stored in other orders, which does not affect the implementation of the present invention.
[0022] In this embodiment, in order to facilitate reading and writing of original feature data, the number of elements included in a single read / write operation needs to match the number of elements included in the original feature data stored at one address. Based on this, the present invention stores the original feature data in two cases according to the relationship between the number of channels of the original feature data and the number of elements included in a single read / write operation: The first storage situation: the value of the number of channels of the original feature data is greater than the number of elements contained in a single read and write operation. According to the number of elements, the data with the same width dimension and height dimension in the original feature data is cut into multiple data blocks, and each data block is stored in each address space of the target storage location. In this embodiment, ideally, the number of elements contained in each data block is equal to NUM_CH, and the number of elements is the data corresponding to a coordinate (Hi, Wi), which is defined as the first type of data block. It is inevitable that the number of elements contained in the data block is less than NUM_CH, which is defined as the second data block, that is, the data length value of the first type of data block is the same as the value of the number of elements, and the data length value of the second type of data block is less than the value of the number of elements. After the data with the same width and height dimensions in the original feature data are cut into first-class data blocks and second-class data blocks according to NUM_CH, each first-class data block is stored in each first-class address space of the target storage location in sequence in the manner of storing one first-class data block in one address space; after each first-class data block is filled, each second-class data block is stored in each second-class address space of the target storage location in sequence in the manner of filling each second-class data block into one address space until there is no remaining storage space; wherein the first and last addresses of each first-class address space are connected, and the last address of the last first-class address space is connected to the first address of the first second-class address space. Figure 8 As shown, the dimension of the original feature data is Hi=Wi=224, and the original feature data uses (Hi, Wi) as the coordinates, and the Ci dimension is conditionally regarded as a whole, and is stored in order from left to right and from top to bottom. The first type of address space is address 0, address 1, etc., and the second type of address space is the address for storing the remaining data. When Ci>NUM_CH, for each position index, the data of the first NUM_CH channels are saved first. After the data of the first NUM_CH channels of all position indexes are stored, the data of the subsequent NUM_CH channels of each position index are stored, and so on, until all the original feature data are stored.
[0023] The second storage situation: The number of channels of the original feature data is less than or equal to the number of elements contained in a single read and write operation. The length value of the padding data is determined according to the difference between the number of channels and the number of elements, and the padding data type is determined. For example, padding can be performed by using Padding 0, and the data with the same width and height dimensions in the original feature data and the padding data are stored in the address spaces of the target storage location. Fig. 9As shown, the dimension of the original feature data is Hi=Wi=224, and the original feature data is taken as the coordinate (Hi, Wi), and the Ci dimension is conditionally regarded as a whole, and is stored in order from left to right and from top to bottom. When Ci≤NUM_CH, Padding0 is performed, and the data with the same width dimension and height dimension and the padding data occupy one address space together.
[0024] Based on the storage method of the original feature data recorded in the above embodiment, when reading the original feature data, what is actually read is the first NUM_CH data at the coordinate (Hi, Wi). For example, when Ci=60, when reading data at address 0 from the off-chip memory, what is actually read is the first NUM_CH=32 data at the coordinate Hi=1, Wi=1, when reading address 1, what is read is the first NUM_CH data at the coordinate Hi=1, Wi=2; when reading address 224*224, what is read is the last 60-NUM_CH=28 data at the coordinate Hi=1, Wi=1, and so on.
[0025] It can be seen from the above that this embodiment divides the original feature data into channels according to the granularity of NUM_CH, and then stores the divided data blocks in sequence. The data storage method matches the data reading method, which is beneficial to improving the efficiency of the image columnization operation of the image columnizer 502 on the original feature data.
[0026] In order to further improve the efficiency of the image columnization operation of the image columnizer 502 on the original feature data, an off-chip memory can be set to actively send data to the image columnizer 502. Accordingly, the image columnizer 502 includes, for example, a data acquisition port set at the input end 703, or a port set to a module that implements write control in the read-write control module 701, and the off-chip memory sends the original feature data to the data acquisition port. Similarly, data transmission must match data reading. This embodiment can also predefine two parameters: single data transmission parameter, data validity parameter, and element width of single read and write operation. The element width of single read and write operation is based on the number of elements included in a single read and write operation, and further limits the bit (bit width) of a single element. Taking the data storage module 702 including multiple memories as an example, combined with the image columnization behavior mode, data is processed by row, and each memory of the corresponding data storage module 702 is essentially a row buffer. Correspondingly, the element width of a single read and write operation is the bit width of a single element in a LineBuffer read / write operation, which can be defined as LB_ELEWIDTH. LB_ELEWIDTH can use a default value of 16, for example. The single data transmission parameter refers to the original feature data from the off-chip memory, and how many elements with how much width are sent each time. When actually writing a program, for example, this parameter can be represented by lb_data_in. The data validity parameter is used to indicate when the data sent according to the single data transmission parameter is valid. When actually writing a program, for example, this parameter can be represented by lb_data_in_valid. After NUM_CH and LB_ELEWIDTH are determined, combined with the image columnization behavior, lb_data_in can be defined as data from an off-chip memory, and NUM_CH elements with a width of LB_ELEWIDTH are sent each time. Correspondingly, lb_data_in_valid indicates when the data of lb_data_in is valid, for example, the value of lb_data_in_valid is 1, indicating validity. Based on the above embodiment, it can be known that when the off-chip memory stores the original feature data, when Ci is less than NUM_CH or Ci cannot be divided by NUM_CH, there will be padding data, so when reading the original feature data from the off-chip memory, the padding data will be read, that is, the data read is not the real original feature data. Based on this, the read-write control parameter at least includes the number of valid data, which is used to indicate the number of valid data in each read data when reading the original feature data from the off-chip memory. Taking NUM_CH data as an example, the number of valid data indicates how many of the NUM_CH data read each time are valid data. When actually writing a program, this parameter can be indicated by params_conv_ci, for example.Based on this, the data acquisition port sends the original feature data in the target storage location to the input terminal 703 according to the single data transmission parameter, and the original feature data block with the first quantity and the width value of the data element is the element bit width of the single read and write operation. The original feature data block is the original feature data sent in the current round of data transmission, which is a part of the original feature data stored in the off-chip memory. When the original feature data is determined to be valid based on the data validity parameter, the data validity parameter is sent to the input terminal 703; the input terminal 703 determines the number of valid data contained in the original feature data based on the number of valid data.
[0027] As can be seen from the above, the present embodiment sets a port for obtaining the original feature data stored according to the agreed specifications from the off-chip memory, and the off-chip memory sends data through the port in an agreed manner, i.e., in accordance with the provisions of the read-write control parameters, so as to facilitate the image serializer 502 to process the original feature data and improve the efficiency of the image serializer 502 in the image columnization operation on the original feature data.
[0028] In the above embodiment, there is no limitation on the structure of the data storage module 702. The data storage module 702 of this embodiment may be composed of a group of memories, that is, the data storage module 702 is a storage array. The memory may be an SRAM (Static Random Access Memory). The data storage module 702 caches the original feature data from the off-chip memory through an array composed of multiple independent high-speed on-chip memories (SRAM). After the read-write control parameters define the convolution operation parameters and the number of elements NUM_CH contained in a single read-write operation and the element width LB_ELEWIDTH of a single read-write operation, it is necessary to further define the storage parameters. The storage parameters include at least the total number of memories, the maximum address supported by a single memory for reading and writing, and the data volume of a single read-write operation. In the actual writing process, LB_CH can be used to represent the total number of memories, LB_DEPTH can represent the maximum address supported by a single memory for reading and writing, and LB_WIDTH can be used to represent the data volume of a single read-write operation. LB_CH indicates how many memories are in the data storage module 702, for example, the default value of 16 can be used; this parameter determines that the maximum of the two dimensions (Hf, Wf) (length and width of the convolution kernel) of the supported convolution is 16. In other words, LB_CH memories indicate that the maximum convolution kernel size supported for the convolution operation is LB_CH. LB_DEPTH indicates the depth of each memory in the data storage module 702, that is, the maximum address + 1 supported by each memory for reading and writing, for example, the default value of 256 can be used. LB_WIDTH is the width of a single memory, that is, the amount of data operated by a single memory read / write at a time, and its value is LB_ELEWIDTH*NUM_CH, that is, the default value can be 32*16=512bit. Considering that the convolution operation or the graphic columnization process is performed in units of lines, it is used to cache the original feature data, and each memory is essentially a row cache, and LineBuffer can be used to represent the memory. In other words, the number of memories included in the storage array is the same as the total number of memories LB_CH in the storage parameters, and the total number of memories is determined according to the maximum convolution kernel size of the convolution operation; each memory is independent of each other, and the read and write data width of each memory is determined according to the number of elements and the element bit width, and the total amount of stored data is determined according to the maximum address supported by a single memory in the storage parameters, and the corresponding read and write data width. For example, please refer to Fig.10, the data storage module 702 includes LB_CH independent SRAMs, the read and write data width of each SRAM is NUM_CH*LB_ELEWIDTH, and each SRAM can store at most LB_DEPTH data with a width of NUM_CH*LB_ELEWIDTH, that is, the maximum address range is 0~LB_DEPTH-1. Since there are LB_CH SRAMs, the lb_wens (write enable signal), lb_waddrs (write address signal), lb_wdatas (write data signal), lb_rens (read enable signal), lb_raddrs (read address signal), and lb_rdatas (read data signal) of the data storage module 702 have LB_CH copies, corresponding to different SRAMs.
[0029] As can be seen from the above, the present invention uses a high-speed on-chip memory to form a storage array to store original feature data, and each storage parameter of the storage array matches the read-write control parameter, which can improve the read-write efficiency of the original feature data.
[0030] In the above embodiment, the structure of the read-write control module 701 is not limited in any way. The present invention also provides an exemplary structure of the read-write control module 701. In this embodiment, Fig.11 As shown, the read-write control module 701 is composed of a read control module and a write control module. The read control module interacts with the write control module and the data storage module 702, and the write control module interacts with the data storage module 702. The read control module and the write control module are introduced below.
[0031] In this embodiment, the write control module is connected to the input terminal 703, and the parameters input to the write control module are defined as write control parameters, that is, the read-write control parameters include write control parameters. The write control module includes at least a counter, a write parameter input port, and a write signal output port. The write parameter input port receives the write control parameters. Among them, for the convenience of description, the counter inside the write control module is defined as a write counter. Based on the write control parameters, according to the sending status of the original feature data and the storage parameters of the data storage module 702, the storage location information of the sub-original feature data currently input in the data storage module 702 is determined, and a write enable signal, a write address signal, and a write data signal are generated according to the storage location information and the sub-original feature data; through the write signal output port, the write enable signal, the write address signal, and the write data signal are sent to the data storage module 702.
[0032] In this embodiment, it can be understood that the original feature data is sent to the image serializer 502 in batches, or in other words, the image serializer 502 reads the original feature data from the off-chip memory in batches or in multiple rounds. For the convenience of description, the part of the original feature data processed at the current moment of the current round is defined as sub-original feature data, and the sub-original feature data is part of the original feature data received at the current moment. Among them, the write control parameters of this embodiment are used to control the amount of single-write data and the total amount of data written to each storage location of the data storage module 702, including but not limited to write start parameters, single data transmission parameters lb_data_in, data valid parameters lb_data_in_valid and write position switching parameters. Among them, the write start parameter indicates that the LineBuffer is about to start writing, which is used as a reset signal and can be represented by params_write_start. The write position switching parameter indicates the relationship between the number of data writes and the storage position, which can be represented by params_lb_wswitch_cnt. Taking the data storage module 702 as a storage array as an example, the write position switching parameter refers to the number of times the data storage module 702 switches the next LineBuffer when writing the original feature data to LB_CH LineBuffers. The data validity parameter indicates whether the sub-original feature data sent according to the single data sending parameter is valid. After the write control parameter determines the write start parameter, the single data sending parameter, the data validity parameter and the write position switching parameter, the write control module receives the write start parameter and resets the timer; whenever the data validity parameter is received, the storage position information of the sub-original feature data currently input in the data storage module 702 is determined according to the write position switching parameter and the count value counted by the timer.
[0033] For example, Fig.12As shown, in this embodiment, the data storage module 702 includes multiple memories, and the write counter may include a first write counter and a second write counter. The first counter can be represented by wline_cnt, and the second counter can be represented by lb_write_index. The single data transmission parameter is that the amount of data of the sub-original feature data sent each time is the same as the read and write data width of a single memory. The write position switching parameter represents the number of valid data written by a single memory, that is, params_lb_wswitch_cnt represents how many times the data is written to switch the next LineBuffer. The Hi dimension of the original feature data in the off-chip memory corresponds to the total number of memories included in the data storage module 702, and the Wi dimension corresponds to the depth of a single SRAM memory, that is, LB_CH LineBuffers can only store LB_CH rows of original feature data at most, and LB_DEPTH data per row. In this way, lb_write_index can reflect the index of the memory SRAM currently being written, and wline_cnt can represent the address of the memory SRAM currently being written. When a data validity parameter is received through the write parameter input port, a write enable signal is generated according to the second count value, indicating that the write enable of the target memory is valid and the write enable of other memories of the data storage module 702 is invalid, a write address signal is generated according to the first count value, and a write data signal is generated according to the sub-original feature data.
[0034] Based on the above write control module, the workflow of wline_cnt and lb_write_index is as follows: Fig.13As shown, before each round of loading the original feature data from the off-chip memory, the value of the write start signal can be set to 1, that is, params_write_start==1, indicating that the write start signal is valid and lasts for one clock cycle. At this time, the first counter and the second counter are reset to 0. Whenever the data valid parameter is received through the write parameter input port, if the first count value of the first write counter is not the write position switching parameter-1, the value of the current first count value is increased by the first preset value; if the first count value is the write position switching parameter-1, the first write counter is reset, and the value of the current second count value of the second write counter is increased by the second preset value, so that the first write counter represents the address of the target memory currently being written, and the second write counter represents the index of the target memory currently being written. Among them, the first preset value and the second preset value can be set according to the actual situation, for example, they can both be 1. That is, whenever there is original feature data, that is, whenever the data valid parameter is valid, that is, lb_data_in_valid==1, it is determined whether the first count value of wline_cnt is equal to params_lb_wswitch_cnt-1. If wline_cnt==params_lb_wswitch_cnt-1, the first count value of wline_cnt is set to 0 so that it can start accumulating from 0 next time, and the second count value of lb_write_index is increased by 1, that is, lb_write_index++; if wline_cnt is different from params_lb_wswitch_cnt-1, the first count value of wline_cnt is increased by 1, that is, wline_cnt++.
[0035] Take the storage array as multiple SRAMs as an example. When lb_write_index is 0 and lb_data_in_valid==1, the write enable of the first SRAM in the corresponding lb_wens is valid, and the write enable of the remaining SRAMs is invalid; the value of the first write address in lb_waddrs is assigned to the value of wline_cnt, and the value of the first write data in lb_wdatas is assigned to the value of lb_data_in; as the sub-original feature data is continuously acquired, wline_cnt gradually increases from 0 to params_lb_wswitch_cnt-1, and accordingly, data is written to the address 0 to params_lb_wswitch_cnt-1 of the first SRAM. Figure 8 and Fig. 9Taking the storage method shown as an example, under the control of the write parameters, the original feature data is read from the off-chip memory. For example, if params_lb_wswitch_cnt is set to 10, the data of the coordinates <1, 1> to <1, 10> will be read and saved in the first SRAM; then the data of the coordinates <2, 1> to <2, 10> will be read and saved in the second SRAM. At this time, lb_write_index becomes 1, wline_cnt becomes 0, and the subsequent params_lb_wswitch_cnt data will be written to the second SRAM, and so on, until the write operation is completed. When the write control operation is completed, lb_write_index indicates how many rows of original feature data are written in the storage array in this round of write operation, and params_lb_wswitch_cnt indicates how many valid data there are in each row.
[0036] From the above, it can be seen that this embodiment specifies how many rows of original feature data are written in each round and how much data is written in each row through write control parameters, and uses a counter to count how many rows of original feature data are currently written and how many valid data are in each row, thereby achieving efficient and accurate caching of the original feature data to the image columnizer 502.
[0037] In this embodiment, if Fig.11As shown, the read-write control module 701 also includes a read control module, and the parameters input to the read control module are defined as read control parameters, that is, the read-write control parameters include read control parameters. It should be noted here that the read-write control parameters include all control parameters that need to be input to the image serializer 502. Different embodiments focus on describing the content of the current embodiment. In order to avoid description, the control parameters used in the reading process are defined as read control parameters, and the control parameters used in the writing process are defined as write control parameters. The same control parameters will be used in the reading and writing processes, that is, the same control parameters will exist in the read control parameters and the write control parameters. This does not mean that the read-write control parameters will include repeated parameters, but after the read-write control parameters are input to the image serializer 502, different functional modules will use them according to their own needs. Among them, the read control module at least includes a read parameter input port, a read data port and a read data control circuit. The read data port is connected to the data storage module 702, and the read data control circuit is connected to the output terminal 704. The read control parameter can be received through the read parameter input port; the read data control circuit is a circuit that is converted into the hardware code of the final feature matrix data by reading the original feature data in the data storage module 702 according to the behavior mode of the convolution through the tool, and the function to be realized is: determine the target position information of the target original feature data covered by the convolution kernel during the sliding process according to the read control parameter, determine the physical position read information corresponding to the data storage module 702 based on the target position information, and read the target original feature data from the data storage module 702 through the read data port according to the physical position read information, and generate feature matrix data according to the storage mode of the original feature data and the target original feature data. Finally, the feature matrix data is output to the specified position, such as the systolic array, through the output terminal 704. Among them, the storage mode of the original feature data refers to the storage mode of storing the original feature data in the off-chip memory and / or the storage mode of the original feature data on the data storage module 702, and the real data is determined to be read according to the storage mode. The location of the real data is determined and the filling data is identified, and the real data is the original feature data. The target position information is the Hi coordinate and Wi coordinate of the original feature data covered by the convolution kernel during the sliding process, and the physical position reading information refers to which memory of the data storage module 702 and which address of the memory to read the data from.
[0038] In order to use the read control circuit to determine the target position information of the target original feature data in the convolution mode with different convolution kernels, different expansion rates, and different filling sizes, the read data control circuit of this embodiment at least includes a data coordinate generation circuit, such as Fig.14As shown, the data coordinate generation circuit is responsible for sequentially generating Hi coordinates and Wi coordinates of the original feature data covered by the convolution kernel during the sliding process for the original feature data cached in the data storage module 702. When determining the coordinate position, since the channel number dimension of the convolution kernel is always the same as the original feature data, the Hi coordinate can be ignored, and the target position information is finally output. In this embodiment, after the write control module completely writes the original feature data to the data storage module 702, the original feature data of Hi=lb_write_index and Wi=params_lb_wswitch_cnt is cached in the data storage module 702. Since the convolution process is the process of sliding multiplication of the convolution kernel on the original feature data, the convolution kernel cannot exceed the right boundary and the lower boundary of the valid feature data. In the convolution process, it may be necessary to selectively fill the upper, lower, left, and right sides of the original feature data, so the right boundary and the lower boundary of the valid feature data are variable. Based on this, the target position information of the present invention includes the right boundary position and the lower boundary position of the target original feature data, wherein the right boundary position is the x coordinate value and the lower boundary position is the y coordinate value. The read parameter input port receives the write position switching parameter and the data write position parameter. Among them, the data write position parameter indicates the data storage amount of the original feature data in the data storage module 702. Taking the storage array as the data storage module as an example, the data write position parameter refers to how many rows of original feature data are written in the storage array in this round of write operation. When the write control module sets the second counter, the data write position parameter can be represented by lb_write_index, that is, the write control module sends lb_write_index to the read control module, and the read control module uses it as the data write position parameter. The right boundary position of the target original feature data is determined according to the write position switching parameter, and the lower boundary position of the target original feature data is determined according to the data write position parameter.
[0039] Exemplarily, in combination with whether the original feature data is read in a filling mode, this embodiment also provides an exemplary determination process of the target position information, and the read parameter input port receives the write position switching parameter, the data write position parameter and the filling mode read parameter. Among them, the filling mode read parameter indicates that the data, the filling position and the filling length are read from the data storage module 702 in a filling mode. Since the filling can be performed up, down, left and right, the filling is the same in the up and down filling, the left and right filling is the same, and the up and down filling length and the left and right filling length may be different, so the filling length may include a first filling length and a second filling length, the first filling length is the filling length in the horizontal direction, and the second filling length is the filling length in the horizontal direction and the filling length in the vertical direction. When the first filling length and the second filling length are the same, the filling length at this time can be defined as a fixed filling length. In actual application, the filling mode read parameter may include params_penable, params_pleft_en, parmas_pright_en, parmas_pup_en, parmas_pdown_en, params_conv_phsize and params_conv_pvsize. Among them, params_penable indicates whether the Padding mode is enabled when reading data from the data storage module 702 or LineBuffer. params_pleft_en indicates whether it is necessary to insert Padding on the left side of the data when reading data from the data storage module 702; parmas_pright_en indicates whether it is necessary to insert Padding on the right side of the data when reading data from the data storage module 702; parmas_pup_en indicates whether it is necessary to insert Padding on the upper side of the data when reading data from the data storage module 702; parmas_pdown_en indicates whether it is necessary to insert Padding on the lower side of the data when reading data from the data storage module 702. params_conv_phsize indicates the size of the single-sided Padding when Padding in the horizontal direction (i.e. left or right); params_conv_pvsize indicates the size of the single-sided Padding when Padding in the vertical direction (i.e. top or bottom). Based on the above-mentioned padding method, the parameters are read to perform Padding (fill with 0) operation on the original feature data, such as Fig.15As shown in the figure, the two parameters params_conv_phsize and params_conv_pvsize indicate how many zeros are padded in the Hi and Wi dimensions of the actual data; params_penable indicates whether to perform padding, and params_pleft_en, parmas_pright_en, parmas_pup_en, and parmas_pdown_en are the sub-switches in the four directions. When the corresponding switch is valid, that is, when the values of params_pleft_en, parmas_pright_en, parmas_pup_en, and parmas_pdown_en are 1, the corresponding position will be padded with the corresponding number of zeros according to the values set by params_conv_phsize and params_conv_pvsize.
[0040] Based on the above read control parameters, when the filling position is left filling and right filling, the right boundary position of the target original feature data is determined according to the write position switching parameter and the double first filling length. When the filling position is left filling or right filling, the right boundary position of the target original feature data is determined according to the write position switching parameter and the first filling length. When the filling position is neither left filling nor right filling, the right boundary position of the target original feature data is determined according to the write position switching parameter; that is, when params_penable is invalid, the right boundary is params_lb_wswitch_cnt, and the lower boundary is lb_write_index. When params_penable is valid, and params_pleft_en and parmas_pright_en are both valid, the right boundary is params_lb_wswitch_cnt+2*params_conv_phsize; when only one of params_pleft_en and parmas_pright_en is valid, the right boundary is params_lb_wswitch_cnt+params_conv_phsize. When params_pleft_en and parmas_pright_en are both invalid, the right boundary is params_lb_wswitch_cnt. When the filling position is upper filling and lower filling, the lower boundary position of the target original feature data is determined according to the data write position parameter and the double second filling length. When the filling position is upper filling or lower filling, the lower boundary position of the target original feature data is determined according to the data write position parameter and the second filling length. When the filling position is neither upper filling nor lower filling, the lower boundary position of the target original feature data is determined according to the data write position parameter. That is, when parmas_pup_en and parmas_pdown_en are both valid, the lower boundary is lb_write_index+2*params_conv_pvsize. When only one of parmas_pup_en and parmas_pdown_en is valid, the lower boundary is lb_write_index+params_conv_pvsize. When both parmas_pup_en and parmas_pdown_en are invalid, the lower boundary is lb_write_index.
[0041] When the above embodiment determines the lower boundary and the right boundary according to different filling modes, the target position information, that is, the corresponding xy coordinates, can be obtained based on the currently received write position switching parameters, data write position parameters and filling mode reading parameters.
[0042] As can be seen from the above, the data coordinate generation circuit of this embodiment generates corresponding xy data coordinates within the lower boundary and the right boundary according to the received read control parameters. In order to facilitate those skilled in the art to understand the principle of the data coordinate generation circuit of the present invention, it may include the following contents:
[0043] The data coordinate generation circuit can set a first read counter, a second read counter, a third read counter, a fourth read counter, a fifth read counter and a sixth read counter, and determine the x coordinate of the original feature data according to the sum of the values of the third read counter and the fifth counter, and determine the y coordinate of the original feature data according to the sum of the values of the fourth read counter and the sixth counter. The counting interval of each time of the third read counter and the fourth counter is +params_conv_dil, and the counting interval of each time of other counters is +1.
[0044] In this embodiment, the first read counter is used to indicate the number of data indexes in the x direction inside the convolution kernel in the current sliding, which can be indicated by x_index_cnt1, the second read counter indicates the number of data indexes in the y direction inside the convolution kernel in the current sliding, which can be indicated by y_index_cnt1, the third read counter indicates the data coordinates in the x direction inside the convolution kernel, which can be indicated by x_index_cnt1_with_dil, the fourth read counter indicates the data coordinates in the y direction inside the convolution kernel, which can be indicated by y_index_cnt1_with_dil, the fifth read counter indicates the overall x coordinate position of the current sliding window in the feature data after Padding, which can be indicated by x_index_cnt2, the sixth read counter indicates the overall y coordinate position of the current sliding window in the feature data after Padding, which can be indicated by y_index_cnt2, and the two together indicate the coordinates of the upper left corner element of the convolution kernel in the feature data after Padding. A seventh counter can also be set, which is used to generate a working signal, which can be indicated by index_en, and when its value is 1, it indicates the start of work. When the read start signal is detected, it indicates that the LineBuffer is about to be read and is used as a reset signal. That is, when params_lb_read_start==1, index_en=1. When y_index_cnt2_finish, index_en changes from 1 to 0. The change process of these seven counters is as follows Fig.18 As shown, Fig.18The "!" in params_lb_read_start indicates the logical NOT operator. In the params_lb_read_start command, all counters are reset to 0 and index_en is set to 1. right_bound is the right boundary, right_bound_nw is the right boundary of the next sliding window, down_bound is the lower boundary, and down_bound_nw is the lower boundary of the next sliding window. The counter starts working when it detects that index_en is 1.
[0045] See also Fig.16 and Fig.17 For example, lb_write_index=6, params_lb_wswitch_cnt=6, the size of the convolution kernel params_conv_ksize is 3*3, the step size params_conv_stride and the expansion rate params_conv_dil are both 1, and through the write operation, 6 rows of original feature data are saved in the first 6 memories of the LB_CH memories of the storage array. Each memory includes 6 data and is set to pad the top, bottom, left and right of the original feature data. The size is 2, that is, the fixed padding length params_pvsize=params_phsize=2. In each slide, such as Fig.16As shown, the convolution kernel uses Z-shaped data indexes inside. The coordinate in the x direction (i.e. x_index_cnt1) needs to be counted params_conv_ksize times in each round, and a total of params_conv_ksize rounds are counted. At the end of each round of counting, an x_index_cnt1_finish (first counter end signal) signal is generated, i.e. x_index_cntl_finish=(index_en&&x_index_cntl==(params_conv_ksize-1)), and x_index_cnt1 returns to 0. Each time a round of counting is completed in the x-direction coordinate, that is, each time an x_index_cnt1_finish signal is generated, the coordinate in the y-direction (that is, y_index_cnt1) is counted once, and a total of params_conv_ksize times are counted. When the coordinate count in the Y-direction is params_conv_ksize times and x_index_cnt1_finish is generated, a y_index_cnt1_finish (second counter end signal) signal is generated, and y_index_cnt1 returns to 0, that is, y_index_cntl_finish=(x_index_cntl_finish&&y_index_cntl==(params_conv_ksize-1)). In this way, the counters x_index_cnt1 and y_index_cnt1 can be used to record the number of xy data indexes inside the convolution kernel during this sliding, and complete the data index change inside the convolution window. In each slide, when the expansion rate is not 1, the difference between the x-direction coordinates and y-direction coordinates of the adjacent internal data of the convolution kernel is params_conv_dil, so the counters x_index_cnt1 and y_index_cnt1 represent the number of data indexes in the x and y directions inside the convolution kernel in this slide. This embodiment uses counters x_index_cnt1_with_dil and y_index_cnt1_with_dil to record the data coordinates in the x and y directions inside the convolution kernel in this slide. For each slide and each slide: after determining the index changes in the x and y directions inside the convolution kernel, this embodiment uses counters x_index_cnt2 and y_index_cnt2 to record the overall x-coordinate position and y-coordinate position of the current sliding window in the feature data after Padding, which together constitute the coordinates of the upper left corner element of the convolution kernel in the feature data after Padding.The initial values of x_index_cnt2 and y_index_cnt2 are both 0. Whenever a y_index_cnt1_finish signal is generated, that is, each time a round of indexing of feature elements in the convolution kernel window is completed, and the next sliding window will not exceed the right boundary, x_index_cnt2 increases params_conv_stride based on the current value. When the y_index_cnt1_finish signal is generated, but the next sliding window will exceed the right boundary, x_index_cnt2 is reset to 0. When the x_index_cnt2_finish (fifth counter end signal) signal is generated, x_index_cnt2_finish=(y_index_cntl_finish&&right_bound_nw>right_bound), it means that the sliding window of the convolution kernel needs to move downward and return to the left boundary. Whenever the x_index_cnt2_finish signal is generated, it means that the sliding window has moved to the position allowed by the right boundary in the X direction, and the next sliding window will not exceed the lower boundary, and y_index_cnt2 increases params_conv_stride based on the current value. When the x_index_cnt2_finish signal is generated, and the next sliding window will exceed the lower boundary, the y_index_cnt2_finish (sixth counter end signal) signal is generated, y_index_cnt2_finish=(x_index_cnt2_finish&&down_bound_nw>down_bound), indicating that the im2col operation for the original feature data written to the data storage module 702 is completely completed. The final output reads the x coordinate of the original feature data as: x_index=x_index_cnt1_with_dil+x_index_cnt2, and the y coordinate of the original feature data as: y_index=y_index_cnt1_with_dil+y_index_cnt2.
[0046] As can be seen from the above, this embodiment uses multiple counters to count the sliding conditions of the convolution operation process, so as to generate corresponding xy data coordinates within the lower boundary and the right boundary according to the received read control parameters.
[0047] In order to use the read control circuit to determine the physical position read information of the target original feature data in the convolution mode with different filling sizes, so as to efficiently and accurately read the cached original feature data, the read data control circuit of this embodiment at least includes a read address generation circuit. When the target position is valid, the read address generation circuit determines the physical address according to the target position information of the target original feature data and the storage method of the original feature data; determines the physical position read information of the data storage module 702 according to the physical address and the target position information, and generates a data read signal according to the physical position read information; the data read signal is sent to the read data port. In this embodiment, the read address generation circuit can receive x_index, y_index, and index_valid in the data coordinate generation circuit, and generate a signal for reading the original feature data in the data storage module. Among them, index_valid is used to indicate that the target position is valid.
[0048] Exemplarily, in combination with whether the original feature data is read in a filling manner, this embodiment also provides an exemplary determination process of the physical address. When the data is read from the data storage module 702 in a filling manner, the filling position is left filling, then the x coordinate in the physical address is the difference between the x coordinate of the target position information and the first filling length; when the filling position is not left filling, then the x coordinate in the physical address is the x coordinate of the target position information; when the filling position is upper filling, then the y coordinate in the physical address is the difference between the y coordinate of the target position information and the second filling length; when the filling position is not upper filling, then the y coordinate in the physical address is the y coordinate of the target position information; when the original feature data is stored without adding a filling element, that is, when the data is read from the data storage module 702 in a filling manner, the x coordinate in the physical address is the x coordinate of the target position information, and the y coordinate in the physical address is the y coordinate of the target position information. In this embodiment, according to params_penable, params_pleft_en, and parmas_pup_en, the x and y coordinates of the physical address are calculated, which can be expressed as real_x_index and real_y_index. When params_penable=1 and params_pleft_en=1, real_x_index=x_index-params_conv_phsize, otherwise, real_x_index=x_index. For example, when the left padding length is 2, when x_index is actually 0 or 1, it does not refer to the feature data stored in the data storage module 702, but the 0 data of Padding; at this time, there is no need to read the data storage module 702. When params_penable=1 and params_pup_en=1, real_y_index=y_index-params_conv_pvsize, otherwise, real_y_index=y_index.
[0049] After the physical address is determined through the above embodiment, considering the existence of padding data, it is necessary to further determine whether the generated physical address is within the real and valid feature data range. If it is within the real and valid feature data range, it is necessary to read the data storage module 702. If it is not within the real and valid feature data range, it is not necessary to read the data storage module 702. For example, x_index_select and y_index_select can be used to represent it. x_index_select means that it is within the valid feature data range in the x-axis direction, and y_index_select means that it is within the valid feature data range in the y-axis direction. The data coordinate generation circuit will continue to generate x_index and y_index, and the read address generation circuit will continue to update real_x_index, real_y_index, x_index_select, and y_index_select according to x_index and y_index. The read address generation circuit reads the original feature data stored in the data storage module 702 according to the generated series of data coordinates and the behavior mode of convolution.
[0050] If the left padding is not enabled, when the x - coordinate value of the physical address is less than the write position switching parameter, it is within the range of valid feature data in the x - axis direction; when the left padding is enabled, when the x - coordinate of the target position information is greater than or equal to the first padding length and the x - coordinate value of the physical address is less than the write position switching parameter, it is within the range of valid feature data in the x - axis direction; if the upper padding is not enabled, when the y - coordinate value of the physical address is less than the data write position parameter, it is within the range of valid feature data in the y - axis direction; when the upper padding is enabled, when the y - coordinate of the target position information is greater than or equal to the second padding length and the x - coordinate value of the physical address is less than the data write position parameter, it is within the range of valid feature data in the y - axis direction. That is to say, in this embodiment, the lower - boundary detection method and the upper - boundary detection method can be determined. For the lower - boundary detection method in the x - direction, it is that the LeftPadding is not enabled or the LeftPadding is enabled AND x_index>=params_conv_phsize; for the lower - boundary detection method in the y - direction, it is that the UpPadding is not enabled or the UpPadding is enabled AND y_index>=params_conv_pvsize. For the upper - boundary detection method in the x - direction, it is real_x_index<params_lb_wswitch_cnt; for the upper - boundary detection method in the y - direction, it is real_y_index<lb_writing_index. For x_index_select, during the lower - boundary detection process, if there is LeftPadding, the LeftPadding part is skipped; during the upper - boundary detection process, if there is RightPadding, the data storage module does not need to be read. For y_index_select, during the lower - boundary detection process, if there is UpPadding, the UpPadding part is skipped; during the upper - boundary detection process, if there is DownPadding, the DownPadding part is skipped. That is to say, when it is satisfied that it is within the range of valid feature data in both the x - axis direction and the y - axis direction, the data read signal is generated according to the storage position of the target original feature data in the data storage module 702 and the corresponding read address as the physical position read information. In the actual application process, the following relational expression can represent the whole process: x_index_select = ((LeftPadding is not enabled) OR (LeftPadding is enabled AND x_index >= params_conv_phsize)) AND (real_x_index < params_lb_wswitch_cnt); y_index_select = ((UpPadding is not enabled) OR (UpPadding is enabled AND y_index >= params_conv_pvsize)) AND (real_y_index < lb_writing_index).
[0051] After the physical address is determined, the storage location of the target original feature data in the data storage module 702 can be determined according to the y coordinate value of the physical address; the read address of the target original feature data in the storage location can be determined according to the x coordinate value of the physical address; the storage location of the target original feature data in the data storage module 702 and the corresponding read address are used as the physical location read information. Taking the data storage module 702 as a storage array, and the storage array includes multiple SRAMs as an example, when both x_index_select and y_index_select are valid, the SRAM in the data storage module 702 needs to be read. Among the LB_CH SRAMs, which SRAM to read is determined according to real_y_index. During the application process, the lb_rens (read enable) of the corresponding SRAM can be set to 1 to achieve this. When the SRAM to be read is determined, the data at the real_x_index address is read. During the actual application process, the corresponding lb_raddrs (read address) can be set to real_x_index.
[0052] In order to enable the read control circuit to support convolution modes with different convolution kernels, different dilation rates, and different padding sizes, the read data control circuit in this embodiment at least includes a padding data generation circuit. The padding data generation circuit generates padding data according to the received target position information, physical address, padding mode parameter, data write position parameter, write position switching parameter, and whether it is within the valid feature data range signal, and outputs a valid padding signal and a padding data signal according to the padding data.
[0053] In this embodiment, the read address generation circuit sends the real_x_index, real_y_index, and x_index_select signals to the padding data generation circuit, and the data coordinate generation circuit sends the x_index and y_index to the padding data generation circuit. The padding data generation circuit generates padding data, such as params_penable, params_pleft_en, params_pright_en, parmas_pup_en, params_pdown_en, such as generating 0 data for Padding, according to the continuously updated x_index, y_index, real_x_index, real_y_index, x_index_select, and the set padding mode read parameters. The generation process of the padding data is as follows: If left padding is enabled, when the x coordinate value of the target position information is less than the fixed padding length, the corresponding padding data is generated. For example, when LeftPadding is enabled (i.e., both params_penable and params_pleft_en are 1), and when x_index is less than params_phsize, 0 data for the corresponding position of LeftPadding is generated. If right padding is enabled, when the x coordinate of the physical address is greater than or equal to the write position switching parameter, the corresponding padding data is generated; for example, when RightPadding is enabled (i.e., both params_penable and params_pright_en are 1), and when real_x_index >= params_lb_wswitch_cnt, 0 data for the position filled by RightPadding is generated. If up padding is enabled, when it is within the range of valid feature data in the x-axis direction and the y coordinate value of the physical address is less than the second padding length, the corresponding padding data is generated; for example, when UpPadding is enabled (i.e., both params_penable and params_pup_en are 1), and when x_index_select is valid and y_index < params_conv_pvsize, 0 data for the corresponding position of UpPadding is generated. If down padding is enabled, when it is within the range of valid feature data in the x-axis direction and the y coordinate value of the physical address is greater than or equal to the data write position parameter, the corresponding padding data is generated. For example, when DownPadding is enabled (i.e., both params_penable and params_pdown_en are 1), and when x_index_select is valid and real_y_index >= lb_writing_index, 0 data for the corresponding position of DownPadding is generated.
[0054] In order to realize that the read control circuit supports convolution modes with different convolution kernels, different expansion rates, and different padding sizes, the read data control circuit of this embodiment includes at least a matrix data generation circuit. The above-mentioned padding data generation circuit will output padding_valid and padding_data signals, padding_valid indicates a valid signal of the padding method currently adopted, and padding_data is the padding data generated according to different padding conditions. The padding data generation circuit will send the generated padding_valid and padding_data signals to the matrix data generation circuit, and the matrix data generation circuit also receives the real feature data from the data storage module 702 read by the read address generation circuit, collects these data, and then sends them in a specified manner, such as the dimension size of NUM_CH. Since the padding data generation circuit and the read data storage module 702 will not be performed at the same time, the matrix data generation circuit will perform simple signal selection based on the respective valid signals for the real feature data from the data storage module 702 read by the read address generation circuit and the padding_data of the padding data generation circuit, and obtain the collected valid signal (lb_data_valid, read data valid signal) and data signal (lb_data, here refers to the read data signal).
[0055] In this embodiment, the read control parameters include at least the number of valid data and the number of elements included in a single read and write operation. Whenever the target original feature data is read, a corresponding number of valid original feature data is selected from the target original feature data according to the number of valid data and stored in the shift register; whenever the amount of valid original feature data in the shift register at the current moment is equal to the number of elements, the valid original feature data of the shift register at the current moment is sent as the feature matrix data at the current moment, and a read data valid signal is output at the same time.
[0056] In order to avoid overflow, the storage space of the shift register is greater than or equal to the preset overflow value. The preset overflow value is determined according to the number of elements. For example, if the number of elements is NUM_CH, the preset overflow value can be NUM_CH*2. Correspondingly, the memory space of the shift register can be NUM_CH*2. After NUM_CH data are gathered in the shift register, they will be sent once, so the shift register will not overflow. For example, if Fig.19As shown, each time lb_data contains NUM_CH data, only params_conv_ci data among NUM_CH data are valid, so splicing is required, that is, NUM_CH valid data need to be gathered before sending. Whenever lb_data_valid is valid, that is, there is new lb_data, params_conv_ci valid data are obtained from lb_data, for example, they can be stored in a shift register of size NUM_CH*2 in order from right to left. At the same time, the old data in the shift register needs to be moved right by params_conv_ci data length.
[0057] In order to accurately control the sending of feature matrix data, the number of valid data stored in the shift register can be recorded by a counter. For the convenience of description, the counter used by the read control module is defined as a read counter, which can be represented by write_cnt. Accordingly, the shift register is connected to the read counter, and the read counter records the number of valid original feature data in the shift register at the current moment. When the read start signal is received, the read counter is reset. When the read data valid signal is valid and the amount of valid original feature data in the shift register at the current moment is not equal to the number of elements, the third count value of the read counter is adjusted to the sum of the current third count value and the number of valid data; when the read data valid signal is valid, the amount of valid original feature data in the shift register at the current moment is equal to the number of elements, and one round of indexing of the feature elements in the convolution kernel window is not completed, the third count value of the read counter is adjusted to the numerical difference between the number of valid data and the number of elements; when the read data valid signal is valid, the amount of valid original feature data in the shift register at the current moment is equal to the number of elements, and one round of indexing of the feature elements in the convolution kernel window is completed, the third count value of the read counter is adjusted to be the numerical difference between the number of valid data and the remaining number of elements based on the current third count value. When the data amount of the effective original feature data of the shift register at the current moment is equal to the number of elements, if the actual data amount of the effective original feature data of the shift register at the current moment is less than the number of elements, the number of padding transmission data is determined according to the difference between the number of elements and the remaining number of elements, and the corresponding padding transmission data is generated according to the number of padding transmission data, and the effective original feature data and padding transmission data of the shift register at the current moment are sent as the feature matrix data at the current moment.
[0058] In this embodiment, output_once_en is defined to indicate whether the number of data in the shift register has been pieced together to be NUM_CH and can be sent once, that is, when output_once_en is valid, it indicates whether the number of data in the shift register has been pieced together to be NUM_CH and can be sent once. y_index_cnt1_finish (the end signal of the second counter) is used to indicate whether the current sliding window is the last time it successfully pieced together. When y_index_cnt1_finish is valid, the current sliding window is the last time it pieced together enough NUM_CH valid data. params_conv_ci_last (the remaining number of elements) is defined to indicate that when the original feature data of a convolution window is pieced together in units of NUM_CH data, the last time it pieced together less than NUM_CH, that is, the remaining number of elements is the number of valid original feature data remaining when the original feature data of a convolution window is pieced together in units of the number of elements, when the number of elements is not pieced together in the last time. For example, when the convolution kernel is 3*3 and Ci is 10, the data covered by one slide is 3*3*10=90. After piecing together NUM_CH=32 data twice, the last time only has 26 data, that is, params_conv_ci_last=26. The logic of write_cnt recording the number of valid data in the shift register is as follows: Fig. 20 As shown, the implementation reflects the amount of data available in the shift register in real time: when params_lb_read_start (read start signal) is valid, write_cnt is reset to 0; when lb_data_valid is valid and output_once_en is invalid, write_cnt increases params_conv_ci; when lb_data_valid is valid, output_once_en is valid, and y_index_cnt1_finish is invalid, write_cnt+=(params_conv_ci-NUM_CH), when lb_data_valid is valid, output_once_en is valid, and y_index_cnt1_finish is valid, write_cnt+=(params_conv_ci-params_conv_ci_last).
[0059] As can be seen from the above, the read control module of this embodiment includes a data coordinate generation circuit, a read address generation circuit, a padding data generation circuit, and a matrix data generation circuit. The data coordinate generation circuit realizes the generation of coordinates of the original feature data according to the im2col mode through custom logic. The coordinates continuously generated by the data coordinate generation circuit are sent to the read address generation circuit and the padding data generation circuit in combination with the read control parameters, realizing the support for convolution modes with different convolution kernels, different expansion rates, and different padding sizes. Finally, the real feature data and padding data read from the on-chip cache are spliced through the matrix data generation circuit to form feature data and then sent out.
[0060] Finally, the present invention also provides an electronic device that can be used as a neural network model accelerator to achieve efficient completion of training tasks and reasoning tasks related to the neural network model. Fig.21 As shown, the electronic device may include an off-chip memory 211, a processor 212, and an image columnizer 502. The off-chip memory 211 stores original feature data; the image columnizer 502 reads the original feature data from the off-chip memory 211 and outputs corresponding feature matrix data; the processor 212 performs matrix multiplication and addition calculations on the feature matrix data and the corresponding feature matrix data to obtain a result matrix. The processor 212 may be, for example, a systolic array or a multiplication-addition tree, or other computing devices capable of performing multiplication and addition operations, which does not affect the implementation of the present invention. In addition, in this embodiment, the processor 212 and the image columnizer 502 may be deployed in a neural network model accelerator, such as an FPGA or a GPU, and the off-chip memory 211 may be a memory outside the neural network model accelerator, that is, the off-chip memory 211 and the neural network model accelerator together constitute an electronic device. Of course, the off-chip memory 211, processor 212, and image serializer 502 can all be deployed in a neural network model accelerator, such as an FPGA or a GPU, that is, the off-chip memory 211 can be a memory inside the neural network model accelerator, that is, the neural network model accelerator alone can constitute an electronic device.
[0061] As can be seen from the above, the image columnizer of this embodiment is integrated into the electronic device as a dedicated hardware module. The memory of the electronic device only needs to store a small amount of original feature data, and there is no need to cache the entire feature matrix data, which avoids directly storing the huge data after the im2col expansion, does not require high memory resources, and can also reduce the data transmission overhead from the off-chip memory to the on-chip memory. During the image columnization operation, the electronic device does not need to rely on an external central processing unit, and the communication and interaction operations between the neural network accelerator and the central processing unit are reduced, which effectively reduces the communication delay, is conducive to improving the execution efficiency of artificial intelligence tasks, and thus effectively reduces the memory resources and computing resources required for artificial intelligence tasks.
[0062] The above is a detailed introduction to an image serializer and electronic device provided by the present invention. The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can refer to each other. Whether the units and algorithm steps of each example described in each disclosed embodiment are executed in electronic hardware or computer software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods for each specific application to implement the described functions, and such implementation should not be considered to exceed the scope of the present invention. Without departing from the principles of the present invention, the present invention can also be improved and modified in a number of ways, and these improvements and modifications also fall within the scope of protection of the present invention.
Claims
1. An image arrayer, characterized in that: It includes an input terminal, a read-write control module, a data storage module and an output terminal; The input end receives read-write control parameters and original feature data; The read-write control module determines a data writing mode and a data reading mode based on the read-write control parameters, and writes the original feature data into the data storage module according to the data writing mode; Under the control of the data reading mode, corresponding data is read from the data storage module according to the convolution operation behavior mode to generate feature matrix data; The output end outputs the feature matrix data.
2. The image arrayer according to claim 1, characterized in that: The read / write control parameter includes at least the number of elements included in a single read / write operation, and the dimension of the original feature data includes the number of channels; Determine a data segmentation parameter according to the number of channels and the number of elements; According to the data segmentation parameters, the original feature data is stored in a target storage location.
3. The image arrayer according to claim 2, characterized in that: The number of channels is less than or equal to the number of elements, and each data and padding data having the same width dimension and height dimension in the original feature data are stored in each address space of the target storage location respectively; the length value of the padding data is the difference between the number of channels and the number of elements; The number of channels is greater than the number of elements. According to the number of elements, the data with the same width dimension and height dimension in the original feature data is cut into multiple data blocks, and each data block is stored in each address space of the target storage location respectively.
4. The image arrayer according to claim 1, characterized in that: The read / write control parameters include at least the number of elements included in a single read / write operation, the element width and storage parameters of a single read / write operation; The data storage module is a storage array, the number of memories included in the storage array is the same as the total number of memories in the storage parameters, and the total number of memories is determined according to the maximum convolution kernel size of the convolution operation; Each memory is independent of each other, and the read and write data width of each memory is determined according to the number of elements and the element bit width, and the total amount of stored data is determined according to the maximum address supported by a single memory in the storage parameters for reading and writing, and the corresponding read and write data width.
5. The image arrayer according to any one of claims 1 to 4, characterized in that: The read-write control module includes a write control module connected to the input end, the read-write control parameters include write control parameters, and the write control module at least includes a write counter, a write parameter input port and a write signal output port; The write parameter input port receives the write control parameter, and the write control parameter controls the amount of data written at a single time and the total amount of data written to each storage location of the data storage module; Based on the write control parameter, according to the write counter statistics the sending status of the original feature data and the storage parameter of the data storage module, determine the storage location information of the currently input sub-original feature data in the data storage module, and generate a write enable signal, a write address signal and a write data signal according to the storage location information and the sub-original feature data; sending the write enable signal, the write address signal and the write data signal to the data storage module through the write signal output port; The sub-original feature data is part of the original feature data received at the current moment.
6. The image arrayer according to claim 5, characterized in that: The data storage module includes a plurality of memories, the write parameter input port receives a single data transmission parameter, and the single data transmission parameter is that the amount of sub-original feature data sent each time is the same as the read and write data width of a single memory; the write counter includes a first write counter and a second write counter; Whenever a data valid parameter is received through the write parameter input port, if the first count value of the first write counter is not the write position switching parameter -1, the value of the current first count value is increased by a first preset value; If the first count value is the write position switching parameter -1, the first write counter is reset, and the value of the current second count value of the second write counter is increased by a second preset value, so that the address of the target memory currently being written is represented by the first write counter, and the second write counter represents the index of the target memory currently being written; The data validity parameter indicates whether the sub-original characteristic data sent according to the single data sending parameter is valid; and the write position switching parameter indicates the relationship between the number of data write times and the storage position.
7. The image arrayer according to any one of claims 1 to 4, characterized in that: The read-write control module includes a read control module, and the read-write control parameter includes a read control parameter; the read control module at least includes a read parameter input port, a read data port and a read data control circuit, the read data port is connected to the data storage module, and the read data control circuit is connected to the output port; The read parameter input port receives the read control parameter; The read data control circuit determines the target position information of the target original feature data covered by the convolution kernel during the sliding process according to the read control parameters, determines the physical position read information corresponding to the data storage module based on the target position information, and reads the target original feature data from the data storage module through the read data port according to the physical position read information, and generates feature matrix data according to the storage method of the original feature data and the target original feature data.
8. The image arrayer according to claim 7, characterized in that: The data reading control circuit at least includes a data coordinate generating circuit for outputting target position information; the target position information includes the right boundary position and the lower boundary position of the target original feature data, wherein the right boundary position is an x coordinate value and the lower boundary position is a y coordinate value; When the read control parameter does not include a fill mode read parameter, the read parameter input port receives a write position switching parameter and a data write position parameter; the write position switching parameter indicates the relationship between the number of data writes and the storage position; the data write position parameter indicates the data storage amount of the original feature data in the data storage module; the right boundary position of the target original feature data is determined according to the write position switching parameter, and the lower boundary position of the target original feature data is determined according to the data write position parameter; When the read control parameter includes a fill mode read parameter, the read parameter input port receives a write position switch parameter, a data write position parameter and the fill mode read parameter; the fill mode read parameter indicates that data is read from the data storage module in a fill mode, including a fill position and a fill length; the fill length includes a first fill length and a second fill length; When the filling position is left filling and right filling, the right boundary position of the target original feature data is determined according to the write position switching parameter and the double first filling length; when the filling position is left filling or right filling, the right boundary position of the target original feature data is determined according to the write position switching parameter and the first filling length; when the filling position is neither left filling nor right filling, the right boundary position of the target original feature data is determined according to the write position switching parameter; When the filling position is upper filling and lower filling, the lower boundary position of the target original feature data is determined according to the data write position parameter and the double second filling length; when the filling position is upper filling or lower filling, the lower boundary position of the target original feature data is determined according to the data write position parameter and the second filling length; when the filling position is neither upper filling nor lower filling, the lower boundary position of the target original feature data is determined according to the data write position parameter.
9. The image arrayer according to claim 7, characterized in that: The read data control circuit at least includes a read address generation circuit; The read address generation circuit determines a physical address according to target position information of the target original feature data and a storage method of the original feature data when the target position is valid; determines physical position read information of the data storage module according to the physical address and the target position information, and generates a data read signal according to the physical position read information; and the data read signal is sent to the read data port.
10. The image arrayer according to claim 9, characterized in that: Determining a physical address according to target location information of the target original feature data and a storage method of the original feature data includes: When data is read from the data storage module in a filling manner, and the filling position is left filling, the x coordinate in the physical address is the difference between the x coordinate of the target position information and the first filling length; when the filling position is not left filling, the x coordinate in the physical address is the x coordinate of the target position information; when the filling position is upper filling, the y coordinate in the physical address is the difference between the y coordinate of the target position information and the second filling length; when the filling position is not upper filling, the y coordinate in the physical address is the y coordinate of the target position information; When the original feature data is stored without adding padding elements, the x coordinate in the physical address is the x coordinate of the target position information, and the y coordinate in the physical address is the y coordinate of the target position information.
11. The image arrayer according to claim 10, characterized in that: Generating a data read signal according to the physical position read information includes: If the left padding is not enabled, when the x-coordinate value of the physical address is less than the write position switching parameter, it is located within the valid feature data range in the x-axis direction; if the left padding is enabled, when the x-coordinate of the target position information is greater than or equal to the first padding length, and the x-coordinate value of the physical address is less than the write position switching parameter, it is located within the valid feature data range in the x-axis direction; If the upper side padding is not enabled, when the y coordinate value of the physical address is less than the data write position parameter, it is located within the valid feature data range in the y-axis direction; if the upper side padding is enabled, when the y coordinate of the target position information is greater than or equal to the second padding length, and the x coordinate value of the physical address is less than the data write position parameter, it is located within the valid feature data range in the y-axis direction; When both the x-axis direction and the y-axis direction are within the valid feature data range, a data read signal is generated according to the storage position of the target original feature data in the data storage module and the corresponding read address as physical position read information.
12. The image arrayer according to claim 7, characterized in that: The read data control circuit includes at least a fill data generation circuit; The filling data generating circuit generates filling data according to the received target position information, physical address, filling mode parameter, data write position parameter, write position switching parameter, and whether it is in the valid characteristic data range signal, and outputs a valid filling signal and a filling data signal according to the filling data; Among them, the generation process of the filling data is as follows: if the left side filling is enabled, when the x-coordinate value of the target position information is less than the fixed filling length, the corresponding filling data is generated; if the right side filling is enabled, when the x-coordinate value of the physical address is greater than or equal to the write position switching parameter, the corresponding filling data is generated; if the upper side filling is enabled, when it is within the valid feature data range in the x-axis direction and the y-coordinate value of the physical address is less than the second filling length, the corresponding filling data is generated; if the lower side filling is enabled, when it is within the valid feature data range in the x-axis direction and the y-coordinate value of the physical address is greater than or equal to the data write position parameter, the corresponding filling data is generated.
13. The image arrayer according to claim 7, characterized in that: The read data control circuit at least includes a matrix data generation circuit; the read control parameters at least include the number of valid data and the number of elements included in a single read and write operation; Whenever the target original feature data is read, a corresponding number of valid original feature data is selected from the target original feature data according to the number of valid data, and stored in a shift register; the storage space of the shift register is greater than or equal to a preset overflow value, and the preset overflow value is determined according to the number of elements; Whenever the amount of valid original feature data of the shift register at the current moment is equal to the number of elements, the valid original feature data of the shift register at the current moment is sent as the feature matrix data at the current moment, and a read data valid signal is outputted at the same time.
14. The image arrayer according to claim 13, characterized in that: The shift register is connected to a read counter, and the read counter records the number of valid original feature data of the shift register at the current moment; When a read start signal is received, the read counter is reset; when a read data valid signal is valid and the amount of valid original feature data of the shift register at the current moment is not equal to the number of elements, the third count value of the read counter is adjusted to be the sum of the current third count value and the number of valid data; When the read data valid signal is valid, the amount of valid original feature data of the shift register at the current moment is equal to the number of elements, and a round of indexing of feature elements in the convolution kernel window has not been completed, the third count value of the read counter is adjusted to be the numerical difference between the number of valid data and the number of elements; When the read data valid signal is valid, the amount of valid original feature data of the shift register at the current moment is equal to the number of elements, and a round of indexing of feature elements in the convolution kernel window is completed, the third count value of the read counter is adjusted based on the current third count value, and the increased value is the difference between the number of valid data and the number of remaining elements; The remaining number of elements refers to the number of valid original feature data remaining when the original feature data of a convolution window is pieced together in units of the number of elements and when the number of elements cannot be pieced together for the last time.
15. An electronic device, characterized in that: comprising an off-chip memory, a processor and an image columnizer as claimed in any one of claims 1 to 14; Wherein, the off-chip memory stores original feature data; The image columnizer reads the original feature data from the off-chip memory and outputs corresponding feature matrix data; The processor performs matrix multiplication and addition calculation on the characteristic matrix data and the corresponding characteristic matrix data to obtain a result matrix.
Citation Information
Patent Citations
Computing device and computing method
CN110688157A
LCEVC video coding device and method based on FPGA
CN116437097A
Data processing method and device, electronic equipment and storage medium
CN116740262A
Design and implementation of efficient access circuit of convolutional neural network accelerator
CN119378610A
Data processing method, electronic device, medium and computer program product
CN119719595A