Method, apparatus, device, and computer-readable storage medium for parallel extraction of image data in multiple convolutional windows

By extracting image data and matrix transposition in multiple convolution windows in parallel, the performance problem caused by traditional serial processing is solved, and more efficient image convolution processing is achieved.

CN112306555BActive Publication Date: 2025-07-04KUNLUNXIN TECHNOLOGY (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201910694475.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-07-30
Publication Date
2025-07-04
Estimated Expiration
2039-07-30

AI Technical Summary

Technical Problem

The traditional image convolution processing method is to extract image data in the convolution window serially, resulting in insufficient processing performance, unable to fully utilize the parallelism of the hardware, and affecting the data conversion efficiency.

Method used

The method of extracting image data in multiple convolution windows in parallel and transposing the parallel matrix is ​​adopted. By dividing the image into multiple groups of convolution windows, the speed of data extraction and transposition is improved by using multiple data processing units to process in parallel.

Benefits of technology

The speed of image convolution processing is accelerated, processing efficiency is improved, hardware parallelism is fully utilized, and hardware resource overhead is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112306555B_ABST
    Figure CN112306555B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method, an apparatus, a device, and a computer-readable storage medium for parallelly extracting image data in multiple convolution windows. The method includes dividing an image into multiple groups of convolution windows, where the multiple groups of convolution windows include a first group of convolution windows and a second group of convolution windows, and each group of convolution windows includes multiple convolution windows. The method further includes using multiple data processing units to parallelly extract the image data in the multiple convolution windows of the first group of convolution windows, and after completing the extraction of the image data in the first group of convolution windows, using multiple data processing units to parallelly extract the image data in the multiple convolution windows of the second group of convolution windows. According to an embodiment of the present disclosure, during the process of convolution data extraction, using multiple data processing units to parallelly extract the image data in multiple convolution windows speeds up the data extraction speed, thereby improving the processing efficiency of image convolution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure generally relate to the field of image data processing technologies, and more particularly to a method, apparatus, device, computer-readable storage medium, and computer program product for parallelly extracting image data in multiple convolutional windows. Background Art

[0002] Machine learning refers to enabling a machine to learn patterns from a large amount of data like a human, so as to generate a machine learning model capable of completing some specific tasks. An artificial neural network is a typical machine learning technology, which creates an artificial neural network modeled on the human brain and allows a computer to learn through a large amount of data by using various machine learning algorithms. Common artificial neural networks include convolutional neural networks (CNNs), recurrent neural networks (RNNs), and so on. Deep learning is also a type of machine learning, but deep learning uses deep deep neural networks (DNNs) to make the model handle more complexly, so that the model can understand the data more deeply.

[0003] A CNN is a feedforward neural network with a convolutional calculation and a deep structure, and has very wide applications in the field of computer vision, especially image processing. From the perspective of a computer, an image is actually a two-dimensional or three-dimensional matrix. What a CNN does is to extract features from a two-dimensional or three-dimensional array by using operations such as convolution and pooling, and recognize the image. A CNN usually consists of an input layer, a convolutional layer, an activation function, a pooling layer, and a fully connected layer.

[0004] With the diversification of neural network models and the increasing demand for computing power, due to factors such as the performance and cost of traditional deep learning hardware platforms (such as general-purpose processors and graphics processing units GPU), the industry has begun to develop deep learning accelerators. One of the hardware cores of a deep learning accelerator is matrix operation, and the operation of the matrix operation module depends on the data supply from the upper level. In order to make full use of the computing power of the matrix operation module, an efficient and flexible data supply method is the focus of hardware design. Summary of the Invention

[0005] According to an exemplary embodiment of the present disclosure, there is provided a method, apparatus, device, and computer-readable storage medium for parallelly extracting image data in multiple convolutional windows.

[0006] In a first aspect of the present disclosure, a method for parallelly extracting image data in multiple convolutional windows is provided. The method includes: dividing an image into multiple groups of convolutional windows, where the multiple groups of convolutional windows include a first group of convolutional windows and a second group of convolutional windows; using multiple data processing units to parallelly extract the image data in the multiple convolutional windows in the first group of convolutional windows; and in response to completing the extraction of the image data in the first group of convolutional windows, using multiple data processing units to parallelly extract the image data in the multiple convolutional windows in the second group of convolutional windows.

[0007] In a second aspect of the present disclosure, an apparatus for parallelly extracting image data in multiple convolutional windows is provided. The apparatus includes: a convolutional window group division module configured to divide an image into multiple groups of convolutional windows, where the multiple groups of convolutional windows include a first group of convolutional windows and a second group of convolutional windows; a first parallel extraction module configured to use multiple data processing units to parallelly extract the image data in the multiple convolutional windows in the first group of convolutional windows; and a second parallel extraction module configured to, in response to completing the extraction of the image data in the first group of convolutional windows, use multiple data processing units to parallelly extract the image data in the multiple convolutional windows in the second group of convolutional windows.

[0008] In a third aspect of the present disclosure, an electronic device is provided, which includes one or more processors and a storage device, where the storage device is used to store one or more programs. When the one or more programs are executed by the one or more processors, the electronic device implements various methods and / or processes according to the embodiments of the present disclosure.

[0009] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, it implements various methods and / or processes according to the embodiments of the present disclosure.

[0010] In a fifth aspect of the present disclosure, a computer program product is provided, which includes a computer program, and when the computer program is executed by a processor, it implements various methods and / or processes according to the embodiments of the present disclosure.

[0011] It should be understood that the content described in the present invention content section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In combination with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more obvious. In the drawings, the same or similar reference numerals represent the same or similar elements, where:

[0013] Figure 1 A schematic diagram showing the convolutional process in a convolutional neural network is shown;

[0014] Figure 2 A flowchart showing a method for parallelly extracting image data in a plurality of convolutional windows according to an embodiment of the present disclosure;

[0015] Figure 3 A schematic diagram showing a process for parallelly extracting image data in a plurality of convolutional windows according to an embodiment of the present disclosure;

[0016] Figure 4 A schematic diagram showing an example architecture of an accelerator device for parallelly processing data according to an embodiment of the present disclosure;

[0017] Figure 5 A schematic diagram showing an example process for extracting convolutional data according to an embodiment of the present disclosure;

[0018] Figure 6 A schematic diagram showing an example process for parallel matrix transpose according to an embodiment of the present disclosure;

[0019] Figure 7 A block diagram showing a device for parallelly extracting image data in a plurality of convolutional windows according to an embodiment of the present disclosure; and

[0020] Figure 8 A block diagram showing an electronic device capable of implementing multiple embodiments of the present disclosure. Detailed Description of Specific Embodiments

[0021] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0022] In the description of the embodiments of the present disclosure, the term "including" and its like should be understood as an open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". There may also be other explicit and implicit definitions hereinafter.

[0023] Traditionally, during the process of image convolution processing, the convolution kernel slides over the image, and each time it extracts the pixel points of a convolution window and outputs them. However, the traditional method extracts the image data in different convolution windows serially, thus unable to perform data conversion efficiently, which affects the processing performance. In addition, the traditional scheme also performs matrix transpose serially when transposing the matrix. Therefore, the deficiencies of the prior art mainly lie in that while ensuring flexibility, the parallelism of the hardware is not fully utilized. Often, only one number or a group of numbers can be operated on each time, and data conversion cannot be performed efficiently, thus limiting the performance of subsequent calculations.

[0024] To this end, embodiments of the present disclosure propose a scheme for parallelly extracting image data in multiple convolution windows. According to embodiments of the present disclosure, during the process of convolution data extraction, multiple data processing units are used to parallelly extract the image data in multiple convolution windows, which speeds up the data extraction speed, thereby improving the processing efficiency of image convolution. In addition, some embodiments of the present disclosure also propose a parallel matrix transpose scheme. By multiple data processing units parallelly extracting multiple columns in a matrix, the matrix transpose speed is accelerated. The following will refer to the attached Figure 1-8 to describe in detail some exemplary implementations of embodiments of the present disclosure.

[0025] Figure 1 FIG. shows a schematic diagram of the convolution process 100 in a convolutional neural network. The convolutional neural network discovers some features of an image through image convolution. For example, it finds the edges of objects in the image, enhances or weakens a certain effect of the image, such as image blurring, sharpening, embossing effect, and so on.

[0026] Figure 1 FIG. shows an exemplary process of convolving an image 110 with a convolution kernel 120, where the convolution kernel 120 can be a 3×3 two-dimensional matrix. It should be understood that multiple convolution kernels can be used to convolve the image. The idea of image convolution is for a single pixel in the input image (such as image 110), and its value is weighted by the values of surrounding neighboring pixels, and the new pixel values generated by this weighted operation can generate a new output image (such as image 130) in order.

[0027] The convolution kernel 120 obtains convolution data by sliding over each convolution window in the image 110. As Figure 1 shown, first, the convolution kernel slides to the first convolution window 111 in the image 110. Through the multiplication and accumulation of the pixel points in the convolution window 111 and the convolution kernel 120 (as shown in 121), a convolution output 131 for the convolution window 111 is generated and stored in the image 130. For example, multiplying element by element and then adding, the obtained value is placed at the first element position of the output image matrix.

[0028] After the convolution of the convolution window 111 is completed, the convolution kernel is slid to the right by a distance of 1. Of course, it is also possible to choose to slide it a greater distance to the right. This distance is called the stride, which can be set in advance. As Figure 1 shown by the arrow 140 in, next, for the second convolution window 112 in the image 110, through the multiplication and accumulation of the pixels in the convolution window 112 and the convolution kernel 120 (as shown in 122), a convolution output 132 for the convolution window 111 is generated and stored in the image 130. Then, the above convolution process is repeated until the convolution kernel 120 slides over all the convolution windows in the image 110, thereby generating the convolved image 130. However, Figure 1 the convolution process described in extracts data serially and performs subsequent calculations, resulting in a slow speed of the convolution process.

[0029] Figure 2 shows a flowchart of a method 200 for parallelly extracting image data in multiple convolution windows according to an embodiment of the present disclosure. It should be understood that the method 200 can be executed by a dedicated accelerator device (such as an artificial intelligence AI chip), or can also be executed by a general-purpose computer or other dedicated computing devices.

[0030] In block 202, the image is divided into multiple groups of convolution windows, where the multiple groups of convolution windows include a first group of convolution windows and a second group of convolution windows. For example, according to the number of available data processing units (e.g., P), the image can be divided into multiple groups of convolution windows (each group of convolution windows includes P convolution windows), so that each group of convolution windows can be processed in parallel by multiple data processing units.

[0031] In block 204, multiple data processing units are used to parallelly extract the image data in multiple convolution windows in the first group of convolution windows. For example, the first group of convolution windows can include P convolution windows, and P data processing units in an acceleration device (such as an AI chip) are used to parallelly extract the image data in the P convolution windows, that is, each processing unit extracts the image data in a corresponding convolution window. In this way, the extraction speed of the image data in the convolution windows is accelerated.

[0032] In block 206, after the extraction of the image data in the first group of convolution windows is completed, multiple data processing units are used to parallelly extract the image data in multiple convolution windows in the second group of convolution windows. Generally, the number of convolution windows in an image can be much larger than the number of data processing units, so it is necessary to extract data in segments in parallel. For example, after P data processing units parallelly extract the image data in P convolution windows, the next P convolution windows are extracted, and the above steps are repeated in sequence until the image data in all the convolution windows in the image is extracted.

[0033] Therefore, according to an embodiment of the present disclosure, during the process of convolutional data extraction, multiple data processing units are used to extract image data in multiple convolutional windows in parallel, which speeds up the data extraction speed, thereby improving the processing efficiency of image convolution.

[0034] Figure 3 FIG. shows a schematic diagram of a process 300 for parallel extraction of image data in multiple convolutional windows according to an embodiment of the present disclosure. As Figure 3 shown, the convolutional windows 311, 312, 313 in the image 310 can be processed in parallel by the data processing units 321, 322, 323 respectively, and then the corresponding data 331, 332, 333 (which can be one-dimensional vectors respectively) in the convolutional windows 311, 312, 313 are mentioned in parallel. It should be understood that for the purpose of clear illustration, Figure 3 the stride of the convolutional windows in is 3, so that the convolutional windows 311, 312, 313 do not overlap. However, the stride can also be set to 1 or other values, so that there can be overlapping pixel points between different convolutional windows. In addition, for simplicity, Figure 3 only an image 310 of one color channel is shown in, however, the image 310 can also include multiple color channels.

[0035] Figure 4 FIG. shows a schematic diagram of an example architecture 400 of an accelerator device for parallel processing of data according to an embodiment of the present disclosure. As Figure 4 shown, the example architecture 400 can include a processor 410, a source memory 420, a target memory 425, a data conversion module 431, a scheduler 470, etc. The data conversion module 431 can act as a coprocessor, which includes an instruction storage unit 430, an instruction decoding unit 440, a control unit 450, a synchronization unit 460, a data reading unit 480, and multiple data processing units 490, where the multiple data processing units 490 can include, for example, P data processing units 491, 492, 493, 494, 499, etc.

[0036] The source memory 420 and the target memory 425 are the input memory and the output memory respectively, which can be off-chip memories (such as double data rate synchronous dynamic random access memory DDR) or on-chip memories (such as static random access memory SRAM), where the source memory 420 and the target memory 425 can be different memories or the same memory.

[0037] The instruction storage unit 430 is used to store the instructions received from the processor 410 for data conversion. The types of instructions can include, but are not limited to: parameter configuration instructions, transpose instructions, convolution data extraction instructions, synchronization instructions, etc. The parameter configuration instructions are used to configure parameters, which include, but are not limited to: data type, scale of the transpose matrix, scale of the convolution image, scale of the convolution kernel, convolution stride, number of edge padding pixels (pad), etc. The transpose instructions are used to configure the starting address of the source memory 420, the starting address of the target memory 425, the transpose data length, etc. The convolution data extraction instructions are used to configure the starting address of the source memory 420, the starting address of the target memory 425, the extracted data length, etc. The synchronization instructions are used to ensure that all instructions before this instruction are executed and the data is written to disk, so as to synchronize each module by the scheduler 470.

[0038] The instruction decoding unit 440 is used to read and parse the instructions from the instruction storage unit 430 when it detects that the instruction storage unit 430 is not empty and the current instruction is executable, and send the parsed content to the control unit 450. The control unit 450 generates corresponding control signals according to the configuration parameters, and the control content includes, but is not limited to: the read request behavior of the data reading unit 480, the behavior of the data processing unit 490, and the behavior of the synchronization unit 460.

[0039] The data reading unit 480 sends a read request to the source memory 420 according to the control signal of the control unit 450, and distributes the read data to multiple data processing units 490. Multiple data processing units 490 extract specific parts from the data from the data reading unit 480 according to the control signal of the control unit 450 and write them to the target memory 425. According to the embodiments of the present disclosure, multiple data processing units 490 can extract image data in multiple convolution windows in parallel, or transpose multiple columns in the matrix in parallel, thereby accelerating the speed of data conversion.

[0040] After receiving the synchronization request, the synchronization unit 460 outputs a synchronization completion signal to the external scheduler 470 after detecting that the current instruction is completed and the data is written to disk. It should be understood that the exemplary architecture 400 of the accelerator device is only an exemplary architecture including multiple data processing units 490, and other acceleration devices with multiple data processing units can also be used in combination with the embodiments of the present disclosure.

[0041] Figure 5 A schematic diagram showing an example process 500 for extracting convolution data according to an embodiment of the present disclosure is shown. As Figure 5 shown, given that the width of the image 510 is W, the height is H, and the channel depth is C, the width of each convolution window is S and the height is R (in Figure 5In the example, the convolution window size is 3×3). The accelerator device for performing image convolution includes a plurality of data processing units 520, such as including P data processing units 521, 522, 523, 529, etc. According to an embodiment of the present disclosure, the plurality of data processing units can extract image data in a plurality of convolution windows in parallel.

[0042] Reference Figure 5 , the data processing unit 521 is used to extract the image data in the convolution window 511. The data processing unit 521 first extracts the first row of data in the first channel (each data processing unit extracts the first row of data in the corresponding convolution window in parallel), then the data processing unit 521 extracts the second row of data in the first channel, and then extracts the third row of data in the first channel. Thus far, Figure 5 the extraction of the data in the first channel in the convolution window 511 in the example is completed. Next, the data processing unit 521 similarly extracts all the image data in the second channel of the convolution window 511, extracts all the image data in the third channel of the convolution window 511, and extracts all the image data in the fourth channel of the convolution window 511, thereby completing the data extraction process for the convolution window 511. As Figure 5 shown, the extracted data 530 includes the data 531 in the first channel (which includes a total of 9 values in three rows of the first channel), the data in the second channel, the data in the third channel, and the data 534 in the fourth channel. According to an embodiment of the present disclosure, since the P data processing units extract data in parallel, the P data processing units can complete the extraction of all the image data in the first P convolution windows in parallel.

[0043] Next, the plurality of data reading units 520 read the data of the subsequent P windows in parallel, and the method is the same as above. Finally, the extraction of the data corresponding to all the convolution windows in the image 510 is completed. Among them, since the P data processing units extract convolution data in parallel, each data processing unit needs to obtain the data of its corresponding convolution window according to the step parameter, and this part of the control behavior can be completed by the control unit.

[0044] In some embodiments, since the extracted convolutional window data is continuously stored in the target memory, the image data in a three-dimensional convolutional window of size C×R×S can be regarded as a one-dimensional vector of length C×R×S on the target memory after being extracted by the data processing unit. Assuming that a total of N convolutional window data are extracted from the image 510, then what is finally stored on the target memory is a two-dimensional matrix with N rows and C×R×S columns. The convolutional kernel can also be regarded as a two-dimensional matrix with F rows and C×R×S columns. After the convolutional kernel is transposed, it becomes a two-dimensional matrix with C×R×S rows and F columns. Then the complex image convolution operation is transformed into the multiplication of two two-dimensional matrices at this time. As shown in the following formula (1), D represents the image data matrix, W represents the weight data matrix. The image data contained in a convolutional window is, for example, the left dotted box (i.e., a one-dimensional vector of length C×R×S), and the weight data contained in a convolutional kernel is, for example, the right dotted box. In this way, the matrix operation efficiency in the convolution operation can be further improved.

[0045]

[0046] Figure 6 FIG. shows a schematic diagram of an example process 600 for parallel matrix transpose according to an embodiment of the present disclosure. As Figure 6 shown, assume that the matrix to be transposed is a matrix 610 of size M×N. Referring to the above Figure 4 described data conversion module, it may include P data processing units working in parallel. Then the matrix 610 is partitioned by P columns as the granularity, that is, the first partition includes the first P columns in the front, the second partition includes the subsequent P columns, and so on.

[0047] As Figure 6 shown, the multiple data processing units 620 include P data processing units, such as data processing units 621, 622, 623, 629, etc. The data reading unit reads one row of the matrix data each time, and each data processing unit will process the corresponding columns in the row data in parallel. For example, the data processing unit 621 will process the data in the first column (column 0), the data processing unit 622 will process the data in the second column (column 1), the data processing unit 623 will process the data in the third column (column 2), and the data processing unit 629 will process the data in the Pth column (column P - 1).

[0048] After the multiple data processing units 620 finish processing the P columns of the first partition in parallel, they continue to process the P columns of data in the next partition until the entire matrix 621 is transposed to generate the transposed matrix 630. As Figure 6As shown, data processing unit 621 transposes the first column in matrix 610 into the first row in matrix 630, data processing unit 622 transposes the second column in matrix 610 into the second row in matrix 630, and data processing unit 629 transposes the Pth column in matrix 610 into the Pth row in matrix 630. In some embodiments, the control unit needs to maintain the write addresses of the target memories of the P data processing units respectively according to the parameters configured by the instruction and the starting address of the target memory.

[0049] Therefore, in the process of convolutional data extraction in the embodiments of the present disclosure, multiple data processing units are used to extract image data in multiple convolutional windows in parallel, which can accelerate the data extraction speed, thereby improving the processing efficiency of image convolution. In addition, in some embodiments of the present disclosure, multiple columns in a matrix are extracted in parallel by multiple data processing units, which can accelerate the matrix transpose speed.

[0050] Figure 7 FIG. shows a block diagram of an apparatus 700 for parallel extraction of image data in multiple convolutional windows according to an embodiment of the present disclosure. As Figure 7 shown, the apparatus 700 includes a convolutional window group partitioning module 710, a first parallel extraction module 720, and a second parallel extraction module 730. The convolutional window group partitioning module 710 is configured to partition an image into multiple groups of convolutional windows, where the multiple groups of convolutional windows include a first group of convolutional windows and a second group of convolutional windows. The first parallel extraction module 720 is configured to use multiple data processing units to parallelly extract image data in multiple convolutional windows in the first group of convolutional windows. The second parallel extraction module 730 is configured to, in response to completion of extraction of image data in the first group of convolutional windows, use multiple data processing units to parallelly extract image data in multiple convolutional windows in the second group of convolutional windows.

[0051] In some embodiments, the first group of convolutional windows includes a first convolutional window and a second convolutional window, and the first parallel extraction module 720 includes: a first data extraction module configured to use a first data processing unit to extract image data in the first convolutional window; and a second data extraction module configured to use a second data processing unit to extract image data in the second convolutional window.

[0052] In some embodiments, the first data extraction module includes: a first extraction module configured to extract the first row of image data in the first channel of the first convolutional window; a second extraction module configured to extract the second row of image data in the first channel of the first convolutional window; and a third extraction module configured to extract the third row of image data in the first channel of the first convolutional window.

[0053] In some embodiments, the first data extraction module further includes: a second channel extraction module configured to, in response to completing the extraction of all image data in the first channel of the first convolution window: extract the first row of image data in the second channel of the first convolution window; extract the second row of image data in the second channel of the first convolution window; and extract the third row of image data in the second channel of the first convolution window.

[0054] In some embodiments, the first data extraction module further includes: a data representation module configured to, in response to completing the extraction of all image data in all channels of the first convolution window, represent all image data in the first convolution window using a one-dimensional vector, where the length of the one-dimensional vector is the product of the number of channels of the image, the number of rows of each convolution window, and the number of columns of each convolution window.

[0055] In some embodiments, the apparatus 700 further includes: a data storage module configured to store all image data in multiple sets of convolution windows in a target memory using a two-dimensional matrix, where the number of rows in the two-dimensional matrix is the number of all convolution windows in the multiple sets of convolution windows, and the number of columns in the two-dimensional matrix is the product of the number of channels of the image, the number of rows of each convolution window, and the number of columns of each convolution window.

[0056] In some embodiments, the apparatus 700 further includes: a block partitioning module configured to partition the matrix into multiple blocks in units of columns, the multiple blocks including a first block and a second block; a first parallel transpose module configured to use multiple data processing units to transpose multiple columns of data in the first block in parallel; and a second parallel transpose module configured to, in response to completing the transposition of multiple columns of data in the first block, use multiple data processing units to transpose multiple columns of data in the second block in parallel.

[0057] In some embodiments, the first parallel transpose module includes: a first matrix transpose module configured to use a first data processing unit among the multiple data processing units to transpose the first column of data in the first block; and a second matrix transpose module configured to use a second data processing unit among the multiple data processing units to transpose the second column of data in the second block.

[0058] In some embodiments, the block partitioning module includes: a second block partitioning module configured to partition the matrix into multiple blocks based on the number of multiple data processing units.

[0059] It should be understood that Figure 7 the convolution window group partitioning module 710, the first parallel extraction module 720, and the second parallel extraction module 730 shown in can be included in a single or multiple electronic devices. Moreover, it should be understood that Figure 7The modules shown can perform the steps and / or actions in the methods and / or processes with reference to the embodiments of the present disclosure.

[0060] Therefore, the embodiments of the present disclosure propose a programmable data conversion method and device applicable to a deep learning accelerator, which can flexibly support matrix transposes and convolution window extractions of images of various scales, and at the same time can make full use of the parallelism characteristics of the hardware to efficiently provide data so as to exert the performance of the matrix operation module. The embodiments of the present disclosure ensure the flexibility of data conversion through programmability, and then use the way of multiple processing units working in parallel to efficiently convert data. In addition, for transposes and convolutions, the embodiments of the present disclosure can reuse the same set of hardware structures to reduce the final implemented hardware overhead.

[0061] Therefore, the benefits of some embodiments of the present disclosure may include but are not limited to: multiple data processing units work in parallel to efficiently complete data conversion work; the processor issues parameter configuration instructions to flexibly configure parameters, which can adapt to data conversions of various scales; through the data conversion method of convolutional data extraction, complex convolution operations can be converted into simple matrix multiplications; transpose and convolutional data extraction can be completed by reusing the same set of hardware structures, saving hardware resources.

[0062] Figure 8 A schematic block diagram of an example device 800 that can be used to implement the embodiments of the present disclosure is shown. It should be understood that device 800 can be a device 700 for implementing the parallel extraction of image data in multiple convolution windows described in the present disclosure. As shown, device 800 includes a central processing unit (CPU) 801, which can execute various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 802 or computer program instructions loaded from a storage unit 808 into a random access memory (RAM) 803. In RAM 803, various programs and data required for the operation of device 800 can also be stored. The CPU 801, ROM 802, and RAM 803 are connected to each other through a bus 804. As Figure 8 shown, an input / output (I / O) interface 805 is also connected to the bus 804.

[0063] Multiple components in device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a disk, an optical disc, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0064] The processing unit 801 executes the various methods and processes described above, such as method 200. For example, in some embodiments, the method may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the CPU 801, one or more actions or steps of the methods described above may be performed. Alternatively, in other embodiments, the CPU 801 may be configured to execute the method by any other suitable means (e.g., by means of firmware).

[0065] The functions described above herein may be performed at least in part by one or more hardware logic components. By way of example and not limitation, the types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0066] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0067] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be either a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0068] In addition, although the various acts or steps are depicted in a particular order, this should be understood as requiring that such acts or steps be performed in the particular order shown or in a sequential order, or that all illustrated acts or steps be performed to achieve the desired result. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the foregoing discussion, these should not be construed as limiting the scope of the present disclosure. Certain features that are described in the context of separate embodiments may also be implemented in combination in a single implementation. Conversely, the various features that are described in the context of a single implementation may also be implemented separately or in any suitable sub-combination in multiple implementations.

[0069] Although embodiments of the present disclosure have been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A method for parallelly extracting image data in multiple convolution windows, comprising: Dividing an image into multiple groups of convolution windows, the multiple groups of convolution windows including a first group of convolution windows and a second group of convolution windows, wherein the first group of convolution windows includes a first convolution window, the first convolution window including R rows of pixels, S columns of pixels, and C channels of pixels, and each of R, S, and C is an integer greater than 1; Using multiple data processing units to parallelly extract the image data in multiple convolution windows in the first group of convolution windows, wherein using multiple data processing units to parallelly extract the image data in multiple convolution windows in the first group of convolution windows includes: extracting first three-dimensional image data by using the first convolution window, wherein the first three-dimensional image data includes R rows of pixels, S columns of pixels, and C channels of pixels, and converting the first three-dimensional image data into first one-dimensional image data, wherein the first one-dimensional image data includes R×S×C pixels arranged in rows, and wherein using multiple data processing units to parallelly extract the image data in multiple convolution windows in the first group of convolution windows includes: extracting second three-dimensional image data by using the first convolution window, wherein the second three-dimensional image data includes R rows of pixels, S columns of pixels, and C channels of pixels; converting the second three-dimensional image data into second one-dimensional image data including R×S×C pixels arranged in rows; And forming a first two-dimensional matrix with R×S×C columns, wherein the first two-dimensional matrix includes the first one-dimensional image data and the second one-dimensional image data, and the first one-dimensional image data and the second one-dimensional image data are arranged in different rows of the two-dimensional matrix; And In response to completing the extraction of the image data in the first group of convolution windows, using the multiple data processing units to parallelly extract the image data in multiple convolution windows in the second group of convolution windows; Obtaining a second two-dimensional matrix of the convolution kernel of the first convolution window from a memory storing data of the convolution kernel, the second two-dimensional matrix having R×S×C columns; Obtaining a third two-dimensional matrix by transposing the second two-dimensional matrix, the third two-dimensional matrix having R×S×C rows; and Multiplying the first two-dimensional matrix by the third two-dimensional matrix.

2. The method according to claim 1, wherein the multiple data processing units include a first data processing unit and a second data processing unit, the first group of convolution windows includes a second convolution window, and wherein using multiple data processing units to parallelly extract the image data in multiple convolution windows in the first group of convolution windows includes: Using the first data processing unit to extract the image data in the first convolution window; And Using the second data processing unit to extract the image data in the second convolution window.

3. The method according to claim 2, wherein using the first data processing unit to extract the image data in the first convolution window includes: Extracting the first row of image data in the first channel of the first convolution window; Extracting the second row of image data in the first channel of the first convolution window; And Extracting the third row of image data in the first channel of the first convolution window.

4. The method according to claim 3, wherein extracting the image data in the first convolutional window using the first data processing unit further comprises: In response to completing the extraction of all the image data in the first channel of the first convolutional window: Extracting the first row of image data in the second channel of the first convolutional window; Extracting the second row of image data in the second channel of the first convolutional window; and Extracting the third row of image data in the second channel of the first convolutional window.

5. The method according to claim 1, further comprising: Dividing the matrix into multiple sub-blocks column by column, the multiple sub-blocks including a first sub-block and a second sub-block; Using the multiple data processing units to transpose multiple columns of data in the first sub-block in parallel; And In response to completing the transposition of the multiple columns of data in the first sub-block, using the multiple data processing units to transpose multiple columns of data in the second sub-block in parallel.

6. The method according to claim 5, wherein using the multiple data processing units to transpose multiple columns of data in the first sub-block in parallel comprises: Using a first data processing unit among the multiple data processing units to transpose the first column of data in the first sub-block; And Using a second data processing unit among the multiple data processing units to transpose the second column of data in the second sub-block.

7. The method according to claim 5, wherein dividing the matrix into multiple sub-blocks column by column comprises: Dividing the matrix into the multiple sub-blocks based on the number of the multiple data processing units.

8. An apparatus for parallelly extracting image data in multiple convolutional windows, comprising: A convolutional window group division module configured to divide an image into multiple groups of convolutional windows, the multiple groups of convolutional windows including a first group of convolutional windows and a second group of convolutional windows, wherein the first group of convolutional windows includes a first convolutional window, the first convolutional window includes R rows of pixels, S columns of pixels, and C channels of pixels, and each of R, S, and C is an integer greater than 1; A first parallel extraction module configured to use multiple data processing units to parallelly extract the image data in multiple convolutional windows in the first group of convolutional windows, wherein using the multiple data processing units to parallelly extract the image data in multiple convolutional windows in the first group of convolutional windows includes: extracting first three-dimensional image data by using the first convolutional window, wherein the first three-dimensional image data includes R rows of pixels, S columns of pixels, and C channels of pixels, and converting the first three-dimensional image data into first one-dimensional image data, wherein the first one-dimensional image data includes R×S×C pixels arranged in rows, wherein using the multiple data processing units to parallelly extract the image data in multiple convolutional windows in the first group of convolutional windows includes: extracting second three-dimensional image data by using the first convolutional window, wherein the second three-dimensional image data includes R rows of pixels, S columns of pixels, and C channels of pixels; converting the second three-dimensional image data into second one-dimensional image data including R×S×C pixels arranged in rows; and a first two-dimensional matrix formed with R×S×C columns, wherein the first two-dimensional matrix includes the first one-dimensional image data and the second one-dimensional image data, and the first one-dimensional image data and the second one-dimensional image data are arranged in different rows of the two-dimensional matrix; and a second parallel extraction module configured to, in response to completing the extraction of the image data in the first set of convolution windows, use the plurality of data processing units to parallelly extract the image data in the plurality of convolution windows in the second set of convolution windows; obtain a second two-dimensional matrix of the convolution kernels of the first convolution window from a memory storing the data of the convolution kernels, the second two-dimensional matrix having R×S×C columns; obtain a third two-dimensional matrix by transposing the second two-dimensional matrix, the third two-dimensional matrix having R×S×C rows; and multiply the first two-dimensional matrix by the third two-dimensional matrix.

9. The apparatus according to claim 8, wherein the plurality of data processing units include a first data processing unit and a second data processing unit, the first set of convolution windows includes a second convolution window, and wherein the first parallel extraction module includes: a first data extraction module configured to use the first data processing unit to extract the image data in the first convolution window; and a second data extraction module configured to use the second data processing unit to extract the image data in the second convolution window.

10. The apparatus according to claim 9, wherein the first data extraction module includes: a first extraction module configured to extract the first row of image data in the first channel of the first convolution window; a second extraction module configured to extract the second row of image data in the first channel of the first convolution window; and a third extraction module configured to extract the third row of image data in the first channel of the first convolution window.

11. The apparatus according to claim 10, wherein the first data extraction module further includes: a second channel extraction module configured to, in response to completing the extraction of all the image data in the first channel of the first convolution window: extract the first row of image data in the second channel of the first convolution window; extract the second row of image data in the second channel of the first convolution window; and extract the third row of image data in the second channel of the first convolution window.

12. The apparatus according to claim 8, further comprising: a block division module configured to divide the matrix into a plurality of blocks in units of columns, the plurality of blocks including a first block and a second block; a first parallel transpose module configured to use the plurality of data processing units to parallelly transpose the multiple columns of data in the first block; and a second parallel transpose module configured to, in response to completing the transposition of the multiple columns of data in the first block, use the plurality of data processing units to parallelly transpose the multiple columns of data in the second block.

13. The apparatus according to claim 12, wherein the first parallel transpose module includes: a first matrix transpose module configured to use the first data processing unit among the plurality of data processing units to transpose the first column of data in the first block; and A second matrix transpose module, configured to transpose the second column data in the second sub-block by using a second data processing unit among the plurality of data processing units.

14. The apparatus according to claim 12, wherein the sub-block division module comprises: A second sub-block division module, configured to divide the matrix into the plurality of sub-blocks based on the number of the plurality of data processing units.

15. An electronic device, the electronic device comprising: One or more processors; And A storage device for storing one or more programs, which when executed by the one or more processors cause the electronic device to implement the method according to any one of claims 1-7.

16. A computer-readable storage medium, having stored thereon a computer program, which when executed by a processor implements the method according to any one of claims 1-7.

17. A computer program product, comprising a computer program, which when executed by a processor implements the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Indirectly accessing sample data to perform multi-convolution operations in parallel processing system

    CN105678378A

  • Convolution operation processing method and related product

    CN108304923A