Control method of continuous convolution structure
Through the depth-first Winograd convolution transformation algorithm and row interleaving calculation method, the convolution neural network architecture is optimized, and the problem of delay and memory requirements of convolutional neural networks under limited hardware resources in the existing technology is solved, and the operation performance is optimized and power consumption is reduced.
Patent Information
- Application Number
- CN202410189244.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-30
- Filing Date
- 2024-02-20
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to realize the real-time calculation and delay reduction of convolutional neural networks under limited hardware resources, especially under low bandwidth conditions, the latency and memory requirements of convolutional neural networks are still not effectively met.
The depth-first Winograd convolution transformation algorithm is adopted, combined with the calculation method of row interleaving, the convolutional neural network architecture is optimized, the number of multipliers is reduced, the chip area and power consumption is reduced, and the delay is minimized.
Through the optimized Winograd convolution algorithm, the operation performance of the convolution structure is optimized, reducing the usage area and power consumption of the multiplication accumulation operator, and reducing the operation delay and memory requirements.
Smart Images

Figure CN120068949A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a Winograd convolution transform algorithm using depth - first, and particularly relates to a control method suitable for a continuous convolution structure with a convolution kernel size of 3*3. Through a line - interleaved algorithm, the use of multiply - accumulate units and the power consumption generated thereby can be effectively reduced. Background Art
[0002] In recent years, with the rapid development of modern artificial intelligence (AI) technology, convolutional neural networks (CNNs) have been widely applied in AI - related technology industries. In the application fields of computational imaging such as super - resolution, noise reduction, and image sharpening, convolutional neural networks (CNNs) have shown excellent performance. However, with the development of convolutional neural network algorithms and the continuous improvement of resolution and frame rate, their computational complexity and hardware resource requirements are still increasing. Therefore, how to optimize convolutional neural network algorithms to complete real - time calculations and reduce computational latency under limited hardware resources is one of the research goals that still need to be worked on so far.
[0003] Generally speaking, for general AI algorithm technologies, the computational workload in the convolutional layer of a convolutional neural network structure is quite dense and huge, so a large amount of computing resources must be relied on. Among them, common algorithms, such as Winograd, are one of the fast convolution algorithms that can be used to effectively reduce the cost of multiplication operations. On the other hand, since the chip memory resources and memory bandwidth are also important indicators, these indicators usually deteriorate as the requirements for image resolution and frame rate in consumer electronics increase. In view of these deficiencies in the prior art, a row operation-based and depth-first image processing method has been proposed, hoping to effectively reduce the latency of the convolutional neural network and its memory requirements under low-bandwidth (or zero-bandwidth) conditions. However, it is known that in such an algorithm, the operations in each convolutional layer are independent of each other and incompatible with the pairwise multiplication rule in the Winograd convolution algorithm. Even if the row operation-based algorithm can be applied to the existing Winograd convolution algorithm, its low latency and low memory requirements still cannot be met. How to integrate the Winograd convolution algorithm with the row operation-based and depth-first image processing method while retaining their advantages is still an unsolved problem in the prior art. In view of this, professionals in this field are indeed in urgent need of developing a novel and creative algorithm that can improve the existing Winograd convolution algorithm, thereby solving the problems existing in the above-mentioned prior art and further optimizing the overall computational efficiency of the convolutional structure. Summary of the Invention
[0004] To solve the above-mentioned many deficiencies, one object of the present invention is to provide an optimized Winograd convolution algorithm. According to the technical solution disclosed in the present invention, such an optimized Winograd convolution algorithm is based on a depth-first process and combines a line interleaved calculation method to minimize the latency in the convolutional neural network architecture.
[0005] Another object of the present invention is to provide a control method suitable for a continuous convolution structure with a kernel size of 3*3. By adopting such a novel control method, the number of multipliers required when executing the algorithm can be effectively reduced, thereby reducing the chip area and power consumption. At the same time, the operation latency is also minimized, and the operation data that needs to be temporarily stored in each convolutional layer is reduced, thereby reducing the additional required area of static random-access memories (SRAM).
[0006] Specifically, according to an embodiment of the present invention, the present invention aims to provide a control method for a continuous convolution structure. In one embodiment, such a control method includes the following steps:
[0007] (a) In a first operation interval, receive the first row and the second row of input data, and perform a Winograd convolution transform to generate the first row of output data;
[0008] (b) In a second operation interval, receive the third row and the fourth row of the input data, and perform a Winograd convolution transform to generate the second row and the third row of the output data; and
[0009] (c) In a third operation interval, receive the fifth row and the sixth row of the input data, and perform a Winograd convolution transform to generate the fourth row and the fifth row of the output data.
[0010] On the other hand, according to another embodiment of the present invention, the control method for the continuous convolution structure also includes the following steps:
[0011] (a) In a first operation interval, receive the first row, the second row, and the third row of input data, and perform a Winograd convolution transform to generate the first row and the second row of output data;
[0012] (b) In a second operation interval, receive the fourth row and the fifth row of the input data, and perform a Winograd convolution transform to generate the third row and the fourth row of the output data; and
[0013] (c) In a third operation interval, receive the sixth row and the seventh row of the input data, and perform a Winograd convolution transform to generate the fifth row and the sixth row of the output data.
[0014] According to the above embodiments, when performing the Winograd convolution transform, the present invention can also selectively generate at least one padding line at the same time, so as to provide dummy data, zero-point data, or duplicate image data through the padding line.
[0015] Alternatively, the at least one padding line can also selectively provide the same image data as the first row of the input data, so as to facilitate the execution of the Winograd convolution transform. Those skilled in the art can optionally design applicable implementation forms according to their needs.
[0016] Generally speaking, according to the control method of the continuous convolution structure disclosed in the present invention, the algorithm can also be preferably applied to other existing known convolution architectures. Once the disclosure content of this application is known, other alternative and modified exemplary examples will be obvious to those skilled in the art. However, the present invention is not limited to the disclosed embodiments, and the patent scope claimed by the present invention requires and covers other alternative examples and modified examples equivalent thereto.
[0017] At the same time, it should be understood that the foregoing technical abstract and the following detailed description are both exemplary and are intended to provide further explanations for the claims of the present invention. Hereinafter, in order to enable those skilled in the art to have a further understanding and recognition of the structural features and achieved effects of the present invention, preferred embodiment figures and detailed descriptions are provided as follows. Brief Description of the Drawings
[0018] Figure 1 It is a schematic diagram when receiving the first row of input data in the control method of the continuous convolution structure disclosed in the present invention.
[0019] Figure 2 It is a schematic diagram when receiving the second row of input data and the first and second padding rows in the control method of the continuous convolution structure disclosed in the present invention.
[0020] Figure 3 It is a schematic diagram when receiving the third row of input data in the control method of the continuous convolution structure disclosed in the present invention.
[0021] Figure 4 It is a schematic diagram when receiving the fourth row of input data and the third, fourth, and fifth padding rows in the control method of the continuous convolution structure disclosed in the present invention.
[0022] Figure 5 It is a schematic diagram when receiving the fifth row of input data in the control method of the continuous convolution structure disclosed in the present invention.
[0023] Figure 6 It is a schematic diagram when receiving the sixth row of input data and the sixth, seventh, and eighth padding rows in the control method of the continuous convolution structure disclosed in the present invention.
[0024] Figure 7 It is a schematic diagram when receiving the seventh row of input data in the control method of the continuous convolution structure disclosed in the present invention.
[0025] Figure 8 It is a schematic diagram when receiving the eighth row of input data in the control method of the continuous convolution structure disclosed in the present invention.
[0026] Figure 9 It is a schematic diagram when receiving the ninth row of input data in the control method of the continuous convolution structure disclosed in the present invention.
[0027] Figure 10 It is a schematic diagram when receiving the tenth row of input data in the control method of the continuous convolution structure disclosed in the present invention.
[0028] Figure 11 It is a flowchart of the steps for obtaining the output sequence in the odd output layer in the control method of the continuous convolution structure disclosed in the present invention.
[0029] Figure 12 It is disclosed according to Figure 11 The flowchart shown, where the schematic diagram of the first operation interval.
[0030] Figure 13 It is disclosed according to Figure 11 The flowchart shown, where the schematic diagram of the second operation interval.
[0031] Figure 14 It is disclosed according to Figure 11 The flowchart shown, where the schematic diagram of the third operation interval.
[0032] Figure 15 It is a flowchart of the steps for obtaining the output sequence in the even output layer in the control method of the continuous convolution structure disclosed in the present invention.
[0033] Figure 16 It is disclosed according to Figure 15 The flowchart shown, where the schematic diagram of the first operation interval.
[0034] Figure 17 It is disclosed according to Figure 15 The flowchart shown, where the schematic diagram of the second operation interval.
[0035] Figure 18 It is disclosed according to Figure 15 The flowchart shown, where the schematic diagram of the third operation interval.
[0036]
Symbol Explanation
[0037] 10: Continuous convolution structure
[0038] Lin: Input layer
[0039] LO1: First output layer
[0040] LO2: Second output layer
[0041] LO3: Third output layer
[0042] LO4: The fourth output layer
[0043] LO5: The fifth output layer
[0044] X1: The first filling row
[0045] X2: The second filling row
[0046] X3: The third filling row
[0047] X4: The fourth filling row
[0048] X5: The fifth filling row
[0049] X6: The sixth filling row
[0050] X7: The seventh filling row
[0051] X8: The eighth filling row
[0052] S1102, S1104, S1106, S1108, S1110, S1112, S1114, S1116: Steps
[0053] S1502, S1504, S1506, S1508, S1510, S1512, S1514, S1516, S1518: Steps Detailed implementation manners
[0054] Embodiments of the present invention will be further illustrated below in conjunction with the relevant drawings. As much as possible, in the drawings and the specification, the same reference numerals represent the same or similar components. In the drawings, for the sake of simplicity and convenience of labeling, the shapes and thicknesses may be exaggerated. It can be understood that the elements not specifically shown in the drawings or described in the specification are in the forms known to those skilled in the art. Those skilled in the art can make various changes and modifications based on the content of the present invention.
[0055] Unless otherwise specified, some conditional clauses or words, such as "can", "could", "might", or "may", generally attempt to express that the embodiments of the present application have, but can also be interpreted as features, elements, or steps that may not be required. In other embodiments, these features, elements, or steps may not be required.
[0056] The description of "an embodiment" or "one embodiment" hereinafter refers to a specific element, structure, or feature associated with at least one embodiment. Therefore, the multiple descriptions of "an embodiment" or "one embodiment" that appear in multiple places hereinafter are not directed to the same embodiment. Furthermore, the specific components, structures, and features in one or more embodiments can be combined in a suitable manner.
[0057] In the specification and claims, certain terms are used to refer to specific elements. However, those skilled in the art should understand that the same element may be referred to by different names. The specification and claims do not distinguish elements by the difference in names, but by the difference in their functions. The term "comprising" mentioned in the specification and claims is an open-ended term and should be interpreted as "comprising but not limited to". In addition, "coupled" herein includes any direct and indirect connection means. Therefore, if it is described in the text that the first element is coupled to the second element, it means that the first element can be directly connected to the second element through electrical connection, wireless transmission, optical transmission or other signal connection means, or can be indirectly electrically or signal-connected to the second element through other elements or connection means.
[0058] The disclosure is specifically described by the following examples, which are only for illustrative purposes. Because for those skilled in the art, without departing from the spirit and scope of the present disclosure, various modifications and refinements can be made. Therefore, the protection scope of the present disclosure shall be subject to the scope defined by the appended claims. Throughout the specification and claims, unless the context clearly dictates otherwise, the meanings of "a" and "the" include such descriptions including "one or at least one" of such element or component. In addition, as used in the present disclosure, unless it is clearly visible from a specific context that multiple elements are excluded, the singular article also includes the description of multiple elements or components. Moreover, when applied in the description herein and in the following claims, unless the context clearly dictates otherwise, the meaning of "therein" can include "therein" and "thereon". The terms used throughout the specification and claims, unless otherwise noted, generally have the ordinary meaning of each term as used in this field, in the content disclosed herein and in the specific context. Certain terms used to describe the present disclosure will be discussed below or elsewhere in this specification to provide additional guidance to practitioners in the description of the present disclosure. The use of examples anywhere in the specification, including the use of examples of any term discussed herein, is only for illustrative purposes and does not limit the scope and meaning of the present disclosure or any illustrative term. Similarly, the present disclosure is not limited to the various embodiments presented in this specification.
[0059] In the following paragraphs, the present invention will provide a control method suitable for a continuous convolution structure. The control method disclosed herein by the present invention is exemplarily applied to a convolution structure with a (3*3) convolution kernel size. However, the control method provided in the following description of the present invention can also be applied to a convolution architecture with other convolution kernel sizes. The present invention is not limited by the following embodiments.
[0060] Please refer to Figure 1As shown, it is a schematic diagram of a continuous convolution structure for image processing before the control method described in the present invention is disclosed. As shown in the figure, the continuous convolution structure 10 includes an input layer Lin, a first output layer LO1, a second output layer LO2, a third output layer LO3, a fourth output layer LO4, and a fifth output layer LO5. Among them, in order to perform the image transformation program, between each level, the Winograd convolution transform can be used to perform image processing. Generally speaking, the Winograd convolution transform may include, for example, but not limited to: Winograd transform, matrix element-wise multiplication, and Fast Fourier Transformation. The algorithm based on the Winograd convolution transform is well-known in the technical field related to image processing, so the present invention will not repeat it here. Generally speaking, in the present invention, the Winograd convolution transform is applied between each image data layer. For example, between the input layer Lin and the first output layer LO1, so as to transform and output the image data of the input layer Lin into the image data of the first output layer LO1. Similarly, between the first output layer LO1 and the second output layer LO2, the Winograd convolution transform is used to transform and output the image data of the first output layer LO1 into the image data of the second output layer LO2. Based on the same principle, the Winograd convolution transform is also used between the second output layer LO2 and the third output layer LO3, between the third output layer LO3 and the fourth output layer LO4, and between the fourth output layer LO4 and the fifth output layer LO5 to perform the transformation and processing of image data.
[0061] As Figure 1 shown, the image data of the input layer Lin is a kind of input data with ten rows (row line “0”, “1”, “2”, “3”, “4”, “5”, “6”, “7”, “8”, “9”). Among them, each row of data is sequentially transmitted to the input layer Lin and received by the input layer Lin. For example, in Figure 1 , the first row “0” of the input data is first transmitted to the input layer Lin and received by the input layer Lin.
[0062] Next, please refer to Figure 2 shown, in Figure 1After the first row "0" of the input data is received, the second row "1" of the input data is then transmitted to and received by the input layer Lin. After the first row "0" and the second row "1" of the input data are received, the present invention provides a first padding row X1 and a second padding row X2 to facilitate the Winograd convolution transform between the input layer Lin and the first output layer LO1, thereby generating a first output layer sequence. According to an embodiment of the present invention, the first padding row X1 or the second padding row X2 is, for example, optionally used to provide zero data; or, according to another embodiment of the present invention, the first padding row X1 or the second padding row X2 may also be, for example, used to provide the same image data as the first row of the input data, that is, the row data of the first row "0" of the input data. Therefore, based on the first padding row X1, the second padding row X2, the first row "0" and the second row "1" of the input data, the Winograd convolution transform can be performed between the input layer Lin and the first output layer LO1, thereby generating a first output layer sequence. As Figure 2 shown, the first output layer sequence generated at this time is the first row "0" of the first output layer LO1, which is shown by a field filled with gray in Figure 2 .
[0063] Next, please refer to Figure 3 shown. At this time, the third row "2" of the input data is then transmitted to and received by the input layer Lin. After that, in Figure 4 , the fourth row "3" of the input data is then transmitted to and received by the input layer Lin. Therefore, based on the first row "0", the second row "1", the third row "2" and the fourth row "3" of the input data, the Winograd convolution transform can be performed between the input layer Lin and the first output layer LO1, thereby generating a first output layer sequence. As Figure 4 shown, the first output layer sequence generated at this time is the second row "1" and the third row "2" of the first output layer LO1, which is also shown by a field filled with gray in Figure 4 .
[0064] As can be seen in the figure, in order to then perform the Winograd convolution transform between the first output layer LO1 and the second output layer LO2, the present invention successively provides a third padding row X3 in the first output layer LO1. According to an embodiment of the present invention, the third padding row X3 is, for example, optionally used to provide zero data or to provide the same image data as the first row of the first output layer sequence, that is, the image data of the first row "0" in the first output layer LO1.
[0065] Therefore, based on the third padding row X3, and the first row "0", second row "1", and third row "2" of the first output layer sequence in the first output layer LO1, the aforementioned Winograd convolution transform can be performed between the first output layer LO1 and the second output layer LO2, thereby generating a second output layer sequence. As Figure 4 shown, the second output layer sequence generated at this time is the first row "0" and second row "1" of the second output layer LO2, which are shown in the Figure 4 with the fields filled in gray in the same way.
[0066] Based on the same calculation rules, in order to perform the Winograd convolution transform between the second output layer LO2 and the third output layer LO3, the present invention successively provides a fourth padding row X4 and a fifth padding row X5 in the second output layer LO2. According to an embodiment of the present invention, the fourth padding row X4 and the fifth padding row X5 are, for example, optionally used to provide zero data, or to provide the same image data as the first row of the second output layer sequence, that is, the image data of the first row "0" in the second output layer LO2.
[0067] Therefore, based on the fourth padding row X4 and the fifth padding row X5, and the first row "0" and second row "1" of the second output layer sequence in the second output layer LO2, the aforementioned Winograd convolution transform can be performed between the second output layer LO2 and the third output layer LO3, thereby generating a third output layer sequence. As Figure 4 shown, the third output layer sequence generated at this time is the first row "0" of the third output layer LO3, which is shown in the Figure 4 with the fields filled in gray in the same way.
[0068] After that, please refer to Figure 5 shown, which is publicly available in Figure 4 afterwards, a schematic diagram of when the fifth row "4" of the input data is then transmitted to the input layer Lin and received by the input layer Lin. Figure 6 For Figure 5 afterwards, a schematic diagram of when the sixth row "5" of the input data is then transmitted to the input layer Lin and received by the input layer Lin. As can be seen in the figure, based on the third row "2", fourth row "3", fifth row "4", and sixth row "5" of the input data, the aforementioned Winograd convolution transform can be performed between the input layer Lin and the first output layer LO1, thereby generating a first output layer sequence. According to Figure 6As shown, the first output layer sequence generated at this time is the fourth row "3" and the fifth row "4" of the first output layer LO1. Then, based on the second row "1", the third row "2", the fourth row "3", and the fifth row "4" of the first output layer sequence in the first output layer LO1, the aforementioned Winograd convolution transform can be performed between the first output layer LO1 and the second output layer LO2, thereby generating a second output layer sequence. As Figure 6 shown, the second output layer sequence generated at this time is the third row "2" and the fourth row "3" of the second output layer LO2. Therefore, based on the same calculation principle, according to the first row "0", the second row "1", the third row "2", and the fourth row "3" of the second output layer sequence in the second output layer LO2, the aforementioned Winograd convolution transform can be performed between the second output layer LO2 and the third output layer LO3, thereby generating a third output layer sequence. As Figure 6 shown, the third output layer sequence generated at this time is the second row "1" and the third row "2" in the third output layer LO3, which is shown in the fields filled with gray in Figure 6 .
[0069] As can be seen in the figure, in order to then perform the Winograd convolution transform between the third output layer LO3 and the fourth output layer LO4, the present invention continuously provides a sixth padding row X6 in the third output layer LO3. According to an embodiment of the present invention, the sixth padding row X6 is, for example, optionally used to provide zero data, or to provide image data identical to the first row of the third output layer sequence, that is, the image data of the first row "0" in the third output layer LO3.
[0070] Therefore, based on the sixth padding row X6, and the first row "0", the second row "1", and the third row "2" of the third output layer sequence in the third output layer LO3, the aforementioned Winograd convolution transform can be performed between the third output layer LO3 and the fourth output layer LO4, thereby generating a fourth output layer sequence. As Figure 6 shown, the fourth output layer sequence generated at this time is the first row "0" and the second row "1" of the fourth output layer LO4.
[0071] Based on the same calculation rule, in order to perform the Winograd convolution transform between the fourth output layer LO4 and the fifth output layer LO5, the present invention continuously provides a seventh padding row X7 and an eighth padding row X8 in the fourth output layer LO4. According to an embodiment of the present invention, the seventh padding row X7 and the eighth padding row X8 are, for example, optionally used to provide zero data, or to provide image data identical to the first row of the fourth output layer sequence, that is, the image data of the first row "0" in the fourth output layer LO4.
[0072] In view of this, based on the seventh padding row X7 and the eighth padding row X8, and the first row "0" and the second row "1" of the fourth output layer sequence in the fourth output layer LO4, the Winograd convolution transform can be performed between the fourth output layer LO4 and the fifth output layer LO5, thereby generating the fifth output layer sequence. As Figure 6 shown, the fifth output layer sequence generated at this time is the first row "0" of the fifth output layer LO5. Therefore, by adopting the algorithm disclosed in the present invention, it can be clearly known that the present invention can successfully achieve the inventive effect of low latency (=5) in a continuous convolution structure.
[0073] After that, please then refer to Figure 7 shown, which is disclosed in Figure 6 After that, a schematic diagram of the seventh row "6" of the input data being then transmitted to the input layer Lin and received by the input layer Lin. Figure 8 For Figure 7 After that, a schematic diagram of the eighth row "7" of the input data being then transmitted to the input layer Lin and received by the input layer Lin. As can be seen from the figure, based on the fifth row "4", the sixth row "5", the seventh row "6", and the eighth row "7" of the input data, the Winograd convolution transform can be performed between the input layer Lin and the first output layer LO1, thereby generating the sixth row "5" and the seventh row "6" of the first output layer sequence. And, according to the same algorithm, the Winograd convolution transform can also be sequentially performed between the first output layer LO1 and the second output layer LO2, between the second output layer LO2 and the third output layer LO3, between the third output layer LO3 and the fourth output layer LO4, and between the fourth output layer LO4 and the fifth output layer LO5, so that the second output layer sequence, the third output layer sequence, the fourth output layer sequence, and the fifth output layer sequence are also sequentially generated. As Figure 8 shown, for better and clearer identification, in the sequence generated by each layer, the fields filled with gray represent the generated row data, including: the sixth row "5" and the seventh row "6" in the first output layer LO1, the fifth row "4" and the sixth row "5" in the second output layer LO2, the fourth row "3" and the fifth row "4" in the third output layer LO3, the third row "2" and the fourth row "3" in the fourth output layer LO4, and the second row "1" and the third row "2" in the fifth output layer LO5.
[0074] After that, please then continue to refer to Figure 9 shown, which is disclosed in Figure 8 After that, a schematic diagram of the ninth row "8" of the input data being then transmitted to the input layer Lin and received by the input layer Lin. Figure 10 For Figure 9Afterwards, the schematic diagram of the tenth line "9" of the input data being then transmitted to and received by the input layer Lin is shown. As can be seen in the figure, according to the seventh line "6", eighth line "7", ninth line "8", and tenth line "9" of the input data, the aforementioned Winograd convolution transformation can be performed between the input layer Lin and the first output layer LO1, thereby generating the eighth line "7" and ninth line "8" of the first output layer sequence. And, according to the same algorithm, the Winograd convolution transformation can also be sequentially performed between the first output layer LO1 and the second output layer LO2, between the second output layer LO2 and the third output layer LO3, between the third output layer LO3 and the fourth output layer LO4, and between the fourth output layer LO4 and the fifth output layer LO5, such that the second output layer sequence, third output layer sequence, fourth output layer sequence, and fifth output layer sequence are also sequentially generated. As Figure 10 shown, in order to preferably provide clearer recognition, in the sequence generated by each layer, the fields filled with gray also represent the generated row data, including: the eighth line "7" and ninth line "8" in the first output layer LO1, the seventh line "6" and eighth line "7" in the second output layer LO2, the sixth line "5" and seventh line "6" in the third output layer LO3, the fifth line "4" and sixth line "5" in the fourth output layer LO4, and the fourth line "3" and fifth line "4" in the fifth output layer LO5.
[0075] Therefore, in view of the above technical solutions disclosed in the present invention, it is quite observable that the control method of the continuous convolution structure disclosed in the present invention, by adopting the row-interleaved Winograd convolution transformation method, thereby improves and optimizes the algorithm. The aforementioned algorithm is based on depth-first and combines the row-interleaved algorithm, enabling the inventive effect of minimizing the delay to be achieved in the convolution structure. At the same time, according to the architecture disclosed in the drawings of the present invention Figures 1 to 10 it can be seen that in each convolution data layer, only a buffer of four rows is needed at a time to temporarily store the image data transformation. Based on this technical feature, the present invention can further significantly save the area required by traditional SRAM.
[0076] Looking further, based on the technical solutions disclosed in the drawings of the present invention Figures 1 to 10 In summary, when generating the output layer sequences of the odd output layers (such as: the first output layer LO1, the third output layer LO3, the fifth output layer LO5, etc.), the control method of the continuous convolution structure disclosed in the present invention includes as shown in the drawings Figure 11The steps shown: S1102, S1104, S1106, S1108, S1110, S1112, S1114, and S1116. In the following related description, in order to elaborate on the calculation method for obtaining the output layer sequence of the odd output layer, the present invention briefly takes the input layer Lin, the first output layer LO1, and the Winograd convolution transform performed between the input layer Lin and the first output layer LO1 as an illustrative example for explanation. At the same time, please refer to Figure 12 , Figure 13 and Figure 14 shown, which respectively correspond to Figure 11 the step description of the process shown in
[0077] For the first operation interval of the control method, first, please refer to Figure 12 and Figure 11 In steps S1102, S1104, and S1106 shown, in step S1102 of the present invention, first, the first row and the second row of an input data are received; then, in step S1104, the first padding row X1 and the second padding row X2 are provided. As described above, the provided first padding row X1 and / or the second padding row X2 are optionally used to provide zero data, or to provide the same image data as the first row of the input data. Therefore, in step S1106, the present invention can generate and output the first row of an output data through the Winograd convolution transform.
[0078] After that, as shown in steps S1108 and S1110 in Figure 11 , the present invention continues with the second operation interval. For the second operation interval of the control method, please also refer to Figure 13 shown, as shown in Figure 13 shown, the present invention then receives the third row and the fourth row of the input data in step S1108. Therefore, in step S1110, the present invention can generate and output the second row and the third row of the output data through the Winograd convolution transform.
[0079] After that, as shown in steps S1112, S1114, and S1116 in Figure 11 , the present invention continues with the third operation interval. For the third operation interval of the control method, please also refer to Figure 14 shown, as shown in Figure 14 shown, the present invention discards or overwrites the first row and the second row of the aforementioned input data in step S1112. After that, in step S1114, the fifth row and the sixth row of the input data are received, so that in step S1116, the present invention can generate and output the fourth row and the fifth row of the output data through the Winograd convolution transform.
[0080] Among them, according to an embodiment of the present invention, when performing the steps S1102, S1108, and S1114 to receive input data, the first row, second row, third row, fourth row, fifth row, and sixth row of the input data can be sequentially temporarily stored in a buffer in the executed image processing flow. In view of this, based on the fact that in each data convolution layer, only a four-line buffer size is required at a time to temporarily store the row data of the input data, it is obvious that by adopting the technical content disclosed in the present invention, the size and requirements of using a traditional SRAM memory can be effectively reduced.
[0081] On the other hand, according to the control method of the continuous convolution structure disclosed in the present invention, on the other hand, when generating the output layer sequence of even output layers (for example: the second output layer LO2, the fourth output layer LO4, the sixth output layer LO6, etc.), the control method of the continuous convolution structure disclosed in the present invention includes the steps shown in the accompanying drawings Figure 15 as follows: S1502, S1504, S1506, S1508, S1510, S1512, S1514, S1516, and S1518. In the following related descriptions, in order to illustrate the calculation method for obtaining the output layer sequence of the even output layer, the present invention briefly takes the first output layer LO1, the second output layer LO2, and the Winograd convolution transform performed between the first output layer LO1 and the second output layer LO2 as an illustrative example for explanation. At the same time, please refer to Figure 16 , Figure 17 and Figure 18 shown, which respectively correspond to Figure 15 the step descriptions of the processes shown in
[0082] Among them, in the embodiment shown in the accompanying drawings Figures 16 to 18 the first output layer LO1 is a convolution layer that provides input data, and the second output layer LO2 is a convolution layer that generates output data.
[0083] For the first operation interval of this control method, first, please refer to Figure 16 and Figure 15 the steps S1502, S1504, and S1506 shown in Figure 16As shown, a third padding line X3 is provided. As described above, the provided third padding line X3 is optionally used to provide zero data or provide the same image data as the first line of the input data. Therefore, in step S1506, the present invention can generate and output the first and second lines of output data through the Winograd convolution transform performed between the first output layer LO1 and the second output layer LO2.
[0084] After that, as Figure 15 shown in steps S1508, S1510, and S1512 of Figure 17 , it can be seen that before receiving the fourth and fifth lines of the input data in step S1510, the present invention screens out or overwrites the first line of the aforementioned input data in step S1508. After screening out or overwriting the first line of the input data, the fourth and fifth lines of the input data can be received (step S1510). Subsequently, in step S1512, the present invention can generate and output the third and fourth lines of the output data through the Winograd convolution transform between the first output layer LO1 and the second output layer LO2.
[0085] After that, according to Figure 15 shown in steps S1514, S1516, and S1518 of Figure 18 , the present invention continues with the third operation interval. For the third operation interval of this control method, please also cooperate with Figure 18 shown. First, the present invention screens out or overwrites the second and third lines of the aforementioned input data in step S1514. After that, in step S1516, the sixth and seventh lines of the input data are received, so that in step S1518, the present invention can generate and output the fifth and sixth lines of the output data through the Winograd convolution transform adopted between the first output layer LO1 and the second output layer LO2.
[0086] Among them, according to an embodiment of the present invention, when performing the steps S1502, S1510, and step S1516 to receive the input data, the first, second, third, fourth, fifth, sixth, and seventh lines of the input data can be temporarily stored in a buffer in sequence in the executed image processing flow. In view of this, as described above, based on the fact that in each data convolution layer, a buffer size of only a four-line buffer is required at a time to temporarily store the row data of the input data, it is obvious that by adopting the technical content disclosed by the present invention, the area size and its requirements of the traditional SRAM memory can be significantly reduced.
[0087] In addition, the inventors of the present application further provided a number of data to confirm and verify that the technical solution disclosed in the present invention is effective. Please refer to Table (I) provided below, which provides a data comparison of the differences between the prior art and the control method disclosed in the present invention when they are respectively applied to a convolutional neural network architecture for image processing calculations.
[0088]
[0089] Table (I)
[0090] It can be clearly seen from the above Table (I) that the technical means disclosed in the present invention, by adopting the Winograd convolution transform of F(2*2, 3*3), can successfully replace the traditional 3*3 convolution transform architecture adopted in the prior art, and moreover, it can also significantly reduce the usage area of the traditional multiply-accumulate arithmetic unit (by about 56%). In addition, the present invention can still maintain quite good convolution transform results without increasing the calculation delay. Thus, it is obvious that the row interleaving algorithm disclosed by the inventors in the present application is effectively applicable to the control method of the continuous convolution structure and can optimize the traditional convolution algorithm.
[0091] In summary, according to the method disclosed in the present invention, taking the application to a convolution architecture with a convolution kernel of (3*3) as an example for illustration, however, it is worth reminding that the present invention is not limited to this convolution architecture. That is to say, other convolution architectures of optional sizes, such as (n*n), can equally apply the method disclosed in the present invention, which is compatible and falls within the scope of the patent protection of the present invention.
[0092] Generally speaking, for those skilled in the art with ordinary knowledge, they can make equivalent modifications or variations based on the present invention without departing from the technical core of the present invention according to different specifications and / or norms. That is to say, the scope of protection claimed by the present invention is not limited to the above. And various modifications or variations and / or circuit implementation manners based on the present invention should still fall within the scope of the claims of the present invention.
[0093] As can be seen from the above embodiments, the present invention provides a Winograd convolution transform with a row interleaving algorithm. By adopting the algorithm of the present invention, the usage area of the multiplier-accumulator that must be used in the traditional art can be effectively saved. At the same time, the overall power consumption required for its circuit layout is also suppressed and reduced. Therefore, by adopting the embodiments and algorithms disclosed in the present invention, compared with the prior art, it can obviously and effectively solve many deficiencies existing in the prior art, and present a more efficient optimization invention effect, and can be widely used in related industries, and successfully overcome many long-existing defects in the prior art. Therefore, it is obvious that the technical solution requested by the applicant in this case indeed has excellent industrial applicability and competitiveness. At the same time, the technical features, methods and means disclosed in the present invention and the achieved effects are significantly different from the current solutions, and it is not easily completed by those skilled in the art, and should have patent requirements.
[0094] The above-described embodiments are only for illustrating the technical ideas and features of the present invention, and their purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. It should not be used to limit the patent scope of the present invention. That is, all equivalent changes or modifications made in accordance with the spirit disclosed in the present invention should still be covered within the scope protected by the claims of the present invention.
Claims
1. A control method for a continuous convolution structure, comprising the following steps: In a first operation interval, receiving a first row and a second row of input data and performing a Winograd convolution transform to generate a first row of output data; In a second operation interval, receiving the third and fourth rows of the input data and performing a Winograd convolution transform to generate the second and third rows of the output data; as well as In a third operation interval, the fifth and sixth rows of the input data are received, and Winograd convolution transformation is performed to generate the fourth and fifth rows of the output data.
2. The control method of the continuous convolution structure according to claim 1, wherein: The step of receiving the first row and the second row of the input data and performing a Winograd convolution transform further includes providing a first padding row and a second padding row to generate the first row of the output data through the Winograd convolution transform.
3. The control method of the continuous convolution structure according to claim 2, wherein: The first filling line or the second filling line selectively provides zero point data, or provides image data that is the same as the first line of the input data.
4. The control method of the continuous convolution structure according to claim 1, wherein: Before receiving the fifth row and the sixth row of the input data, the method further includes deleting or covering the first row and the second row of the input data.
5. The control method of the continuous convolution structure according to claim 1, wherein: The first row, the second row, the third row, the fourth row, the fifth row, and the sixth row of the input data may be temporarily stored in a buffer.
6. A control method for a continuous convolution structure, comprising the following steps: In a first operation interval, receiving a first row, a second row, and a third row of input data, and performing a Winograd convolution transform to generate a first row and a second row of output data; In a second operation interval, receiving the fourth and fifth rows of the input data and performing a Winograd convolution transform to generate the third and fourth rows of the output data; as well as In a third operation interval, the sixth and seventh rows of the input data are received, and Winograd convolution transformation is performed to generate the fifth and sixth rows of the output data.
7. The control method of the continuous convolution structure according to claim 6, wherein: The step of receiving the first row, the second row and the third row of the input data and performing a Winograd convolution transform further includes providing a third padding row to generate the first row and the second row of the output data through the Winograd convolution transform.
8. The control method of the continuous convolution structure according to claim 7, wherein: The third fill line selectively provides zero point data, or provides image data that is the same as the first line of the input data.
9. The control method of the continuous convolution structure according to claim 6, wherein: Before receiving the fourth row and the fifth row of the input data, the method further includes deleting or covering the first row of the input data.
10. The control method of the continuous convolution structure according to claim 6, wherein: Before receiving the sixth row and the seventh row of the input data, the method further includes deleting or covering the second row and the third row of the input data.
11. The control method of the continuous convolution structure according to claim 6, wherein: The first row, the second row, the third row, the fourth row, the fifth row, the sixth row, and the seventh row of the input data may be temporarily stored in a buffer.