Image Scaling Device, Method, Equipment and Storage Medium Based on Data Flow Architecture

By adopting an image scaling device based on data flow architecture on the AI ​​chip, pre-caches and directly uses the interpolation coefficients of the image scaling algorithm, the impact of traditional image scaling operations on the performance and power consumption of the AI ​​chip is solved, and more efficient image scaling operations are achieved.

CN114820313BActive Publication Date: 2025-06-13SHENZHEN CORERAIN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210446502.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-26
Publication Date
2025-06-13
Estimated Expiration
2042-04-26

AI Technical Summary

Technical Problem

Traditional image scaling operations require real-time calculation of interpolation coefficients, resulting in an impact on AI chip performance and power consumption.

Method used

Using an image scaling device based on the data flow architecture, the interpolation coefficients required by the image scaling algorithm are calculated offline in advance, and combined with the programmable address generation module and the cache module, the target feature data and interpolation coefficient are directly read and output for calculation.

Benefits of technology

The process of calculating interpolation coefficients in real time is avoided, and the process of AI chip realizing image scaling operations is simplified, thereby reducing the impact on AI chip performance and power consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114820313B_ABST
    Figure CN114820313B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses an image scaling device, method, device and storage medium based on a data flow architecture. The device includes: a first programmable address generation module, a second programmable address generation module, a first on-chip cache module, a second on-chip cache module, and a calculation module. By using the second on-chip cache module to pre-cache the required interpolation coefficients obtained by offline calculation, and then using two programmable address generation modules to sequentially generate read addresses for the two on-chip cache modules based on their respective configurations according to the image scaling algorithm, the corresponding target feature data and target interpolation coefficients can be directly read from the first on-chip cache module and the second on-chip cache module respectively, and output to the calculation module to calculate each feature data of the output feature map. The process of real-time calculating interpolation coefficients is avoided, the process of implementing image scaling operation by the AI chip is simplified, and thus the impact on the performance and power consumption of the AI chip is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the technical field of data processing, and in particular, to an image scaling device, method, equipment and storage medium based on a data stream architecture. Background Art

[0002] With the rapid development of deep learning, neural network algorithms have been widely applied to machine vision applications, such as image recognition and image classification. Aiming at the problems of complex neural network algorithms, large computational volume and long inference running time, AI chips have been customized to accelerate the operation of neural network algorithms. The image scaling (resize) operation is a common operation in neural network algorithms, which scales the feature map to a specified size. Common image scaling methods include bilinear interpolation and nearest neighbor interpolation (Nearest Neighbor, NN), etc. Among them, the existing scaling process using the bilinear interpolation method is as follows: the boundary points of the output feature map are overlapped with the boundary points of the input feature map, and then all the remaining points of the output feature map are evenly placed in the area determined by the boundary points. In this way, some output points are inserted at equal intervals between every two points of each input feature map according to the scaling ratio factor. The size of each point of the output feature map is only related to the four input feature points adjacent to it on the input feature map. By calculating the relevant interpolation coefficients, the values of the output feature points can be calculated.

[0003] The traditional method usually uses the CPU or controls the arithmetic logic unit (ALU) through instructions to perform operations to implement the bilinear resize operation. The implementation method is relatively simple. It directly uses an ALU for calculation and stores the parameters, instructions and data in the storage unit. The implementation process is as follows: the main control unit starts to execute the image scaling operation, the ALU reads the index of the current calculated feature point on the feature map and the scaling ratio factor (scale) from the storage unit, the ALU calculates four bilinear coefficients according to the index and scale, the ALU stores the four calculated bilinear coefficients back into the storage unit, the ALU reads the data and bilinear coefficients from the storage unit, the ALU uses the data and bilinear coefficients to complete the calculation, the ALU stores the calculated result back into the storage unit, and repeats the above process to complete the resize operation for all feature points on the feature map.

[0004] It can be seen that the traditional resize operation needs to call a floating-point multiplier to calculate the bilinear interpolation coefficients in real time, and also needs to read and write the storage unit repeatedly, thus affecting the performance and power consumption of the AI chip. Similarly, other image scaling methods such as NN also have these problems. Summary of the Invention

[0005] An embodiment of the present invention provides an image scaling device, method, device and storage medium based on a data flow architecture to avoid the real-time calculation process of interpolation coefficients, simplify the image scaling operation process, and thus reduce the impact on the performance and power consumption of the AI chip.

[0006] In a first aspect, an embodiment of the present invention provides an image scaling device based on a data flow architecture. The device includes: a first programmable address generation module, a second programmable address generation module, a first on-chip cache module, a second on-chip cache module, and a calculation module; wherein,

[0007] The first programmable address generation module is used to sequentially generate a first read address of the first on-chip cache module based on a first external configuration and the used image scaling algorithm;

[0008] The second programmable address generation module is used to sequentially generate a second read address of the second on-chip cache module based on a second external configuration and the image scaling algorithm;

[0009] The first on-chip cache module is used to cache the feature data of the input feature map and sequentially read the target feature data according to each of the first read addresses;

[0010] The second on-chip cache module is used to cache the interpolation coefficients required for the image scaling algorithm obtained by pre-offline calculation and sequentially read the target interpolation coefficients according to each of the second read addresses;

[0011] The calculation module is used to sequentially calculate based on the target feature data and the corresponding target interpolation coefficients according to the image scaling algorithm to obtain each feature data of the output feature map.

[0012] Optionally, the first programmable address generation module includes an externally configurable first register for storing the first external configuration; the second programmable address generation module includes an externally configurable second register for storing the second external configuration.

[0013] Optionally, the calculation module is further used to output each feature data of the obtained output feature map to a result cache space for caching.

[0014] Optionally, the image scaling algorithm includes a nearest neighbor interpolation algorithm; correspondingly, the first on-chip cache module is further used to directly use the target feature data as an interpolation result and directly output the interpolation result to the result cache space for caching.

[0015] Optionally, the result cache space includes one of the first on-chip cache module and the external storage module.

[0016] In a second aspect, an embodiment of the present invention further provides an image scaling method based on a data flow architecture, the method including:

[0017] The first programmable address generation module sequentially generates a first read address of the first on-chip cache module for each feature point of the output feature map based on the used image scaling algorithm, and the second programmable address generation module sequentially generates a second read address of the second on-chip cache module for each feature point of the output feature map based on the image scaling algorithm; wherein, the first on-chip cache module is used to cache the feature data of the input feature map, and the second on-chip cache module is used to cache the interpolation coefficients required by the image scaling algorithm calculated offline in advance;

[0018] The first on-chip cache module sequentially reads and outputs target feature data according to each of the first read addresses, and the second on-chip cache module sequentially reads and outputs target interpolation coefficients according to each of the second read addresses;

[0019] The calculation module receives the target feature data and the target interpolation coefficients, and sequentially calculates based on the target feature data and the corresponding target interpolation coefficients according to the image scaling algorithm to obtain the respective feature data of the output feature map.

[0020] Optionally, before the first programmable address generation module sequentially generates a first read address of the first on-chip cache module for each feature point of the output feature map based on the used image scaling algorithm, and the second programmable address generation module sequentially generates a second read address of the second on-chip cache module for each feature point of the output feature map based on the image scaling algorithm, it further includes:

[0021] Obtain the first size of the input feature map to be scaled, the second size of the output feature map, and the used image scaling algorithm according to a preset neural network algorithm model;

[0022] Configure the first programmable address generation module and the second programmable address generation module according to the first size, the second size, and the calculation mode of the image scaling algorithm.

[0023] Optionally, before configuring the first programmable address generation module and the second programmable address generation module according to the first size, the second size, and the calculation mode of the image scaling algorithm, it further includes:

[0024] The interpolation coefficient is pre-calculated offline according to the first dimension and the second dimension and cached in the second on-chip cache module.

[0025] In a third aspect, an embodiment of the present invention further provides a computer device, which includes:

[0026] One or more processors;

[0027] A memory for storing one or more programs;

[0028] When the one or more programs are executed by the one or more processors, the one or more processors implement the image scaling method based on the data flow architecture provided in any embodiment of the present invention.

[0029] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the image scaling method based on the data flow architecture provided in any embodiment of the present invention.

[0030] An embodiment of the present invention provides an image scaling device based on a data flow architecture, including a first programmable address generation module, a second programmable address generation module, a first on-chip cache module, a second on-chip cache module, and a calculation module. By using the second on-chip cache module to pre-cache the interpolation coefficients required for the image scaling algorithm obtained by offline calculation, and using the first on-chip cache module to pre-cache the feature data of the input feature map, and then the first programmable address generation module and the second programmable address generation module generate the read addresses of the first on-chip cache module and the second on-chip cache module in sequence based on their respective configurations according to the used image scaling algorithm, so that the corresponding target feature data and target interpolation coefficients can be directly read from the first on-chip cache module and the second on-chip cache module in sequence and output to the calculation module to calculate each feature data of the output feature map. This avoids the process of real-time calculating interpolation coefficients, simplifies the process of the AI chip implementing the image scaling operation, and thus reduces the impact on the performance and power consumption of the AI chip. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 It is a schematic structural diagram of an image scaling device based on a data flow architecture provided in Embodiment 1 of the present invention;

[0032] Figure 2 It is a flowchart of an image scaling method based on a data flow architecture provided in Embodiment 2 of the present invention;

[0033] Figure 3 It is a schematic structural diagram of a computer device provided in Embodiment 3 of the present invention. DETAILED DESCRIPTION

[0034] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. In addition, it should be noted that for the sake of description, only parts related to the present invention rather than all structures are shown in the drawings.

[0035] Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the steps as sequential processes, many of the steps can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the steps can be rearranged. The process can be terminated when its operations are completed, but it can also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, and the like.

[0036] In addition, terms such as "first", "second", etc. may be used herein to describe various directions, actions, steps, or elements, etc., but these directions, actions, steps, or elements are not limited by these terms. These terms are only used to distinguish the first direction, action, step, or element from another direction, action, step, or element. For example, without departing from the scope of the present application, the first programmable address generation module can be referred to as the second programmable address generation module, and similarly, the second programmable address generation module can be referred to as the first programmable address generation module. Both the first programmable address generation module and the second programmable address generation module are programmable address generation modules, but they are not the same programmable address generation module. Terms such as "first", "second", etc. cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, features defined with "first", "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0037] Embodiment 1

[0038] Figure 1 FIG. is a schematic structural diagram of an image scaling device based on a data flow architecture provided in Embodiment 1 of the present invention. This embodiment is applicable to the case where a feature map needs to be scaled during the implementation of a neural network algorithm by an AI chip. As Figure 1As shown in the figure, the device includes: a first programmable address generation module 11, a second programmable address generation module 12, a first on-chip cache module 13, a second on-chip cache module 14, and a calculation module 15. Among them, the first programmable address generation module 11 is used to sequentially generate a first read address of the first on-chip cache module 13 based on a first external configuration and the used image scaling algorithm. The second programmable address generation module 12 is used to sequentially generate a second read address of the second on-chip cache module 14 based on a second external configuration and the image scaling algorithm. The first on-chip cache module 13 is used to cache the feature data of the input feature map and sequentially read the target feature data according to each of the first read addresses. The second on-chip cache module 14 is used to cache the interpolation coefficients required for the image scaling algorithm obtained by pre-offline calculation and sequentially read the target interpolation coefficients according to each of the second read addresses. The calculation module 15 is used to sequentially calculate based on the target feature data and the corresponding target interpolation coefficients according to the image scaling algorithm to obtain the respective feature data of the output feature map.

[0039] Specifically, the first programmable address generation module 11 and the second programmable address generation module 12 are general programmable address generation modules, which are address generation modules specifically designed for the on-chip cache module of the data flow architecture AI chip. When the AI chip needs to execute a specific image scaling algorithm, the first programmable address generation module 11 can receive the first external configuration and complete initialization to prepare to continuously generate the first read address of the first on-chip cache module 13 in the order of the calculation process of the used image scaling algorithm. At the same time, the second programmable address generation module 12 can receive the second external configuration and complete initialization to prepare to continuously generate the second read address of the second on-chip cache module 14 in the order of the calculation process of the used image scaling algorithm, and the first external configuration and the second external configuration can be the same.

[0040] The first on-chip cache module 13 caches the feature data of the input feature map. Specifically, it can be cached according to the data cache mode designed for the AI chip. At the same time, the first on-chip cache module 13 can also be used to cache the feature data of the finally calculated output feature map. The second on-chip cache module 14 caches the interpolation coefficients required for the image scaling algorithm pre-calculated offline. For some image scaling algorithms, each feature point of the output feature map can correspond to a set of interpolation coefficients. For example, when using the bilinear interpolation algorithm, each feature point can correspond to a set of four interpolation coefficients. Each set of interpolation coefficients is only related to the position information (i.e., coordinates) of the feature point and has nothing to do with the specific feature data. Therefore, after the sizes of the input feature map and the output feature map are given, the interpolation coefficients required for the calculation corresponding to each feature point of the output feature map can be directly calculated in advance. Specifically, it can be calculated offline in advance through the driver program of the AI chip and cached in the second on-chip cache module 14 according to the data cache mode designed for the AI chip. In addition, when the AI chip performs the convolution operation, the second on-chip cache module 14 can also be used to cache the pre-calculated convolution weight data.

[0041] After the first programmable address generation module 11 and the second programmable address generation module 12 are ready to start generating read data addresses, the data stream can drive the first programmable address generation module 11 and the second programmable address generation module 12, so that the first programmable address generation module 11 automatically calculates the corresponding first read address for each feature point of the output feature map according to the first external configuration, and the second programmable address generation module 12 automatically calculates the corresponding second read address for each feature point of the output feature map according to the second external configuration. The first programmable address generation module 11 and the second programmable address generation module 12 can continuously generate addresses in real time. At the same time, whether they generate new addresses is driven by the data stream, that is, if there is data cached in the first on-chip cache module 13 and the data can be output, the first programmable address generation module 11 generates the first read address for the data to be output. If there is data cached in the second on-chip cache module 14 and the data can be output, the second programmable address generation module 12 generates the second read address for the data to be output.

[0042] After the first programmable address generation module 11 generates the first read address, as long as the data corresponding to the first read address is cached in the first on-chip cache module 13 and the calculation module 15 can process new data, the first on-chip cache module 13 can immediately read the data from the first read address as the current target feature data and output it. Similarly, after the second programmable address generation module 12 generates the second read address, as long as the data corresponding to the second read address is cached in the second on-chip cache module 14 and the calculation module 15 can process new data, the second on-chip cache module 14 can immediately read the data from the second read address as the current target interpolation coefficient and output it.

[0043] The calculation module 15 can be the calculation module of the AI chip. When data continuously enters the calculation module 15, the calculation module 15 can complete the multiply-accumulate calculation on the data according to the calculation method specified by the used image scaling algorithm. At the same time, the calculation module 15 can also be used for the convolution operation calculation of the AI chip. Specifically, after the calculation module 15 receives the target feature data output by the first on-chip cache module 13 and the target interpolation coefficient output by the second on-chip cache module 14, it can complete the operation process according to the used image scaling algorithm. This process is carried out continuously, and finally the operation of the entire output feature map is completed.

[0044] For the AI chip with a data flow architecture, usually the entire image scaling operation runs continuously in a pipelined form, that is, the first programmable address generation module 11 and the second programmable address generation module 12 respectively generate N (positive integer) addresses continuously, and the first on-chip cache module 13 and the second on-chip cache module 14 immediately read and output the data after the corresponding first address is generated and received. The calculation module 15 immediately performs calculations after receiving the first group of data. In the whole process, each module is working simultaneously. After the first on-chip cache module 13 and the second on-chip cache module 14 read the data of the corresponding first address, they immediately read the data of the second address, and then sequentially read the data of the addresses received later. Similarly, the calculation module 15 also operates in a similar way.

[0045] On the basis of the above technical solution, optionally, the first programmable address generation module 11 includes an externally configurable first register, and the first register is used to store the first external configuration; the second programmable address generation module 12 includes an externally configurable second register, and the second register is used to store the second external configuration.

[0046] Specifically, a set of externally configurable registers can be provided in both the first programmable address generation module 11 and the second programmable address generation module 12. When the AI chip needs to execute a specific algorithm, the corresponding register configuration can be generated according to the calculation mode of the algorithm and the data caching mode of the on-chip cache module of the AI chip, and configured to the registers in the first programmable address generation module 11 and the second programmable address generation module 12. Then, the first programmable address generation module 11 and the second programmable address generation module 12 can continuously generate data read addresses according to their respective configurations. Specifically, during the image scaling operation, the first size of the input feature map to be scaled and the second size of the output feature map, as well as the image scaling algorithm used, can be obtained first according to the preset neural network algorithm model. Then, the first external configuration of the first programmable address generation module 11 and the second external configuration of the second programmable address generation module 12 can be generated according to the first size, the second size, and the calculation mode of the image scaling algorithm. Then, the first programmable address generation module 11 and the second programmable address generation module 12 can be configured through the driver program of the AI chip, specifically, the first register and the second register can be configured. Among them, for the AI chip to work, the user needs to provide the network model to be run, which is usually stored in the host computer (PC or server). The user program running in the host computer transmits the network model to the driver program of the AI chip through PCIE or other buses, and the driver program can obtain parameters such as the input and output map sizes from the network model.

[0047] Furthermore, the first programmable address generation module 11 and the second programmable address generation module 12 can be used to implement the basic function in the following form:

[0048]

[0049] Among them, y represents the generated read address, floor represents rounding down, x represents the index of the currently calculated feature point on the output feature map, and A, B, C, D, E, and T are all register configuration parameters (integer parameters and can be negative). The functions completed by the first programmable address generation module 11 and the second programmable address generation module 12 are combinations of the above basic function functions at multiple layers. By configuring the first register and the second register, various image scaling operations in various software frameworks can be implemented based on the above basic function functions. Exemplarily, when applying the bilinear interpolation algorithm, if the width of the input feature map is 16, the width of the output feature map is 32, the channel size is 64, and the align_corner attribute of the algorithm is true, the register configuration for width direction addressing can be generated as follows: A = 0, B = 16 - 1 = 15, C = 0, D = 32 - 1 = 31, E = 0, T = 64. Then, a similar register configuration can be performed for the height direction addressing process. After completing the register configuration of the first programmable address generation module 11 and the second programmable address generation module 12, the first programmable address generation module 11 and the second programmable address generation module 12 can be initialized to prepare for starting to generate read data addresses.

[0050] Based on the above technical solution, optionally, the calculation module 15 is further configured to output each feature data of the obtained output feature map to a result cache space for caching. Further optionally, the result cache space includes one of the first on-chip cache module 13 and the external storage module. Specifically, the calculation module 15 can output the obtained calculation result to the external storage module DDR or the on-chip storage unit of the AI chip (specifically, it can be the first on-chip cache module 13) after each calculation process is completed for subsequent use.

[0051] Based on the above technical solutions, optionally, the image scaling algorithm includes the nearest neighbor interpolation algorithm; correspondingly, the first on-chip cache module 13 is further configured to directly use the target feature data as the interpolation result and directly output the interpolation result to the result cache space for caching. Specifically, for the image scaling operation of the nearest neighbor interpolation algorithm, the above device structure and configuration method can also be adopted. However, since the nearest neighbor interpolation algorithm directly takes the nearest data around as the interpolation result and does not require calculation, when the device provided in this embodiment executes the nearest neighbor interpolation algorithm, the calculation module 15 does not work, and the second programmable address generation module 12 and the second on-chip cache module 14 may also not work. After the first programmable address generation module 11 adopts a register configuration similar to the above, it continuously generates a first read address to the first on-chip cache module 13. The first on-chip cache module 13 can directly use the first read address to read data as the interpolation result and does not output it to the calculation module 15, but directly stores it from the internal storage of the first on-chip cache module 13 to the result cache space. Optionally, the result cache space includes one of the first on-chip cache module 13 and the external storage module. Specifically, if the nearest neighbor interpolation algorithm operation needs to be combined with other calculation operations for simultaneous processing, the first on-chip cache module 13 can also output the read data to the calculation module 15 for processing other calculation operations. Thus, different image scaling modes of various software frameworks can be realized, improving the versatility of the AI chip.

[0052] The image scaling device based on the data flow architecture provided in the embodiment of the present invention includes a first programmable address generation module, a second programmable address generation module, a first on-chip cache module, a second on-chip cache module, and a calculation module. By using the second on-chip cache module to pre-cache the interpolation coefficients required for the image scaling algorithm obtained by offline calculation, and using the first on-chip cache module to pre-cache the feature data of the input feature map, and then the first programmable address generation module and the second programmable address generation module sequentially generate the read addresses of the first on-chip cache module and the second on-chip cache module based on their respective configurations according to the used image scaling algorithm, so that the corresponding target feature data and target interpolation coefficients can be directly read from the first on-chip cache module and the second on-chip cache module respectively and output to the calculation module to calculate each feature data of the output feature map. The process of real-time calculating the interpolation coefficients is avoided, the process of the AI chip implementing the image scaling operation is simplified, and thus the influence on the performance and power consumption of the AI chip is reduced.

[0053] Embodiment 2

[0054] Figure 2The flowchart of the image scaling method based on the data flow architecture provided in the second embodiment of the present invention. This embodiment is applicable to the situation where feature maps need to be scaled during the implementation of neural network algorithms by an AI chip. This method can be applied to the image scaling device based on the data flow architecture provided in any embodiment of the present invention, and has the corresponding method flow and beneficial effects of the device. As Figure 2 shown, the specific steps are as follows:

[0055] S21. The first programmable address generation module sequentially generates the first read address of the first on-chip cache module for each feature point of the output feature map based on the used image scaling algorithm, and the second programmable address generation module sequentially generates the second read address of the second on-chip cache module for each feature point of the output feature map based on the image scaling algorithm; wherein, the first on-chip cache module is used to cache the feature data of the input feature map, and the second on-chip cache module is used to cache the interpolation coefficients required for the image scaling algorithm calculated offline in advance.

[0056] S22. The first on-chip cache module sequentially reads and outputs the target feature data according to each of the first read addresses, and the second on-chip cache module sequentially reads and outputs the target interpolation coefficients according to each of the second read addresses.

[0057] S23. The calculation module receives the target feature data and the target interpolation coefficients, and sequentially performs calculations based on the target feature data and the corresponding target interpolation coefficients according to the image scaling algorithm to obtain the respective feature data of the output feature map.

[0058] On the basis of the above technical solution, optionally, before the first programmable address generation module sequentially generates the first read address of the first on-chip cache module for each feature point of the output feature map based on the used image scaling algorithm, and the second programmable address generation module sequentially generates the second read address of the second on-chip cache module for each feature point of the output feature map based on the image scaling algorithm, it further includes: obtaining the first size of the input feature map to be scaled, the second size of the output feature map, and the used image scaling algorithm according to a preset neural network algorithm model; configuring the first programmable address generation module and the second programmable address generation module according to the first size, the second size, and the calculation mode of the image scaling algorithm.

[0059] Based on the above technical solution, optionally, before configuring the first programmable address generation module and the second programmable address generation module according to the first dimension, the second dimension, and the calculation mode of the image scaling algorithm, the method further includes: pre-offline calculating the interpolation coefficient according to the first dimension and the second dimension, and caching the interpolation coefficient into the second on-chip cache module.

[0060] Specifically, for relevant content, reference can be made to the description of the above embodiments, and details will not be repeated here.

[0061] In the technical solution provided by the embodiment of the present invention, by using the second on-chip cache module to pre-cache the interpolation coefficient required by the used image scaling algorithm obtained through offline calculation, and using the first on-chip cache module to pre-cache the feature data of the input feature map, and then through the first programmable address generation module and the second programmable address generation module according to their respective configurations, sequentially generate the read addresses of the first on-chip cache module and the second on-chip cache module based on the used image scaling algorithm, so that the corresponding target feature data and target interpolation coefficient can be directly read from the first on-chip cache module and the second on-chip cache module respectively, and output to the calculation module to calculate each feature data of the output feature map. This avoids the process of real-time calculating the interpolation coefficient, simplifies the process of the AI chip implementing the image scaling operation, and thus reduces the impact on the performance and power consumption of the AI chip.

[0062] Embodiment III

[0063] Figure 3 FIG. 3 is a schematic structural diagram of a computer device provided in Embodiment III of the present invention, showing a block diagram of an exemplary computer device suitable for implementing the embodiment of the present invention. Figure 3 The displayed computer device is only an example, and should not bring any limitation to the functions and usage scope of the embodiments of the present invention. As Figure 3 shown, the computer device includes a processor 31, a memory 32, an input device 33, and an output device 34; the number of processors 31 in the computer device can be one or more, Figure 3 taking one processor 31 as an example, the processor 31, the memory 32, the input device 33, and the output device 34 in the computer device can be connected through a bus or other means, Figure 3 taking the connection through the bus as an example.

[0064] The memory 32, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the image scaling method based on the data flow architecture in the embodiments of the present invention. The processor 31 executes various functional applications and data processing of the computer device by running the software programs, instructions, and modules stored in the memory 32, that is, implements the above-mentioned image scaling method based on the data flow architecture.

[0065] The memory 32 may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the computer device, etc. In addition, the memory 32 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some instances, the memory 32 may further include a memory remotely set relative to the processor 31, and these remote memories can be connected to the computer device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0066] The input device 33 can be used to obtain a neural network algorithm model preset by a user, and generate key signal inputs related to user settings and function control of the computer device, etc. The output device 34 can be used to transmit calculation results to subsequent modules, etc.

[0067] Embodiment Four

[0068] Embodiment Four of the present invention further provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute an image scaling method based on a data flow architecture when executed by a computer processor. The method includes:

[0069] Sequentially generating, by a first programmable address generation module, a first read address of a first on-chip cache module for each feature point of an output feature map based on the used image scaling algorithm, and sequentially generating, by a second programmable address generation module, a second read address of a second on-chip cache module for each feature point of the output feature map based on the image scaling algorithm; wherein, the first on-chip cache module is used to cache the feature data of an input feature map, and the second on-chip cache module is used to cache the interpolation coefficients required for the image scaling algorithm calculated offline in advance;

[0070] Sequentially reading and outputting target feature data by the first on-chip cache module according to each of the first read addresses, and sequentially reading and outputting target interpolation coefficients by the second on-chip cache module according to each of the second read addresses;

[0071] The computing module receives the target feature data and the target interpolation coefficient, and sequentially calculates based on the target feature data and the corresponding target interpolation coefficient according to the image scaling algorithm to obtain each feature data of the output feature map.

[0072] A storage medium can be any of various types of memory devices or storage devices. The term "storage medium" is intended to include: installation media such as CD-ROMs, floppy disks, or magnetic tape devices; computer system memory or random access memory such as DRAM, DDR RAM, SRAM, EDO RAM, Rambus RAM, etc.; non-volatile memory such as flash memory, magnetic media (such as hard disks or optical storage); register or other similar types of memory elements, etc. The storage medium can also include other types of memory or combinations thereof. Additionally, the storage medium can be located in the computer system in which the program is executed, or can be located in a different second computer system that is connected to the computer system via a network (such as the Internet). The second computer system can provide program instructions to the computer for execution. The term "storage medium" can include two or more storage media that can reside in different locations (such as in different computer systems connected via a network). The storage medium can store program instructions (such as specifically implemented as a computer program) executable by one or more processors.

[0073] Of course, the computer-executable instructions of a storage medium provided by an embodiment of the present invention are not limited to the method operations as described above, and can also execute related operations in the image scaling method based on the data flow architecture provided by any embodiment of the present invention.

[0074] A computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0075] The program code contained on a computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination of the above.

[0076] Through the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software and the necessary general hardware. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as a floppy disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk, or optical disc of a computer, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present invention.

[0077] Note that the above is only the preferred embodiment of the present invention and the applied technical principle. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein. Various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, it can also include more other equivalent embodiments, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. An image scaling device based on a data flow architecture, characterized in that, it includes: a first programmable address generation module, a second programmable address generation module, a first on-chip cache module, a second on-chip cache module, and a calculation module; wherein, the first programmable address generation module is used to sequentially generate a first read address of the first on-chip cache module based on a first external configuration and the used image scaling algorithm; the second programmable address generation module is used to sequentially generate a second read address of the second on-chip cache module based on a second external configuration and the image scaling algorithm; the first on-chip cache module is used to cache the feature data of the input feature map and sequentially read target feature data according to each of the first read addresses; the second on-chip cache module is used to cache the interpolation coefficients required for the image scaling algorithm obtained by pre-offline calculation and sequentially read target interpolation coefficients according to each of the second read addresses; the calculation module is used to sequentially calculate based on the target feature data and the corresponding target interpolation coefficients according to the image scaling algorithm to obtain each feature data of the output feature map.

2. The image scaling device based on a data flow architecture according to claim 1, characterized in that, the first programmable address generation module includes an externally configurable first register for storing the first external configuration; the second programmable address generation module includes an externally configurable second register for storing the second external configuration.

3. The image scaling device based on a data flow architecture according to claim 1, characterized in that, the calculation module is further used to output each feature data of the obtained output feature map to a result cache space for caching.

4. The image scaling device based on a data flow architecture according to claim 1, characterized in that, the image scaling algorithm includes a nearest neighbor interpolation algorithm; correspondingly, the first on-chip cache module is further used to directly use the target feature data as an interpolation result and directly output the interpolation result to the result cache space for caching.

5. The image scaling device based on a data flow architecture according to claim 3 or 4, characterized in that, the result cache space includes one of the first on-chip cache module and an external storage module.

6. An image scaling method based on a data flow architecture, characterized in that, it includes: sequentially generating a first read address of the first on-chip cache module for each feature point of the output feature map by a first programmable address generation module based on the used image scaling algorithm, and sequentially generating a second read address of the second on-chip cache module for each feature point of the output feature map by a second programmable address generation module based on the image scaling algorithm; wherein, the first on-chip cache module is used to cache the feature data of the input feature map, and the second on-chip cache module is used to cache the interpolation coefficients required for the image scaling algorithm obtained by pre-offline calculation; The target feature data is sequentially read and output by the first on-chip cache module according to each of the first read addresses, and the target interpolation coefficients are sequentially read and output by the second on-chip cache module according to each of the second read addresses; The calculation module receives the target feature data and the target interpolation coefficients, and sequentially performs calculations based on the target feature data and the corresponding target interpolation coefficients according to the image scaling algorithm to obtain the respective feature data of the output feature map.

7. The image scaling method based on a data flow architecture according to claim 6, wherein, Before the first programmable address generation module sequentially generates the first read addresses of the first on-chip cache module for each feature point of the output feature map based on the used image scaling algorithm, and the second programmable address generation module sequentially generates the second read addresses of the second on-chip cache module for each feature point of the output feature map based on the image scaling algorithm, it further includes: Obtaining the first size of the input feature map to be scaled, the second size of the output feature map, and the used image scaling algorithm according to a preset neural network algorithm model; Configuring the first programmable address generation module and the second programmable address generation module according to the first size, the second size, and the calculation mode of the image scaling algorithm.

8. The image scaling method based on a data flow architecture according to claim 7, wherein, Before configuring the first programmable address generation module and the second programmable address generation module according to the first size, the second size, and the calculation mode of the image scaling algorithm, it further includes: Pre-calculating the interpolation coefficients offline according to the first size and the second size and caching them in the second on-chip cache module.

9. A computer device, wherein, comprising: One or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the image scaling method based on a data flow architecture according to any one of claims 6-8.

10. A computer-readable storage medium having a computer program stored thereon, wherein, When the program is executed by a processor, it implements the image scaling method based on a data flow architecture according to any one of claims 6-8.

Citation Information

Patent Citations

  • High-parallelism and low-delay image scaling and clipping processing method based on zynq platform

    CN112017107A

  • Data processing method and device, electronic equipment and medium

    CN112631955A