Accelerator determination method, image processing method, equipment and storage medium

By adopting an accelerator determination method based on FPGA in the image compression algorithm, and using the fusion factor list and parameter mapping optimization hardware implementation, the problems of long image compression time and high GPU power consumption are solved, and low power consumption and high efficiency image compression are achieved.

CN120088345APending Publication Date: 2025-06-03UNIV OF SCI & TECH OF CHINA
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510084648.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

In deep learning-based image compression algorithms, high-resolution images and a large number of parameters lead to long image compression times, and the high power consumption of the GPU is not suitable for power-sensitive end-side devices.

Method used

The accelerator determination method based on field programmable gate array (FPGA) is adopted to optimize the hardware implementation of the image compression algorithm through the determination of the fusion factor list and parameter mapping, reducing hardware power consumption and improving data processing efficiency.

Benefits of technology

Under storage constraints and bandwidth constraints, the target fusion factor is determined and parameter mapping is performed to reduce the number of accesses to DRAM, and the balance of memory access and computing time is achieved, thereby reducing hardware power consumption and improving the efficiency of image compression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088345A_ABST
    Figure CN120088345A_ABST
Patent Text Reader

Abstract

The invention provides an accelerator determination method based on a field-programmable gate array, an image processing method, equipment and a storage medium, which can be applied to the technical field of image compression. The accelerator determination method based on the field programmable gate array comprises the steps that a fusion factor list is determined according to the number of fusion units corresponding to a network model, the fusion factor list comprises a plurality of fusion factors, and the fusion factors represent the number of the fusion units needed for fusing the fusion units in the network model into a fusion structure; under the storage constraint and bandwidth constraint of the field programmable gate array, determining a target fusion factor from the plurality of fusion factors according to the image data block parameters and the fusion network parameter sets respectively corresponding to the plurality of fusion factors; and according to the target fusion factor and the fusion network parameter set corresponding to the target fusion factor, performing parameter mapping on the field programmable gate array to obtain an accelerator for realizing image compression.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image compression technology, and more particularly to an accelerator determination method, an image processing method, a device, and a storage medium. Background Art

[0002] In the process of compressing an image using a deep learning-based image compression algorithm, due to the high resolution of the processed image and the large number of parameters in the algorithm, the time required for image compression is relatively long. In the related art, the powerful parallel computing ability of a Graphics Processing Unit (GPU) is used for floating-point operations to process a large amount of data.

[0003] In the process of implementing the concept of the present disclosure, at least the following problems exist in the related art: The computing power of the GPU is strong, but the power consumption is relatively high, which is not suitable for power-sensitive edge devices. There is an urgent need for a method that can reduce the hardware power consumption of image compression and improve the data processing efficiency. Summary of the Invention

[0004] In view of the above problems, the present disclosure provides an accelerator determination method, an image processing method, a device, and a storage medium.

[0005] According to a first aspect of the present disclosure, an accelerator determination method based on a field programmable gate array is provided. The method includes: determining a fusion factor list according to the number of fusion units corresponding to a network model, where the fusion factor list includes a plurality of fusion factors, and the fusion factor represents the number of fusion units required to fuse the fusion units in the network model into a single fusion structure; under the storage constraint and bandwidth constraint of the field programmable gate array, determining a target fusion factor from the plurality of fusion factors according to the image data block parameters and the fusion network parameter sets respectively corresponding to the plurality of fusion factors; performing parameter mapping on the field programmable gate array according to the target fusion factor and the fusion network parameter set corresponding to the target fusion factor to obtain an accelerator for implementing image compression.

[0006] According to an embodiment of the present disclosure, under the storage constraint and bandwidth constraint of the field programmable gate array, determining a target fusion factor from the plurality of fusion factors according to the image data block parameters and the fusion network parameter sets respectively corresponding to the plurality of fusion factors includes: determining a plurality of first fusion factors that satisfy the bandwidth constraint from the plurality of fusion factors according to the image data block parameters and the fusion network parameter sets respectively corresponding to the plurality of fusion factors; determining a plurality of second fusion factors that satisfy the storage constraint from the plurality of first fusion factors according to the image data block parameters and the fusion network parameter sets respectively corresponding to the plurality of fusion factors; and determining the largest fusion factor among the plurality of second fusion factors as the target fusion factor.

[0007] According to an embodiment of the present disclosure, determining, from a plurality of fusion factors, a plurality of first fusion factors that satisfy a bandwidth constraint according to an image data block parameter and a set of fusion network parameters respectively corresponding to the plurality of fusion factors includes: starting from the smallest fusion factor, determining a model bandwidth overhead corresponding to the fusion factor according to the image data block parameter and the set of fusion network parameters respectively corresponding to the plurality of fusion factors until a marked fusion factor with a model bandwidth overhead that satisfies the bandwidth constraint is obtained; determining the fusion factors greater than the marked fusion factor in the fusion factor list as the first fusion factors to obtain a plurality of first fusion factors.

[0008] According to an embodiment of the present disclosure, determining a model bandwidth overhead corresponding to a fusion factor according to an image data block parameter and a set of fusion network parameters respectively corresponding to the plurality of fusion factors includes: determining an input bandwidth and an output bandwidth according to the image data block parameter and the set of fusion network parameters of each fusion structure; determining a model bandwidth overhead corresponding to the fusion factor according to the input bandwidth and the output bandwidth of each fusion structure.

[0009] According to an embodiment of the present disclosure, the set of fusion network parameters includes: the number of fusion structures, the number of input channels of each fusion structure, and the number of output channels of each fusion structure; determining a model bandwidth overhead corresponding to the fusion factor according to the image data block parameter and the set of fusion network parameters respectively corresponding to the plurality of fusion factors further includes: determining a model input bandwidth corresponding to the fusion factor according to the image data block parameter, the number of fusion structures, the number of input channels, and a preset inference time; determining a model output bandwidth corresponding to the fusion factor according to the image data block parameter, the number of fusion structures, the number of output channels, and the preset inference time; determining a model bandwidth overhead corresponding to the fusion factor according to the model input bandwidth corresponding to the fusion factor and the model output bandwidth corresponding to the fusion factor.

[0010] According to an embodiment of the present disclosure, the image data block parameters include the image data block length, the image data block width, and the number of image data blocks; the fusion network parameter set includes the parameter sets of each fusion structure, and the parameter set includes: the number of fusion network layers, the number of output channels of each fusion network layer, the number of input channels of each fusion network layer, the convolution kernel length of each fusion network layer, the convolution kernel width of each fusion network layer, the residual storage flag of each fusion network layer, and the overlapping data width of the blocks; determining, from a plurality of first fusion factors, a plurality of second fusion factors that satisfy the storage constraint according to the image data block parameters and the fusion network parameter sets respectively corresponding to a plurality of fusion factors, including: for each fusion structure, determining the overlapping data overhead of the blocks according to the image data block length, the number of fusion network layers, the number of output channels of each fusion network layer, and the overlapping data width of the blocks; determining the image data overhead according to the image data block length, the image data block width, and the number of input channels of each fusion network layer; determining the weight overhead according to the number of fusion network layers, the number of output channels of each fusion network layer, the number of input channels of each fusion network layer, the convolution kernel length of each fusion network layer, and the convolution kernel width of each fusion network layer; determining the residual storage overhead according to the number of output channels of each fusion network layer, the number of input channels of each fusion network layer, and the residual storage flag of each fusion network layer; determining the storage overhead of the fusion structure according to the overlapping data overhead of the blocks, the image data overhead, the weight overhead, and the residual storage overhead; and determining the maximum value of the storage overheads of each fusion structure as the model storage overhead corresponding to the fusion factor.

[0011] According to an embodiment of the present disclosure, the method further includes: in the case where the matrix width of the image data is greater than or equal to the matrix height of the image data, determining the matrix height of the image data as the image data block length and determining the preset length as the image data block length; in the case where the matrix width of the image data is less than the matrix height of the image data, determining the preset length as the image data block length and determining the matrix height of the image data as the image data block length.

[0012] The second aspect of the present disclosure provides an image processing method based on a field-programmable gate array. The method includes: in response to a received start instruction from the outside, calling an accelerator to process multiple image data blocks to obtain compressed image data; wherein, the accelerator is obtained by performing parameter mapping on the field-programmable gate array according to a target fusion factor and a fusion network parameter set corresponding to the target fusion factor; the target fusion factor is determined from a fusion factor list according to the image data block parameters and the fusion network parameter sets respectively corresponding to multiple fusion factors under the storage constraint and bandwidth constraint of the field-programmable gate array; the fusion factor list is based on the number of fusion units corresponding to the network model, wherein the fusion factor list includes multiple fusion factors, and the fusion factor represents the number of fusion units required to fuse the fusion units in the network model into a single fusion structure.

[0013] The third aspect of the present disclosure provides an electronic device, including: one or more processors; a memory for storing one or more computer programs, wherein the above one or more processors execute the above one or more computer programs to implement the steps of the above method.

[0014] The fourth aspect of the present disclosure provides a computer-readable storage medium, on which a computer program or instruction is stored, and when the computer program or instruction is executed by a processor, the steps of the above method are implemented.

[0015] According to the embodiments of the present disclosure, a fusion factor list is determined according to the number of fusion units in the network model, and the bandwidth constraint of the field-programmable gate array is used as a constraint on the execution efficiency of algorithms corresponding to different fusion factors. Under the dual constraints of the storage constraint and bandwidth constraint of the field-programmable gate array, the target fusion factor is determined, taking into account both storage resources and the efficiency of the accelerator. According to the target fusion factor and the fusion network parameter set corresponding to the target fusion factor, parameter mapping is performed on the field-programmable gate array, so that the field-programmable gate array can reduce the number of accesses to DRAM within the allowable range of hardware resources, achieve a balance between memory access and calculation time, and thus reduce the power consumption of the hardware. Description of the Drawings

[0016] Through the following description of the embodiments of the present disclosure with reference to the drawings, the above content and other objects, features, and advantages of the present disclosure will become clearer. In the drawings:

[0017] Figure 1 Schematically shows a flowchart of a method for determining an accelerator based on a field-programmable gate array according to an embodiment of the present disclosure.

[0018] Figure 2 Schematically shows a structural diagram of a network model according to an embodiment of the present disclosure.

[0019] Figure 3 Schematically shows a flowchart for determining a target fusion factor according to an embodiment of the present disclosure.

[0020] Figure 4 Schematically shows a block logic diagram of image data according to an embodiment of the present disclosure.

[0021] Figure 5 Schematically shows a processing logic diagram of image data according to an embodiment of the present disclosure.

[0022] Figure 6 Schematically shows an architecture diagram of an accelerator according to an embodiment of the present disclosure.

[0023] Figure 7 Schematically shows a block diagram of an electronic device suitable for implementing the method for determining a field programmable gate array-based accelerator according to an embodiment of the present disclosure. Detailed implementation manners

[0024] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, obviously, one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present disclosure.

[0025] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising" and the like used herein indicate the presence of the described features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components.

[0026] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0027] In cases where expressions similar to "at least one of A, B, and C, etc." are used, generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C).

[0028] In the technical solutions of the present disclosure, the user information involved (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties. Moreover, for the processing of relevant data such as collection, storage, use, processing, transmission, provision, disclosure, and application, all comply with relevant laws, regulations, and standards, adopt necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for users to choose to authorize or reject.

[0029] In the scenario of making automated decisions using personal information, the methods, devices, and systems provided by the embodiments of the present disclosure all provide corresponding operation entrances for users to choose to agree or reject the results of automated decisions; if the user chooses to reject, then the expert decision-making process is entered. The expression "automated decision" here refers to the activity of automatically analyzing and evaluating an individual's behavior habits, hobbies, or economic, health, credit status, etc. through a computer program and making a decision. The expression "expert decision" here refers to the activity of a person who specializes in a certain field of work, has specialized experience, knowledge, and skills, and has reached a certain professional level to make a decision.

[0030] The image compression algorithm based on deep learning means that the traditionally hand-designed filter is replaced by a Convolutional Neural Network (CNN) and end-to-end training is carried out. Experimental results show that this method outperforms the current state-of-the-art traditional image compression algorithms in terms of performance. Different from application fields such as image classification, the network architecture of the image compression algorithm has uniqueness: it does not contain pooling layers and fully connected layers, the resolution of the input picture is relatively high, for example, the ultra-high definition resolution of 4K or even higher, and the image size does not shrink with the increase of the network depth. The number of parameters is huge and the network structure is complex. For example, it contains many residual connections and skip structures. These characteristics lead to the fact that the image compression algorithm based on deep learning requires a large amount of time for calculation and accessing external memory, and the execution efficiency of the algorithm is low.

[0031] Currently, the common implementation methods for improving the execution efficiency of the algorithm mainly include two types: one is to perform model inference for image compression based on GPU, and the other is to perform fixed-point inference based on a dedicated convolutional neural network accelerator, that is, converting the floating-point model into a fixed-point model for inference. However, the first method is not applicable to the end-side application scenarios with strict power consumption requirements due to its high power consumption; the second method is difficult to fully meet the computational requirements of the non-linear network structure in the image compression algorithm, that is, in the process of image compression inference, it is impossible to directly perform inference on the non-linear network structure based on the fixed-point model.

[0032] The present disclosure aims to design an accelerator based on a Field Programmable Gate Array (FPGA) to implement an image compression algorithm based on deep learning through this accelerator.

[0033] The current mainstream processing mode of network models based on field programmable gate arrays is a layer-by-layer processing mode, that is, the output of each layer of the network model needs to be stored in the off-chip dynamic random access memory (DRAM) of the FPGA, and when this data is needed, the controller of the FPGA loads this data from the DRAM to the on-chip. This mode brings a large number of DRAM access operations. Especially when high-definition images are output, it becomes the main bottleneck affecting computing performance, and frequent off-chip access also leads to high power consumption.

[0034] The present disclosure circumvents the DRAM access of the intermediate layer through layer fusion technology. Therefore, an embodiment of the present disclosure provides a method for determining an accelerator based on a field programmable gate array, including: determining a fusion factor list according to the number of fusion units corresponding to the network model, where the fusion factor list includes multiple fusion factors, and the fusion factor represents the number of fusion units required to fuse the fusion units in the network model into a fusion structure; under the storage constraint and bandwidth constraint of the field programmable gate array, determining a target fusion factor from multiple fusion factors according to the image data block parameters and the fusion network parameter sets respectively corresponding to the multiple fusion factors; performing parameter mapping on the field programmable gate array according to the target fusion factor and the fusion network parameter set corresponding to the target fusion factor to obtain an accelerator for implementing image compression.

[0035] Figure 1 The flowchart of the method for determining an accelerator based on a field programmable gate array according to an embodiment of the present disclosure is schematically shown.

[0036] As Figure 1 shown, the method for determining an accelerator based on a field programmable gate array in this embodiment includes operation S110 to operation S130.

[0037] In operation S110, a fusion factor list is determined according to the number of fusion units corresponding to the network model, where the fusion factor list includes multiple fusion factors, and the fusion factor represents the number of fusion units required to fuse the fusion units in the network model into a fusion structure.

[0038] In operation S120, under the storage constraint and bandwidth constraint of the field programmable gate array, a target fusion factor is determined from multiple fusion factors according to the image data block parameters and the fusion network parameter sets respectively corresponding to the multiple fusion factors.

[0039] In operation S130, according to the target fusion factor and the set of fusion network parameters corresponding to the target fusion factor, parameter mapping is performed on the field programmable gate array to obtain an accelerator for implementing image compression.

[0040] According to an embodiment of the present disclosure, the network model represents the network structure for implementing the image compression algorithm. In the network model for implementing image compression, although non-linear structures such as residual blocks and long skip connections are generally present, overall, most network models are generally composed of repetitive structures and are not an irregular superposition of convolutional layers. Among them, the repetitive structures in the network model can be regarded as fusion units. The fusion unit may include one or more convolutional layers, and the fusion unit may also include one or more accumulation layers. There are multiple repetitive fusion units in the network model.

[0041] According to an embodiment of the present disclosure, a fusion factor list is determined according to the number of fusion units corresponding to the network model. For example, when the number of fusion units in the network model is 8, the fusion factor list may be {1, 2, 3, 4, 5, 6, 7, 8}. The structure diagram of the network model can be recognized by a model for image recognition to determine the number of fusion units in the network model.

[0042] According to an embodiment of the present disclosure, the fusion factor represents the number of fusion units in the fusion structure. Each fusion structure is formed by fusing fusion units according to the fusion factor and the data flow direction. For example, when the number of fusion units in the network model is 4, and the data flow of the fusion units is in the order of fusion unit 1, fusion unit 2, fusion unit 3, and fusion unit 4, the fusion factor of 2 means that fusion unit 1 and fusion unit 2 are used as one fusion structure, and fusion unit 3 and fusion unit 4 are used as one fusion structure.

[0043] According to an embodiment of the present disclosure, in order to reduce the power consumption of the computer during the implementation of the method, the method for determining the fusion factor list can be programmed as: in the initial fusion factor list, delete the fusion factors with the same number of formed fusion structures to obtain the fusion factor list. For example, when the number of fusion units in the network model is 7, the initial fusion factor list may be {1, 2, 3, 4, 5, 6, 7}. Among them, the fusion factors 5 and 6, and the fusion factor 4, generate the same number of fusion structures. Delete the larger fusion factors, that is, the fusion factors 5 and 6, to obtain the fusion factor list {1, 2, 3, 4, 7}.

[0044] According to an embodiment of the present disclosure, it is possible to round up the ratio of the number of fusion units to the fusion factor to determine whether the number of fusion structures formed by the fusion factor is the same. For example, when the number of fusion units in the network model is 7, for fusion factor 5 and fusion factor 6, compared with fusion factor 4, rounding up (7 / 4), rounding up (7 / 5), and rounding up (7 / 6), the obtained values are the same. Preferably, the smallest fusion factor 4 is retained.

[0045] According to an embodiment of the present disclosure, the storage constraint of the field programmable gate array is a preset value of the internal available storage resources of the field programmable gate array; the bandwidth constraint of the field programmable gate array mainly involves the limitations of data transmission rate and bandwidth, and the bandwidth constraint directly affects the performance and efficiency of the algorithm. To reduce the calculation errors caused by insufficient storage space in the field programmable gate array, the storage constraint can be set to be lower than the actual available storage resource value of the field programmable gate array. For example, the storage constraint can be set to 90% of the field programmable gate array.

[0046] According to an embodiment of the present disclosure, the target fusion factor represents the fusion factor that can make the image compression algorithm execute with higher efficiency on the field programmable gate array. The target fusion factor is the largest fusion factor that satisfies the storage constraint and the bandwidth constraint. The image data block parameter is the parameter of the data block that is transferred to the field programmable gate array for operation each time, and the image data block parameter is used to represent the size of the data block. The fusion network parameter set represents the set of model parameters of the fused network model corresponding to the fusion factor.

[0047] According to an embodiment of the present disclosure, the storage logic of the data in the network model calculation process can be identified by using the target fusion factor. For example, when the number of fusion units in the network model is 4, the fusion units transfer data in the order of fusion unit 1, fusion unit 2, fusion unit 3, and fusion unit 4, and the target fusion factor is 2. Then, fusion unit 1, fusion unit 2, fusion unit 3, and fusion unit 4 are respectively identified, so that the intermediate results of the fusion structure composed of fusion unit 1 and fusion unit 2, and the intermediate results of the fusion structure composed of fusion unit 3 and fusion unit 4 are stored on the FPGA, that is, the output data of fusion unit 1 and fusion unit 3 are stored on the FPGA, and the output data of fusion unit 2 and fusion unit 4 are stored in the DRAM.

[0048] According to an embodiment of the present disclosure, an FPGA configuration file is generated according to the fusion network parameter set corresponding to the target fusion factor, and the logical resources and storage resources inside the FPGA are allocated, etc., to map the network parameters to the hardware resources of the FPGA, so as to obtain an accelerator that can be used to implement image compression, where the configuration file can include a bitstream file.

[0049] According to an embodiment of the present disclosure, a fusion factor list is determined according to the number of fusion units in the network model, and the bandwidth constraint of the field programmable gate array is used as the constraint on the execution efficiency of algorithms corresponding to different fusion factors. Under the dual constraints of the storage constraint and the bandwidth constraint of the field programmable gate array, a target fusion factor is determined, taking into account both the storage resources and the efficiency of the accelerator. According to the target fusion factor and the fusion network parameter set corresponding to the target fusion factor, parameter mapping is performed on the field programmable gate array, so that the field programmable gate array can reduce the number of accesses to the DRAM within the allowable range of hardware resources, achieve a balance between memory access and computing time, and thus reduce the power consumption of the hardware.

[0050] Figure 2 Schematically shows a structural diagram of a network model according to an embodiment of the present disclosure.

[0051] As Figure 2 shown, in the first-level forward transform network structure of the iwave image compression algorithm, where Liftingstructure is a repetitive structure, and one Lifting structure is a fusion unit 201. The Liftingstructures are not in a linear structure, and each Lifting structure depends on the output of the previous level as the input. Each Lifting structure contains repetitive residual blocks, and each residual block may include multiple convolutional layers conv. The forward transform composed of 4 Lifting structures is a fusion structure 202.

[0052] According to an embodiment of the present disclosure, under the storage constraint and the bandwidth constraint of the field programmable gate array, according to the image data block parameters and the fusion network parameter sets respectively corresponding to multiple fusion factors, a target fusion factor is determined from multiple fusion factors, including: determining multiple first fusion factors that satisfy the bandwidth constraint from multiple fusion factors according to the image data block parameters and the fusion network parameter sets respectively corresponding to multiple fusion factors; determining multiple second fusion factors that satisfy the storage constraint from multiple first fusion factors according to the image data block parameters and the fusion network parameter sets respectively corresponding to multiple fusion factors; and determining the largest fusion factor among multiple second fusion factors as the target fusion factor.

[0053] According to an embodiment of the present disclosure, the current mainstream method for processing network models based on field-programmable gate arrays is a layer-by-layer processing mode. In the layer-by-layer mode, due to the high latency of accessing DRAM, the acceleration performance of the accelerator is often limited by memory access rather than the array computing power. If the number of fused layers is too small, in the most extreme case, the fusion factor is 1 and the fusion unit is a single network layer, which is equivalent to the layer-by-layer mode. At this time, the on-chip storage requirement is the smallest, but it may break the memory access-computation time balance, that is, it does not meet the bandwidth constraint. On the other hand, if the number of fused layers is too large, that is, the fusion factor is large, although the bandwidth constraint is met, the storage constraint is not met. Therefore, an optimal balance point needs to be found between the two.

[0054] According to an embodiment of the present disclosure, first, according to the image data block parameters and the fusion network parameter sets respectively corresponding to multiple fusion factors, multiple first fusion factors that meet the bandwidth constraint are determined from the multiple fusion factors. Among the multiple first fusion factors that meet the bandwidth constraint, multiple second fusion factors that meet the storage constraint are determined, so as to determine the target fusion factor from the second fusion factors, thereby avoiding the calculation of the storage requirements for all fusion factors and improving the processing efficiency.

[0055] According to an embodiment of the present disclosure, determining multiple first fusion factors that meet the bandwidth constraint from multiple fusion factors according to the image data block parameters and the fusion network parameter sets respectively corresponding to the multiple fusion factors includes: starting from the smallest fusion factor, according to the image data block parameters and the fusion network parameter sets respectively corresponding to the multiple fusion factors, determining the model bandwidth overhead corresponding to the fusion factor until a marked fusion factor whose model bandwidth overhead meets the bandwidth constraint is obtained; determining the fusion factors in the fusion factor list that are greater than the marked fusion factor as the first fusion factors to obtain multiple first fusion factors.

[0056] According to an embodiment of the present disclosure, the model bandwidth overhead is the bandwidth overhead of the network model in the field-programmable gate array during the actual inference process when constructing the network model according to the fusion factor. Starting from the smallest fusion factor, the model bandwidth overhead corresponding to the fusion factor is determined in ascending order until a fusion factor whose model bandwidth overhead meets the bandwidth constraint is obtained, and this fusion factor is determined as the marked fusion factor. Determining the fusion factors in the fusion factor list that are greater than the marked fusion factor as the first fusion factors to obtain multiple first fusion factors.

[0057] According to an embodiment of the present disclosure, the larger the fusion factor, the smaller the corresponding model bandwidth overhead. By determining the model bandwidth overhead corresponding to the fusion factor from small to large, the data calculation amount can be reduced and the efficiency of determining the accelerator can be improved.

[0058] According to an embodiment of the present disclosure, determining a model bandwidth overhead corresponding to a fusion factor according to an image data block parameter and a set of fusion network parameters respectively corresponding to a plurality of fusion factors includes: determining an input bandwidth and an output bandwidth according to the image data block parameter and the set of fusion network parameters of each fusion structure; determining the model bandwidth overhead corresponding to the fusion factor according to the input bandwidth and the output bandwidth of each fusion structure. Each fusion factor determines the fusion structure that constitutes the fused network model.

[0059] According to an embodiment of the present disclosure, by determining the input bandwidth and the output bandwidth of each fusion structure, the model bandwidth overhead of the network model constituted by the fusion structure is determined, so that the model bandwidth overhead of the network model can be roughly estimated, and thus the data processing efficiency of the accelerator can be evaluated during the accelerator determination stage to select a fusion factor with lower power consumption.

[0060] According to an embodiment of the present disclosure, the set of fusion network parameters includes: the number of fusion structures, the number of input channels of each fusion structure, and the number of output channels of each fusion structure; determining the model bandwidth overhead corresponding to the fusion factor according to the image data block parameter and the set of fusion network parameters respectively corresponding to a plurality of fusion factors further includes: determining the model input bandwidth corresponding to the fusion factor according to the image data block parameter, the number of fusion structures, the number of input channels, and a preset inference time; determining the model output bandwidth corresponding to the fusion factor according to the image data block parameter, the number of fusion structures, the number of output channels, and the preset inference time; determining the model bandwidth overhead corresponding to the fusion factor according to the model input bandwidth corresponding to the fusion factor and the model output bandwidth corresponding to the fusion factor.

[0061] According to an embodiment of the present disclosure, the number of fusion structures is the number of fusion structures in the network model. The number of input channels of each fusion structure is the number of channels of the data matrix input to the fusion structure, and the number of output channels of each fusion structure is the number of channels of the data matrix output by the last layer structure of the fusion structure. For the same image data, the preset inference time is the average time required for a network model without layer fusion to infer one frame of picture.

[0062] According to an embodiment of the present disclosure, when all the fusion structures corresponding to the fusion factor are exactly the same, the model input bandwidth can be expressed by formula (1):

[0063] (1);

[0064] Wherein, is the length of the image data block, is the number of fusion structures, is the width of the image data block, is the number of input channels of the fusion structure, is the number of image data blocks, is the preset inference time.

[0065] According to an embodiment of the present disclosure, when the respective fusion structures corresponding to the fusion factor are not completely the same, the input data volume of each fusion structure is calculated according to the product of the image data block length, the number of fusion structures, the number of input channels of the fusion structure, and the number of image data blocks, and the quotient of the sum of the input data volumes of each fusion structure and the preset inference time is used as the model input bandwidth.

[0066] According to an embodiment of the present disclosure, when the respective fusion structures corresponding to the fusion factor are completely the same, the model output bandwidth can be expressed by formula (2):

[0067] (2);

[0068] wherein, is the number of output channels of the fusion structure.

[0069] According to an embodiment of the present disclosure, when the respective fusion structures corresponding to the fusion factor are not completely the same, the input data volume of each fusion structure is calculated according to the product of the image data block length, the number of fusion structures, the number of output channels of the fusion structure, and the number of image data blocks, and the quotient of the sum of the input data volumes of each fusion structure and the preset inference time is used as the model input bandwidth.

[0070] According to an embodiment of the present disclosure, the model bandwidth overhead corresponding to the fusion factor can be expressed by formula (3):

[0071] (3);

[0072] According to an embodiment of the present disclosure, through the above method for determining the model bandwidth overhead, the performance and power consumption of the accelerator can be evaluated in the early stage of the design phase, so as to guide the selection of the fusion factor and the design of the accelerator, and finally realize an efficient and low-power accelerator.

[0073] According to an embodiment of the present disclosure, the image data block parameters include the image data block length, the image data block width, and the number of image data blocks; the fusion network parameter set includes the parameter sets of the respective fusion structures, and the parameter set includes: the number of fusion network layers, the number of output channels of each fusion network layer, the number of input channels of each fusion network layer, the length of the convolution kernel of each fusion network layer, the width of the convolution kernel of each fusion network layer, the residual storage identifier of each fusion network layer, and the overlapping data width of the blocks.

[0074] According to an embodiment of the present disclosure, determining, from a plurality of first fusion factors, a plurality of second fusion factors that satisfy storage constraints according to image data block parameters and a set of fusion network parameters corresponding to a plurality of fusion factors respectively includes: for each fusion structure, determining the chunk overlap data overhead according to the image data block length, the number of fusion network layers, the number of output channels of each fusion network layer, and the chunk overlap data width; determining the image data overhead according to the image data block length, the image data block width, and the number of input channels of each fusion network layer; determining the weight overhead according to the number of fusion network layers, the number of output channels of each fusion network layer, the number of input channels of each fusion network layer, the convolution kernel length of each fusion network layer, and the convolution kernel width of each fusion network layer; determining the residual storage overhead according to the number of output channels of each fusion network layer, the number of input channels of each fusion network layer, and the residual storage flag of each fusion network layer; determining the storage overhead of the fusion structure according to the chunk overlap data overhead, the image data overhead, the weight overhead, and the residual storage overhead; and determining the maximum value among the storage overheads of each fusion structure as the model storage overhead corresponding to the fusion factor.

[0075] According to an embodiment of the present disclosure, for each fusion structure, the chunk overlap data overhead can be represented by formula (4):

[0076] (4);

[0077] Wherein, N is the number of network layers included in the fusion structure, and the network layer includes a convolutional layer and an accumulation layer, that is, the number of fusion network layers; is the number of output channels of the nth fusion network layer, where 2 is the chunk overlap data width.

[0078] According to an embodiment of the present disclosure, for each fusion structure, the image data overhead can be represented by formula (5):

[0079] (5);

[0080] Wherein, is the number of output channels of the first fusion network layer of the fusion structure.

[0081] According to an embodiment of the present disclosure, for each fusion structure, the weight overhead can be represented by formula (6):

[0082] (6);

[0083] Wherein, is the convolution kernel length of the nth fusion network layer; is the convolution kernel width of the nth fusion network layer; is the number of input channels of the nth fusion network layer.

[0084] According to an embodiment of the present disclosure, for each fusion structure, the residual storage overhead can be expressed by formula (7):

[0085] (7);

[0086] where L is the number of residual storage identifiers in the fusion structure, is the number of output channels of the m-th fusion network layer.

[0087] According to an embodiment of the present disclosure, for each fusion structure, the storage overhead can be expressed by formula (8):

[0088] + + (8);

[0089] According to an embodiment of the present disclosure, when the respective fusion structures corresponding to the same fusion factor are different, the maximum value among the storage overheads of the respective fusion structures is determined as the model storage overhead corresponding to the fusion factor.

[0090] According to an embodiment of the present disclosure, by separately calculating the block overlapping data overhead, the image data overhead, the weight overhead, and the residual storage overhead, the storage requirements of each fusion structure and the entire model can be more accurately evaluated, so as to determine the optimal fusion factor under storage constraints by comparing the model storage overheads of different fusion factors, providing guidance for the hardware parameter mapping of the FPGA to improve the image processing efficiency of the accelerator.

[0091] According to an embodiment of the present disclosure, this method of determining the model bandwidth overhead and the model storage overhead enables the finally determined accelerator to support both the fusion of convolutional layers and the fusion of accumulation layers, enhancing the versatility and flexibility of the accelerator, and is particularly suitable for complex networks including residual connections and skip structures.

[0092] Figure 3 Schematically shows a flowchart of determining a target fusion factor according to an embodiment of the present disclosure.

[0093] As Figure 3 shown, the flowchart of determining the target fusion factor in this embodiment includes operations S301 to S312.

[0094] In operation S301, the fusion factors in the fusion factor list are sorted, and the smallest fusion factor in the sequence is determined.

[0095] In operation S302, according to the image data block parameters and the fusion network parameter sets respectively corresponding to multiple fusion factors, the model bandwidth overhead corresponding to the fusion factor is determined.

[0096] In operation S303, it is determined whether the model bandwidth overhead corresponding to the fusion factor satisfies the bandwidth constraint.

[0097] In the case where the model bandwidth overhead corresponding to the fusion factor does not satisfy the bandwidth constraint, operation S304 is executed; in the case where the model bandwidth overhead corresponding to the fusion factor satisfies the bandwidth constraint, operation S305 is executed.

[0098] In operation S304, a larger fusion factor is selected.

[0099] In operation S305, the fusion factor that satisfies the bandwidth constraint is determined as the marked fusion factor.

[0100] In operation S306, the fusion factors in the fusion factor list that are greater than the marked fusion factor are determined as the first fusion factors.

[0101] In operation S307, according to the image data block parameters and the fusion network parameter sets respectively corresponding to multiple fusion factors, the model storage overhead corresponding to the first fusion factor is determined.

[0102] In operation S308, it is determined whether the model storage overhead corresponding to the fusion factor satisfies the storage constraint.

[0103] In the case where the model storage overhead corresponding to the fusion factor satisfies the storage constraint, operation S309 is executed; in the case where the model storage overhead corresponding to the fusion factor does not satisfy the storage constraint, operation S311 is executed.

[0104] In operation S309, it is determined whether all the first fusion factors have been traversed.

[0105] In operation S310, a larger fusion factor is selected.

[0106] In operation S311, the fusion factors that satisfy the storage constraint, the marked fusion factor, and the fusion factors that are greater than the marked fusion factor and less than the fusion factor that satisfies the storage constraint are determined as the second fusion factors.

[0107] In operation S312, according to the image data block parameters and the fusion network parameter sets respectively corresponding to multiple second fusion factors, the model storage overhead corresponding to the fusion factor is determined.

[0108] According to an embodiment of the present disclosure, the method for determining an FPGA-based accelerator further includes: when the matrix width of the image data is greater than or equal to the matrix height of the image data, determining the matrix height of the image data as the image data block length and determining a preset length as the image data block length; when the matrix width of the image data is less than the matrix height of the image data, determining the preset length as the image data block length and determining the matrix height of the image data as the image data block length.

[0109] According to an embodiment of the present disclosure, the matrix width of the image data may be the width in the row direction of the matrix of the image data, and the matrix height of the image data may be the width in the column direction of the matrix of the image data. The preset length is preferably 8 bits.

[0110] Figure 4 Schematically shows a block division logic diagram of image data according to an embodiment of the present disclosure.

[0111] As Figure 4 shown, when the matrix width (col) of the image data is greater than or equal to the matrix height (row) of the image data, in the matrix width direction of the image data, the image data is sliced according to a preset image data block width to obtain a plurality of image data blocks.

[0112] When the matrix width (col) of the image data is less than the matrix height (row) of the image data, in the matrix width direction of the image data, the image data is sliced according to a preset image data block width to obtain a plurality of image data blocks.

[0113] As Figure 4 shown, Tile is the image data block. In the actual calculation process, each image data block may be further divided into multiple blocks. Overlap represents the block overlap data, and the block overlap data is the data between the Tile and the next adjacent Tile during the process of processing the image data in blocks. This block division method helps to reduce the on-chip storage overhead of the block feature map, and only the Overlap in one direction needs to be considered for layer fusion, reducing the storage amount of the Overlap data. Thereby improving the data processing efficiency of the accelerator.

[0114] Dynamically selecting the block division method and dividing the Tile only along one direction reduces the on-chip storage overhead of the block feature map, simplifies the overlap processing of layer fusion, and is particularly suitable for high-resolution image processing.

[0115] Figure 5 Schematically shows a processing logic diagram of image data according to an embodiment of the present disclosure.

[0116] As Figure 5As shown, the fusion unit includes nine convolutional layers. Among them, the top layer of the three-dimensional matrix in the figure is the input data matrix, and the remaining layers are the outputs of each network layer. L1c~L4c, L6c, L7c, and L9c are the outputs of the convolutional layers, and L5a and L8a are the outputs of the two accumulation layers. Image data is read from off-chip DRAM. The data flow between each fusion network layer shows three coordinate axes: the row direction, column direction, and layer calculation direction of the feature map. Each square represents a pixel. Among them, the two columns of data on the right boundary of the image data block are used as the Overlap data of the next image data block.

[0117] During the calculation process, the number of columns of pre-computed image data can be selected according to the number of fusion network layers in the fusion unit. For example, Figure 5 As shown, when the total number of convolutional layers and accumulations is 9, 9 columns of data on the left side of the image data can be selected as the calculation data. Among them, the unit of each column of data is a pixel to determine the Overlap data of the first Tile.

[0118] First, load part of the data onto the chip. After the convolution / accumulation calculations from L1 to L9, retain the two columns of pink diagonal data on the right boundary in the overlap buffer. For the convolutional layer, during the calculation of the Tile, in order to ensure that the size remains unchanged after convolution, two columns of data need to be filled on its left boundary. For the left boundary of Tilen, its calculation depends on the adjacent pink diagonal data on the left and top. Through layer fusion, these adjacent data are calculated by Tilen-1 and saved in the finite-sized Ooverlap buffer, and are updated in a timely manner to avoid data loss.

[0119] The overall data forms a parallelepiped along the layer direction, and the tile of the next layer is shifted one pixel to the left. This block layer fusion data flow has the following three major advantages: 1. It solves the problem that traditional layer fusion cannot adapt to the situation where the image size does not decrease with depth in the image compression algorithm. 2. It eliminates the overlap in one direction instead of directly discarding it, thus avoiding accuracy loss. 3. It supports the fusion of non-linear structures, not only supporting the fusion of convolutional layers but also the fusion of accumulation layers.

[0120] According to an embodiment of the present disclosure, an FPGA-based image processing method includes: in response to a received start instruction from the outside, calling an accelerator to process a plurality of image data blocks to obtain compressed image data; wherein, the accelerator is obtained by performing parameter mapping on the field programmable gate array according to a target fusion factor and a fusion network parameter set corresponding to the target fusion factor; the target fusion factor is determined from a fusion factor list according to image data block parameters and fusion network parameter sets respectively corresponding to a plurality of fusion factors under the storage constraint and bandwidth constraint of the field programmable gate array; the fusion factor list is based on the number of fusion units corresponding to the network model, wherein the fusion factor list includes a plurality of fusion factors, and the fusion factor represents the number of fusion units required to fuse the fusion units in the network model into a single fusion structure.

[0121] According to an embodiment of the present disclosure, before starting the accelerator, the layout of data in the DRAM is determined in advance, specifically including the number of network layers, bias, weights, input pictures, and network parameters, and each data is accessed according to the starting address, where the network parameters include the input pad length, convolution stride, convolution kernel size, number of input data rows, number of input data columns, number of column blocks of the input picture, number of row blocks of the input picture, number of column blocks of the output picture, number of row blocks of the output picture, number of valid array rows, number of valid array columns, number of input channel blocks, number of output channel blocks, accumulation layer flag, and fusion layer flag.

[0122] According to an embodiment of the present disclosure, in response to a received start instruction from the outside, calling an accelerator to process a plurality of image data blocks included in the image data to obtain compressed image data, where the compressed image data is the output result of the network model.

[0123] According to an embodiment of the present disclosure, calling an accelerator to process a plurality of image data blocks to obtain compressed image data is divided into three major modules in terms of function: a calculation module, a storage module, and a control logic. The control logic generates control signals to control the data flow and configure the parameters for each layer of the network to operate. The calculation module includes a convolution operation module, an accumulation module, and a matrix transpose module. Among them, the convolution operation array is used to complete the multiplication and addition operation of convolution in batches and output the result of the partial sum. The array adopts the structure of a systolic array with fixed weight data flow.

[0124] Since the maximum convolution kernel size in the algorithm is 3x3 and the maximum number of channels of the transformation network is 16. From the perspective of improving the array utilization rate and reducing power consumption, the array size is set to 18x16: 16 columns of processing elements (PEs) support convolution calculations for 16 output channels in parallel, and 18 rows of PEs support 3x3 convolutions for 2 input channels. The storage module mainly includes 7 parts.

[0125] The computing module includes a convolution operation module, an accumulation module, and a matrix transpose module. Among them, the convolution operation array is used to complete the multiplication and addition operation of convolution in batches and output the result of the partial sum. The array adopts a systolic array structure with fixed weights and data flow. Since the maximum size of the convolution kernel in the algorithm is 3x3 and the maximum number of channels in the transformation network is 16, from the perspective of improving the array utilization rate and reducing power consumption, the array size is set to 18x16: 16 columns of PEs support the convolution calculation of 16 output channels in parallel, and 18 rows of PEs support the 3x3 convolution of 2 input channels.

[0126] Figure 6 FIG. schematically shows an architecture diagram of an accelerator according to an embodiment of the present disclosure.

[0127] As Figure 6 shown, the accelerator includes a controller (Controller). The controller is the command center of the entire system and is responsible for generating control signals to control the data flow and configure the parameters for the operation of each layer of the network. In_buf: input buffer, used to cache the picture data adapted to the calculation of the PE array. Each row of PEs corresponds to a ping-pong buffer. W_buf: weight buffer, used to cache the weight data adapted to the operation of the systolic array. Each row of PEs corresponds to a ping-pong buffer.

[0128] weight buffer: weight register, used to cache the weight data of all fusion layers from off-chip.

[0129] Processing element PE: the core for performing convolution operations. Each PE can perform one or more multiplication and addition operations. Multiple PEs are shown in the picture, and they can work in parallel to accelerate the convolution operation. Activation buffers (Act buffer0 and Act buffer1) The activation buffers are used to cache the intermediate results of the tiles or block layer fusions from off-chip. When one buffer is read, the other can be written, thereby improving the data processing efficiency.

[0130] The accumulator is used to accumulate the outputs of the PEs, and these outputs are the results of the convolution operations. The accumulator can process the outputs of multiple PEs and add them together to form the final convolution result.

[0131] The post-processing module is used to perform any necessary operations after the convolution operation, such as applying activation functions, normalization, etc. Scalar random access memory (Scalar RAM) The scalar RAM is used to store scalar values, which may be used for control signals, configuration parameters, etc. Bias random access memory (Bias RAM) is used to store the bias terms in the convolution operation.

[0132] The residual buffer is used to cache the inter-layer data for the accumulation of the residual structure, which helps with gradient flow and improves model performance. The overlap buffer is used to cache the overlapping data (Overlap) of the input data to ensure that the convolution operation can cover the entire input area.

[0133] DRAM is an external memory used to store a large amount of data and model parameters. Transposition is used to perform a transpose operation on the data. Line split is used to split the image data.

[0134] After starting the accelerator operation, the controller controls the import of the network parameters of the current layer from the DRAM. It is determined whether to update the tile data from the DRAM and store the on-chip results back to the DRAM according to the fusion layer flag of the network parameters, where the fusion layer flag is determined based on the target fusion factor, that is, different identifiers are used to identify the last network layer and other network layers of the fusion structure, so that the controller can determine whether to update the tile data from the DRAM and store the on-chip results back to the DRAM.

[0135] Calculate the start addresses and read data lengths of the feature map, weights, and biases of the current layer according to other network parameters. During the convolution calculation process, the controller is mainly divided into 5 processes: reading the weights and feature map from the DRAM to the on-chip cache; reading the weights from the weight buffer into the PE; reading the feature map data into the array for calculation; the array outputs part of the data and accumulates / post-processes the data between channels to the output cache; storing the data in the output memory back to the activation buffer or the DRAM.

[0136] The specific operation steps of the accelerator module are as follows:

[0137] Step 1: Initial data loading: Read all the weight data of the first convolutional layer into the weight register (weight buffer), and at the same time read the entire tile data into Act buffer0.

[0138] Step 2: Weight and input data loading: Load the weights of the input channels and output channels of the current layer from the weight buffer. And copy these weight data to each w_buf. Then, load the blockn data of the input channels from Act buffer0 and copy it to each in_buf. This step ensures the readiness of the input data and weight data for the convolution calculation.

[0139] Step 3: Convolution calculation initialization: First, fix the weights in w_buf to the weight buffer of the array. Then, load the feature map data from in_buf into the array to complete the convolution calculation. The partial sum results output by the boundary PEs are written into Act buffer1. This step marks the official start of the convolution calculation.

[0140] Step 4: Multi-channel convolution calculation. Repeat the process of Step 3 until all input channels are traversed. During this process, the partial results of int48 output and the intermediate data temporarily stored in Act buffer1 are accumulated by the accumulator and then written back to Act buffer1. At the same time, read the new blockn+1 data from Act buffer0 and update it to the other block cache of the ping-pong in_buf. This step ensures the continuity of the multi-channel convolution calculation.

[0141] Step 5: Post-processing and result storage: Read the partial sum results of int48 cached in Act buffer1. Among them, the bit width of the feature map usually affects the calculation accuracy and efficiency. For a feature map with a bit width of int16, although it can save storage space and computing resources, in operations such as convolution and activation, the intermediate results may exceed the representation range of int16, resulting in overflow. Therefore, the intermediate results are represented by the wide bit width of int48. Send them to the post-processing module. After bias processing, multiplying by the scaling factor, and processing with the non-linear activation function, the final result of blockn is obtained and written back to Act buffer1. The results on the right boundary are written into the overlap buffer. This step completes the post-processing and storage of the convolution results.

[0142] Repeat the process of Step 2 to Step 5 until the calculation of all blockns in the current tile is completed. This step ensures that the convolution calculation and post-processing of all blocks within the tile are completed.

[0143] Step 6: Cache swapping and continued calculation: Swap the read and write roles of Act buffer1 and Act buffer0. Read the feature map data from Act buffer1 to participate in the calculation, and repeat the process of Step 2 to Step 6. The convolution results are written into Act buffer0. This step realizes the dynamic management of the cache and the continuity of the calculation.

[0144] Repeat the above process until the inference of a fused structure is completed. Subsequently, reload a new tile from off-chip DRAM to participate in the calculation until the calculation of the entire image data is completed.

[0145] Figure 7A block diagram of an electronic device suitable for implementing a field-programmable gate array-based accelerator determination method according to an embodiment of the present disclosure is schematically shown.

[0146] As Figure 7 shown, the electronic device 700 according to an embodiment of the present disclosure includes a processor 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage section 708 into a random access memory (RAM) 703. The processor 701 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 701 may also include on-board memory for caching purposes. The processor 701 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.

[0147] In the RAM 703, various programs and data required for the operation of the electronic device 700 are stored. The processor 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. The processor 701 performs various operations of the method flow according to an embodiment of the present disclosure by executing the program in the ROM 702 and / or the RAM 703. It should be noted that the program may also be stored in one or more memories other than the ROM 702 and the RAM 703. The processor 701 may also perform various operations of the method flow according to an embodiment of the present disclosure by executing the program stored in the one or more memories.

[0148] According to an embodiment of the present disclosure, the electronic device 700 may further include an input / output (I / O) interface 705, and the input / output (I / O) interface 705 is also connected to the bus 704. The electronic device 700 may further include one or more of the following components connected to the input / output (I / O) interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A driver 710 is also connected to the input / output (I / O) interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the driver 710 as needed so that a computer program read from it can be installed into the storage section 708 as needed.

[0149] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the methods according to the embodiments of the present disclosure are implemented.

[0150] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include one or more memories other than the above-described ROM 702 and / or RAM 703 and / or ROM 702 and RAM 703.

[0151] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiments of the present disclosure may be written in any combination of one or more programming languages. Specifically, these computing programs may be implemented using high-level procedures and / or object-oriented programming languages, and / or assembly / machine languages. The programming languages include but are not limited to, such as Java, C++, python, "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).

[0152] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.

[0153] Those skilled in the art will appreciate that the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present disclosure.

[0154] The embodiments of the present disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although the embodiments have been described separately above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present disclosure.

Claims

1. A method for determining an accelerator based on a field programmable gate array, characterized in that: The method comprises: Determine a fusion factor list according to the number of fusion units corresponding to the network model, wherein the fusion factor list includes a plurality of fusion factors, and the fusion factor represents the number of fusion units required to fuse the fusion units in the network model into a fusion structure; Under the storage constraint and bandwidth constraint of the field programmable gate array, according to the image data block parameters and the fusion network parameter sets respectively corresponding to the multiple fusion factors, determine the target fusion factor from the multiple fusion factors; According to the target fusion factor and the fusion network parameter set corresponding to the target fusion factor, parameter mapping is performed on the field programmable gate array to obtain an accelerator for implementing image compression.

2. The method according to claim 1, characterized in that The method of determining a target fusion factor from a plurality of fusion factors according to image data block parameters and fusion network parameter sets respectively corresponding to the plurality of fusion factors under the storage constraint and bandwidth constraint of the field programmable gate array comprises: Determining, from the plurality of fusion factors, a plurality of first fusion factors satisfying the bandwidth constraint according to the image data block parameters and the fusion network parameter sets respectively corresponding to the plurality of fusion factors; Determining, from the plurality of first fusion factors, a plurality of second fusion factors satisfying the storage constraint according to the image data block parameters and the fusion network parameter sets respectively corresponding to the plurality of fusion factors; The largest fusion factor among the plurality of the second fusion factors is determined as the target fusion factor.

3. The method according to claim 2, characterized in that The step of determining, from the plurality of fusion factors, a plurality of first fusion factors satisfying the bandwidth constraint according to the image data block parameters and the fusion network parameter sets respectively corresponding to the plurality of fusion factors, comprises: Starting from the smallest fusion factor, according to the image data block parameters and the fusion network parameter sets corresponding to the multiple fusion factors, the model bandwidth overhead corresponding to the fusion factor is determined until a marked fusion factor whose model bandwidth overhead satisfies the bandwidth constraint is obtained; A fusion factor in the fusion factor list that is greater than the marker fusion factor is determined as the first fusion factor to obtain a plurality of first fusion factors.

4. The method according to claim 3, characterized in that The determining, according to the image data block parameters and the fusion network parameter sets corresponding to the multiple fusion factors, the model bandwidth overhead corresponding to the fusion factors, comprises: Determining an input bandwidth and an output bandwidth according to the image data block parameters and the fusion network parameter sets of each fusion structure; The model bandwidth overhead corresponding to the fusion factor is determined according to the input bandwidth and the output bandwidth of each fusion structure.

5. The method according to claim 4, characterized in that The fusion network parameter set includes: the number of fusion structures, the number of input channels of each fusion structure, and the number of output channels of each fusion structure; The step of determining the model bandwidth overhead corresponding to the fusion factor according to the image data block parameters and the fusion network parameter sets respectively corresponding to the multiple fusion factors further includes: Determining a model input bandwidth corresponding to a fusion factor according to the image data block parameters, the number of fusion structures, the number of input channels, and a preset inference time; Determining a model output bandwidth corresponding to a fusion factor according to the image data block parameters, the number of fusion structures, the number of output channels, and the preset inference time; The model bandwidth overhead corresponding to the fusion factor is determined according to the model input bandwidth corresponding to the fusion factor and the model output bandwidth corresponding to the fusion factor.

6. The method according to claim 2, characterized in that The image data block parameters include image data block length, image data block width and image data block quantity; The fusion network parameter set includes parameter sets of each fusion structure, and the parameter set includes: the number of fusion network layers, the number of output channels of each fusion network layer, the number of input channels of each fusion network layer, the convolution kernel length of each fusion network layer, the convolution kernel width of each fusion network layer, the residual storage identifier of each fusion network layer, and the block overlap data width; The determining, from the plurality of first fusion factors, a plurality of second fusion factors satisfying the storage constraint according to the image data block parameters and the fusion network parameter sets respectively corresponding to the plurality of fusion factors, comprises: For each fusion structure, Determine the block overlapping data overhead according to the image data block length, the number of fusion network layers, the number of output channels of each fusion network layer and the block overlapping data width; Determining image data overhead according to the image data block length, the image data block width, and the number of input channels of each fusion network layer; Determine the weight overhead according to the number of fused network layers, the number of output channels of each fused network layer, the number of input channels of each fused network layer, the convolution kernel length of each fused network layer, and the convolution kernel width of each fused network layer; Determining residual storage overhead according to the number of output channels of each fused network layer, the number of input channels of each fused network layer, and the residual storage identifier of each fused network layer; Determine the storage overhead of the fusion structure according to the block overlap data overhead, the image data overhead, the weight overhead and the residual storage overhead; The maximum value of the storage costs of the fusion structures is determined as the model storage cost corresponding to the fusion factor.

7. The method according to claim 6, characterized in that The method further comprises: In a case where the matrix width of the image data is greater than or equal to the matrix height of the image data, determining the matrix height of the image data as the image data block length, and determining the preset length as the image data block length; In a case where the matrix width of the image data is smaller than the matrix height of the image data, the preset length is determined as the image data block length, and the matrix height of the image data is determined as the image data block length.

8. An image processing method based on a field programmable gate array, characterized in that: The method comprises: In response to a received start instruction from the outside, calling the accelerator to process the plurality of image data blocks to obtain compressed image data; The accelerator is obtained by performing parameter mapping on a field programmable gate array according to a target fusion factor and a fusion network parameter set corresponding to the target fusion factor; the target fusion factor is determined from a fusion factor list according to image data block parameters and fusion network parameter sets corresponding to multiple fusion factors under storage constraints and bandwidth constraints of the field programmable gate array; the fusion factor list is based on the number of fusion units corresponding to a network model, wherein the fusion factor list includes multiple fusion factors, and the fusion factor represents the number of fusion units required to fuse the fusion units in the network model into a fusion structure.

9. An electronic device, comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Image compression method and image compression architecture

    CN121883622A

  • Image compression method and image compression architecture

    CN121883622B