FPGA chip configuration method and device, program product, medium and electronic equipment

By dynamically configuring the processing unit in the FPGA chip to adapt to the computing needs of the deep learning model, the problem of low resource utilization caused by the fixed GPU architecture is solved, and efficient utilization of hardware resources and performance improvement is achieved.

CN120256381AActive Publication Date: 2025-07-04北京汤谷软件技术有限公司
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510747915.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-07-04
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

In the prior art, GPUs with fixed architectures cannot achieve optimal resource utilization when facing diversified deep learning tasks, and may cause waste of computing power or computing power bottlenecks, making it difficult to efficiently provide deep learning models with hardware resources that dynamically adapt to their computing needs.

Method used

By obtaining the operation index of each converter module in the deep learning model, and combining the total amount of programmable hardware resources in the FPGA chip, multiple processing units are dynamically configured to adapt to the operation amount of the converter module, and the reasonable allocation and flexible adjustment of hardware resources are achieved.

Benefits of technology

It improves the efficiency of hardware resources utilization, avoids resource waste, significantly improves the performance and efficiency of the training and inference process of deep learning models, and has good scalability and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256381A_ABST
    Figure CN120256381A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of chip configuration, and provides an FPGA chip configuration method and device, a program product, a medium and electronic equipment. The method comprises the steps that the operation index of each converter module in a deep learning model is acquired, and the operation index is used for representing the operand of the converter module; obtaining the total amount of programmable hardware resources in an FPGA chip, wherein the FPGA chip is used for executing an operation task on the deep learning model; a plurality of processing units are configured in the FPGA chip according to the operation indexes of the converter modules and the total hardware resource amount, and the processing units are matched with the operation amount of the converter modules. Through the technical scheme provided by the invention, hardware resources capable of dynamically adapting to the operation requirements of the deep learning model can be efficiently provided for the deep learning model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of chip configuration, and particularly relates to a configuration method, device, program product, medium, and electronic device for an FPGA chip. Background Art

[0002] With the rapid development of deep learning technology, the complexity and computing requirements of deep learning models have been continuously increasing, and the requirements for their supporting hardware have also become higher and higher. As the most important computing acceleration hardware in the field of deep learning, the GPU (Graphics Processing Unit) has become the preferred platform for training and inferring large-scale neural network models by virtue of its powerful parallel computing ability and high-bandwidth memory access. However, once the GPU is designed or produced, its hardware architecture is fixed. Different deep learning models have different architectures, and there are also significant differences in the requirements for hardware resources. Therefore, when facing diverse deep learning tasks, a GPU with a fixed architecture may not be able to achieve the optimal resource utilization rate, and may even encounter problems such as computing power waste or computing power bottlenecks. Based on this, how to efficiently provide hardware resources that can dynamically adapt to the computing requirements of deep learning models has become a technical problem to be solved urgently. Summary of the Invention

[0003] Embodiments of this application provide a configuration method, device, computer program product, computer-readable storage medium, and electronic device for an FPGA chip, which can thereby efficiently provide hardware resources that can dynamically adapt to the computing requirements of deep learning models.

[0004] Other features and advantages of this application will become apparent through the following detailed description, or be partially learned through the practice of this application.

[0005] According to the first aspect of the embodiments of this application, a configuration method for an FPGA chip is provided. The method includes: obtaining the operation indices of each transformer module in a deep learning model, where the operation indices are used to characterize the amount of operations of the transformer module; obtaining the total amount of programmable hardware resources in the FPGA chip, where the FPGA chip is used to execute the computing task of the deep learning model; and configuring a plurality of processing units in the FPGA chip according to the operation indices of each transformer module and the total amount of hardware resources, where the processing units are adapted to the amount of operations of the transformer module.

[0006] In some embodiments of the present application, based on the foregoing solution, configuring a plurality of processing units in the FPGA chip according to the operation exponents of the respective converter modules and the total amount of hardware resources includes: allocating hardware resources to the respective converter modules according to the operation exponents of the respective converter modules and the total amount of hardware resources, wherein the allocated amount of hardware resources for each converter module is positively correlated with the operation exponent of each converter module, and the sum of the allocated amounts of hardware resources for the respective converter modules is less than or equal to the total amount of hardware resources; configuring a plurality of processing units in the FPGA chip based on the hardware resources allocated to the respective converter modules.

[0007] In some embodiments of the present application, based on the foregoing solution, allocating hardware resources to the respective converter modules according to the operation exponents of the respective converter modules and the total amount of hardware resources includes: determining the total sum of the operation exponents of the respective converter modules, and determining a first proportion of the operation exponent of each converter module in the total sum of the operation exponents; allocating hardware resources to each converter module based on the corresponding first proportion of each converter module, wherein the absolute value of the difference between the second proportion of the allocated amount of hardware resources for each converter module in the total amount of hardware resources and the first proportion is less than a preset proportion.

[0008] In some embodiments of the present application, based on the foregoing solution, configuring a plurality of processing units in the FPGA chip based on the hardware resources allocated to the respective converter modules includes: generating a first chip configuration file based on the hardware resources allocated to the respective converter modules, the first chip configuration file being used to describe the state and connection of the internal hardware resources of the FPGA chip; loading the first chip configuration file into the FPGA chip to configure a plurality of processing units in the FPGA chip.

[0009] In some embodiments of the present application, based on the foregoing solution, each converter module includes a plurality of network modules. After configuring a plurality of processing units in the FPGA chip, the method further includes: configuring processing subunits corresponding one-to-one to the network modules in each converter module in the processing unit corresponding to each converter module, the processing subunits being used to execute the operation tasks for the network modules.

[0010] In some embodiments of the present application, based on the foregoing solution, based on the hardware resources allocated to each converter module, a plurality of processing units are configured in the FPGA chip, and processing subunits corresponding one-to-one to the network modules in each converter module are configured in the processing unit corresponding to each converter module, including: respectively determining the number of network modules included in each converter module; generating a second chip configuration file based on the hardware resources allocated to each converter module and the number of network modules included in each converter module, where the second chip configuration file is used to describe the state and connection of the internal hardware resources of the FPGA chip; loading the second chip configuration file into the FPGA chip to configure a plurality of processing units in the FPGA chip, and configure processing subunits corresponding one-to-one to the network modules in each converter module in the processing unit corresponding to each converter module.

[0011] In some embodiments of the present application, based on the foregoing solution, the converter module includes a Transformer block, and the plurality of network modules include a plurality of parallel attention network modules and a plurality of parallel expert network modules.

[0012] In some embodiments of the present application, based on the foregoing solution, the method further includes: executing a training operation task of the deep learning model through the FPGA chip; or, executing an inference operation task of the deep learning model through the FPGA chip.

[0013] According to a second aspect of the embodiments of the present application, there is provided a configuration device for an FPGA chip, the device including: a first acquisition unit, configured to acquire the operation index of each converter module in the deep learning model, where the operation index is used to characterize the operation amount of the converter module; a second acquisition unit, configured to acquire the total amount of programmable hardware resources in the FPGA chip, where the FPGA chip is used to execute the operation task of the deep learning model; a configuration unit, configured to configure a plurality of processing units in the FPGA chip according to the operation index of each converter module and the total amount of hardware resources, and the processing unit adapts to the operation amount of the converter module.

[0014] According to a third aspect of the embodiments of the present application, there is provided a computer program product, the computer program product including computer instructions, where the computer instructions are stored in a computer-readable storage medium and are adapted to be read and executed by a processor so that a computer device having the processor executes to implement the operations performed by the method as described in the first aspect above.

[0015] According to a fourth aspect of the embodiments of the present application, there is provided a computer-readable storage medium storing at least one computer program instruction, and the at least one computer program instruction is loaded and executed by a processor to implement the operations performed by the method described in the first aspect above.

[0016] According to a fifth aspect of the embodiments of the present application, there is provided an electronic device including one or more processors and one or more memories, and at least one computer program instruction is stored in the one or more memories, and the at least one computer program instruction is loaded and executed by the one or more processors to implement the operations performed by the method described in the first aspect above.

[0017] In the present application, based on the reconfigurable characteristics of the FPGA chip architecture, by obtaining the operation exponents of each transformer module in the deep learning model and combining with the total amount of programmable hardware resources in the FPGA chip, an efficient configuration of the FPGA chip is achieved. Compared with the method of using a chip with a fixed architecture to execute the deep learning model operation task in the prior art, the technical solution proposed in the present application can give full play to the programmable advantages of the FPGA chip, dynamically configure the processing units on the FPGA chip according to the actual operation requirements of different transformer modules, so that each transformer module obtains hardware resources matching its operation volume. This on-demand allocation mechanism enables the FPGA chip to flexibly adjust its internal hardware resources, improve the utilization efficiency of hardware resources, avoid resource waste, and then efficiently provide hardware resources that dynamically adapt to the operation requirements of the deep learning model, significantly improving the performance and efficiency of the deep learning model inference or training process.

[0018] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the accompanying drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings. In the drawings: Figure 1 A flowchart showing a method for configuring an FPGA chip in an embodiment of the present application is shown; Figure 2 A block diagram showing a configuration device of an FPGA chip in an embodiment of the present application is shown; Figure 3 A schematic structural diagram of an electronic device in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0020] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0021] In addition, the described features, structures or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will realize that the technical solutions of the present application may be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be used. In other cases, well-known methods, devices, implementations or operations are not shown or described in detail to avoid obscuring aspects of the present application.

[0022] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices. It should also be noted that in the drawings, for the sake of simplicity of the drawings, some components that do not affect the explanation of the technical solutions of the present application are adaptively omitted.

[0023] The flowcharts shown in the drawings are only illustrative and do not necessarily include all the contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps may be decomposed, and some operations / steps may be combined or partially combined. Therefore, the actual execution order may be changed according to the actual situation.

[0024] In the description of the present application, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present application, unless otherwise specified, the meaning of "a plurality" is two or more.

[0025] To enable those skilled in the art to better understand the present application, the technical concepts and application backgrounds involved in the present application will be briefly described first.

[0026] FPGA Chip: An FPGA (Field Programmable Gate Array) chip is a programmable logic device. It consists of a large number of programmable logic units (such as look-up tables LUTs, flip-flops, etc.), programmable interconnect resources, and I / O interfaces. Users can program the FPGA through a hardware description language (such as VHDL or Verilog) to implement various digital logic functions.

[0027] Deep Learning Model: A deep learning model refers to an artificial neural network model composed of multiple (usually three or more) non-linear processing units. It can automatically extract and represent the features in data through learning a large amount of data, and achieve complex pattern recognition and prediction tasks. The deep learning model is a method in the field of machine learning and is widely used in fields such as image recognition, speech recognition, and natural language processing.

[0028] Transformer Module: The transformer module is a basic component unit in the deep learning model. It can be used to process sequential data and is widely used in fields such as natural language processing and computer vision.

[0029] Currently, with the rapid development of deep learning technology, the complexity and computing requirements of deep learning models are constantly increasing, and the hardware requirements for their supporting systems are also getting higher and higher. As the most important computing acceleration hardware in the field of deep learning, the GPU (Graphics Processing Unit) has become the preferred platform for training and inferring large-scale neural network models due to its powerful parallel computing ability and high-bandwidth memory access. However, once the GPU is designed or produced, its hardware architecture is fixed. Different deep learning models have different architectures, and there are also significant differences in the requirements for hardware resources. Therefore, when facing diverse deep learning tasks, a fixed-architecture GPU may not be able to achieve the optimal resource utilization rate, and may even result in computing power waste or computing power bottlenecks. In this case, this application proposes a configuration scheme for FPGA chips to efficiently provide hardware resources that can dynamically adapt to the computing requirements of deep learning models.

[0030] The following elaborates on the implementation details of the technical solution of the embodiments of this application: Refer to Figure 1 , which shows the flowchart of the configuration method of the FPGA chip in the embodiments of this application. The configuration method of this FPGA chip can be executed by a device with computing and processing capabilities. Refer to Figure 1 shown, the configuration method of this FPGA chip includes at least steps 110 to 130, which are introduced in detail as follows: In step 110, obtain the operation indices of each transformer module in the deep learning model, where the operation indices are used to characterize the amount of operations of the transformer module.

[0031] In this application, the transformer module can be a Transformer block (also known as TransformerBlock). In different deep learning models, the number of transformer modules is generally different. These differences mainly depend on factors such as the design goals, computing resources, expected applications, and performance requirements of the deep learning model. For example, the number of transformer modules in different versions of the large language model GPT-3 ranges from 12 to 96. Also, the large language model DeepSeek-7B contains 32 transformer modules, the large language model DeepSeek-33B contains 48 transformer modules, and the large language model DeepSeek-67B contains 80 transformer modules.

[0032] In each transformer module, there are generally multiple network modules. Among them, the network module is mainly composed of multiple parallel attention network modules (i.e., multi-head attention, Multi-Head Attention) and a feed-forward network (FeedForward Network, FFN). In the feed-forward network, multiple parallel expert network modules (Mixture of Experts, MOE) can generally be introduced to increase the model capacity. In addition, each transformer module usually also has auxiliary structures such as normalization and residual connections.

[0033] In this application, the number and type of network modules (such as self-attention, feed-forward network, gating mechanism, expert network, etc.) included in different transformer modules will vary due to different model design goals and implementation methods.

[0034] In this application, the operation exponent refers to the order-of-magnitude description of the amount of computation required by the transformer module during forward inference or training. It can be used to characterize the amount of computation of the transformer module, that is, to measure the complexity and computational overhead of the transformer module. For example, it can be expressed by FLOPs (Floating Point Operations, floating-point operation counts). Also, for example, it can be expressed by the number of multiply-accumulate operations (MACs). In addition, the operation exponent can also be expressed by any index that can reflect the computational load of the module, such as computational throughput requirements, or the number of parameters, or energy consumption requirements.

[0035] Continue to refer to Figure 1 , in step 120, obtain the total amount of programmable hardware resources in the FPGA chip, where the FPGA chip is used to execute the computational tasks of the deep learning model.

[0036] In this application, the programmable hardware resources in the FPGA chip mainly include look-up tables (LUTs), flip-flops (FFs), digital signal processing units (DSPs, DSP Slices / Blocks), on-chip memories (BRAMs, Block RAMs), and I / O resources. Among them, the look-up table is the basic unit for implementing any logic function, the flip-flop is used for sequential logic storage, the digital signal processing unit is used for efficient arithmetic operations such as multiplication and addition, the on-chip memory is used for data caching and storage, and the I / O resources include high-speed serial / parallel interfaces, etc., which mainly affect the external data exchange capability.

[0037] In this application, the total amount of programmable hardware resources in the FPGA chip can be obtained according to the specific model of the FPGA chip. For example, the total number of LUTs (i.e., the metric of the look-up table), the total number of FFs (i.e., the metric of the flip-flop), the total number of DSP units (i.e., the metric of the digital signal processing unit), the total capacity of BRAMs (i.e., the metric of the on-chip memory), and the number of I / Os (i.e., the metric of the I / O resources) can be obtained according to the specific model of the FPGA chip.

[0038] In this application, the FPGA chip is used as the hardware acceleration platform for the deep learning model, that is, the FPGA chip is used to implement the inference and / or training operations of the deep learning model.

[0039] Continue to refer to Figure 1 In step 130, according to the operation exponents of the respective converter modules and the total amount of hardware resources, a plurality of processing units are configured in the FPGA chip, and the processing units are adapted to the operation amount of the converter modules.

[0040] In this application, a plurality of processing units (PUs) can be dynamically configured inside the FPGA chip according to the operation exponents of the respective converter modules and the total amount of hardware resources of the FPGA chip, and the hardware resource scale of these processing units is adapted to the operation exponents of the respective converter modules. For example, converter modules with larger operation exponents are allocated more logic units (which can include look-up tables and flip-flops) and DSP units to configure the processing units, and converter modules with smaller operation exponents are allocated fewer hardware resources to configure the processing units, so as to achieve a reasonable allocation of hardware resources.

[0041] In this application, it should be noted that the processing units configured inside the FPGA chip are a set of hardware resources configured on the FPGA chip for specific converter modules and capable of independently or collaboratively completing specific computing tasks. Each processing unit may include resources such as lookup tables (LUTs), flip-flops (FFs), multipliers, registers, on-chip RAM, etc. These resources are combined and optimized according to the actual operation requirements of the converter module to achieve efficient data processing capabilities.

[0042] In this application, based on the repeatable configuration characteristic of the FPGA chip architecture, by obtaining the operation indices of each converter module in the deep learning model and combining with the total amount of programmable hardware resources in the FPGA chip, an efficient configuration of the FPGA chip is achieved. Compared with the method of using a chip with a fixed architecture to execute the deep learning model operation task in the prior art, the technical solution proposed in this application can give full play to the programmable advantages of the FPGA chip. According to the actual operation requirements of different converter modules, the processing units on the FPGA chip are dynamically configured, so that each converter module obtains hardware resources matching its operation volume. This on-demand allocation mechanism enables the FPGA chip to flexibly adjust the internal hardware resources, improve the utilization efficiency of the hardware resources, avoid resource waste, and then efficiently provide the deep learning model with hardware resources that dynamically adapt to its operation requirements, significantly improving the performance and efficiency of the deep learning model inference or training process. In addition, the technical solution proposed in this application also has good scalability and adaptability, and can support the deployment of deep learning models with different scales and structures on the same FPGA platform, greatly enhancing the application value of the FPGA chip in the field of intelligent computing.

[0043] In this application, in step 130 as shown in Figure 1 The step of configuring multiple processing units in the FPGA chip according to the operation indices of each converter module and the total amount of hardware resources can be executed according to the following steps 131 to 132: Step 131, allocate hardware resources to each converter module respectively according to the operation indices of each converter module and the total amount of hardware resources, wherein the hardware resource allocation amount of each converter module is positively correlated with the operation index of each converter module, and the sum of the hardware resource allocation amounts of each converter module is less than or equal to the total amount of hardware resources.

[0044] Step 132, configure multiple processing units in the FPGA chip based on the hardware resources allocated to each converter module.

[0045] In this application, the total amount of programmable hardware resources in the FPGA chip (such as the total number of logic units, i.e., the total number of LUTs and the total number of FFs) can be used as an overall constraint condition. According to the operation exponents of each converter module, weighted allocation, proportional allocation or other optimization algorithms are adopted to allocate corresponding amounts of hardware resources to each converter module respectively. Specifically, the higher the operation exponent of a converter module, the more hardware resources it is allocated to ensure that its high operation requirements can be met; the converter module with a lower operation exponent is allocated fewer hardware resources to achieve reasonable utilization of hardware resources. It should be noted that the sum of the hardware resources allocated to all modules cannot exceed the total amount of hardware resources of the FPGA chip, so as to avoid the failure of FPGA chip configuration or the degradation of FPGA chip performance caused by over-allocation of resources.

[0046] In this application, it should be noted that the sum of the hardware resource allocation amounts of each converter module can be less than the total amount of the hardware resources. Specifically, when allocating hardware resources to the converter module, the allocation amount is not necessarily the larger the better. Because if the allocation amount is too large, during the operation of the converter module, the allocated hardware resources cannot be fully utilized (i.e., the resource utilization rate is low), resulting in waste of hardware resources. Therefore, the most ideal hardware resource allocation amount should be: there are no idle resources during the operation of the converter module for the allocated hardware resources, and there is no queuing situation for the operation tasks in the converter module.

[0047] Furthermore, after completing the hardware resource allocation of each converter module, multiple processing units can be correspondingly configured inside the FPGA chip according to the amount and type of hardware resources allocated to each converter module. The amount and type of hardware resources of each processing unit (reflected in scale, composition and functional characteristics) are all matched with the operation requirements of the converter module it serves. For example, for a converter module with a higher operation exponent (i.e., a larger amount of operations), a larger-scale processing unit can be configured to enable it to have higher parallel computing capabilities and bandwidth support, while for a module with a smaller operation exponent (i.e., a smaller amount of operations), a smaller-scale or simplified processing unit is configured. This configuration process includes the specific mapping and connection of hardware resources such as internal logic units, DSP units, and storage units in the FPGA to ensure that each processing unit can efficiently execute the operation tasks of its corresponding converter module.

[0048] In this application, through the technical solutions in the above step 131 to step 132, the FPGA chip can dynamically construct a hardware acceleration structure with strong adaptability and high resource utilization rate for the operation requirements of different converter modules, thereby improving the overall operation efficiency and performance of the deep learning model.

[0049] In this application, when allocating hardware resources to each converter module according to the operation exponents of the respective converter modules and the total amount of hardware resources, the following steps 1311 to 1312 can be executed: Step 1311: Determine the total sum of the operation exponents of the respective converter modules, and determine the first proportion of the operation exponent of each converter module in the total sum of the operation exponents.

[0050] Step 1312: Allocate hardware resources to each converter module based on the first proportion corresponding to each converter module, where the absolute value of the difference between the second proportion of the hardware resource allocation amount of each converter module in the total amount of hardware resources and the first proportion is less than a preset proportion.

[0051] In this application, when analyzing the operation requirements of the converter modules, first, the total sum of the operation exponents of all converter modules can be counted. The total sum S is obtained by adding up the operation exponents of each module. Then, the first proportion of each converter module is calculated, and its formula (1) can be: where, represents the first proportion of the i th converter module; represents the operation exponent of the i th converter module; represents the total sum of the operation exponents of the respective converter modules.

[0052] It can be understood that the first proportion of the converter module can reflect the weight of the converter module in the overall operation load. Through the first proportion, the operation requirements of each converter module can be quantified, providing a basis for subsequent hardware resource allocation, ensuring that the hardware resource allocation matches the actual operation requirements, and thus avoiding waste or uneven distribution of hardware resources.

[0053] Furthermore, after knowing the first proportion of each converter module, the limited FPGA hardware resources (such as look-up tables LUTs and flip-flops FFs) can be allocated proportionally, that is, the proportion of the actual hardware resources allocated to each converter module (the second proportion) is as close as possible to the proportion of the operation requirements (the first proportion) of each converter module. Specifically, the hardware resources can be allocated to the converter module according to its first proportion, the second proportion (the amount of resources actually allocated / the total amount of hardware resources) is calculated, and the allocation accuracy is verified to ensure that the difference between the second proportion and the first proportion of each converter module is less than a preset proportion (i.e., the tolerance threshold, such as 1%). By controlling the deviation between the actual allocation proportion and the theoretical requirement proportion, it can be ensured that each module will neither be performance-limited due to insufficient hardware resources nor cause waste due to excessive hardware resources in the hardware resource allocation.

[0054] To enable those skilled in the art to better understand this application, the following will be described with reference to a specific embodiment.

[0055] Suppose there are three transformer modules A, B, and C in a deep learning model, and their floating-point operation requirements per second (i.e., operation exponents) are as follows: Transformer module A: requires 60,000 floating-point operations per second (FLOPS); Transformer module B: requires 30,000 floating-point operations per second (FLOPS); Transformer module C: requires 10,000 floating-point operations per second (FLOPS). The total number of lookup tables available for allocation on the FPGA chip is 99,999 LUT. The preset proportion is 1%.

[0056] First step, calculate the total operation exponent as 60,000 + 30,000 + 10,000 = 100,000 floating-point operations per second.

[0057] Second step, calculate the first proportion of each transformer module: The first proportion of transformer module A is 60,000 / 100,000 = 0.6; The first proportion of transformer module B is 30,000 / 100,000 = 0.3; The first proportion of transformer module C is 10,000 / 100,000 = 0.1.

[0058] Third step, allocate hardware resources (i.e., lookup tables) according to the first proportion of each transformer module: The number of LUTs allocated to transformer module A = 99,999 × 0.6 = 59,999.4 → rounded to 60,000; The number of LUTs allocated to transformer module B = 99,999 × 0.3 = 29,999.7 → rounded to 30,000; The number of LUTs allocated to transformer module C = 99,999 × 0.1 = 9,999.9 → rounded to 10,000. Since the sum of the hardware resource allocation amounts for each transformer module exceeds 1 LUT, adjustment is required. At this time, a resource allocation correction algorithm can be used to preferentially reduce transformer module A, which has the most LUTs allocated, so that the allocation sum is equal to 99,999, that is, the number of LUTs allocated to transformer module A is adjusted to 59,999.

[0059] Finally, since the second proportion of transformer module A is 59,999 / 99,999 ≈ 0.59999, and the absolute value of the difference |0.59999 - 0.6| = 0.00001 < 0.01, the allocation requirements are met.

[0060] In this application, it should be noted that the above examples are only used to help those skilled in the art better understand the content of this application. In fact, in the specific process of hardware resource allocation, relevant parameters such as the operation index, the amount of hardware resource allocation, and the number of converter modules may be more complex. This application does not use the above examples as a limitation of the protection scope, and any equivalent transformation or modification made within the spirit and principle of this application shall be covered within the protection scope of this application.

[0061] In this application, by calculating the sum of operation indices and the first ratio, and performing allocation according to the first ratio during hardware resource allocation, the operation requirements of each converter module can be scientifically quantified. In the actual process of hardware resource allocation, rounding and fine-tuning are used to meet this condition. In addition, by ensuring that the absolute value of the difference between the actual allocation ratio (the second ratio) of each converter module and its theoretical operation requirement ratio (the first ratio) is less than a preset tolerance threshold (such as 1%), the deviation between the actual allocation ratio and the theoretical ratio can be precisely controlled, ensuring the accuracy of hardware resource allocation and making the hardware resource allocation both fair and efficient.

[0062] In this application, based on the hardware resources allocated to each converter module, multiple processing units can be configured in the FPGA chip, which can be executed according to steps 1321 to 1322: Step 1321: Generate a first chip configuration file based on the hardware resources allocated to each converter module, where the first chip configuration file is used to describe the status and connection of the internal hardware resources of the FPGA chip.

[0063] Step 1322: Load the first chip configuration file into the FPGA chip to configure multiple processing units in the FPGA chip.

[0064] In this application, the configuration file can be used to describe the status and connection of the internal hardware resources of the FPGA chip. It can be a bitstream file (which is essentially a "bitstream" describing how each programmable switch (such as look-up table LUT, connection, register, logic cell / logic element, etc.) inside the FPGA is connected and operates) or a similar hardware description file.

[0065] It should be noted that the "status" mentioned above refers to the internal specific configuration of each programmable resource (such as LUT, register, etc.). For example, a 4-input LUT can implement 16 input combinations, and each combination can be configured to output 0 or 1, and these 16 bits constitute the "status" of this LUT.

[0066] It should also be noted that the "wiring" mentioned above refers to a large number of programmable switches (switches) and wire networks inside the FPGA, which can flexibly connect different units such as LUTs and registers together. The configuration file determines which units have signal paths and which do not.

[0067] The state and wiring of all programmable resources inside the FPGA refer to the specific functions (states) of each configurable functional unit (LUT, register, etc.) in the FPGA, and which paths are used to connect these units (wiring).

[0068] In this application, after generating the first chip configuration file, the generated first chip configuration file can be loaded into the FPGA chip through the configuration interface of the FPGA (such as JTAG, SPI, configuration flash, etc.), thereby completing the dynamic configuration of the processing units in the hardware resources. The loading process can be completed through a dedicated configuration tool or host computer software. When the first configuration file is successfully loaded, the hardware resources inside the FPGA chip complete the physical configuration and wiring according to the description of the configuration file, realizing multiple independent or cooperative processing units. These processing units can run in parallel and execute the functions and logics of their respective corresponding converter modules. Through this dynamic hardware configuration method, the parallel processing ability and reconfigurable characteristics of the FPGA chip can be fully utilized, thereby improving the overall performance of executing arithmetic tasks for deep learning models and the utilization rate of hardware resources.

[0069] In this application, each converter module includes multiple network modules. As mentioned above, the multiple network modules can include multiple parallel attention network modules and multiple parallel expert network modules.

[0070] Further, after the above step 1322, that is, after configuring multiple processing units in the FPGA chip, the following step 140 can also be executed: Step 140, configure processing subunits corresponding one-to-one to the network modules in each converter module in the processing unit corresponding to each converter module, and the processing subunits are used to execute the arithmetic tasks of the network modules.

[0071] In this application, inside each processing unit of the FPGA chip, a dedicated processing subunit can be configured for each network module according to the number and type of network modules inside the converter module. In this way, it can be ensured that the arithmetic tasks of each network module are independently completed by dedicated hardware resources, avoiding resource contention and improving the parallelism of deep learning model operations.

[0072] That is to say, necessary hardware resources (such as LUT, register, BRAM, DSP, etc.) can be allocated to each processing subunit, and the hardware structure of the processing subunit can be customized according to the operation characteristics of the corresponding network module (such as matrix multiplication, convolution, activation function, etc.) to achieve efficient operation acceleration of the network module.

[0073] In this application, the data flow path between processing subunits can ensure that the operation results can be efficiently transmitted between each processing subunit to achieve pipelined processing or parallel processing. Specifically, during the operation stage of the FPGA chip, each processing subunit can execute operation tasks (such as forward inference, feature transformation, etc.) on the network module it is responsible for according to the input data and control signals. The operation results can be directly output or transmitted to the next-level processing subunit or processing unit to form a complete data processing flow.

[0074] In this application, by further dividing each processing unit inside the transformer module into multiple processing subunits and corresponding one-to-one with the network module, finer-grained hardware acceleration and parallel processing can be achieved, thereby greatly improving the parallel operation ability of the FPGA chip, significantly enhancing the performance and efficiency of the FPGA chip in complex neural networks or multi-layer data processing tasks, and giving full play to the advantages of the reconfigurability and high parallelization of the FPGA chip. In addition, by further dividing the transformer module into multiple network modules and configuring independent processing subunits for each network module, the resource allocation of each processing subunit can also be flexibly adjusted according to the complexity of the actual network module to achieve the optimal utilization of hardware resources.

[0075] In this application, based on the hardware resources allocated to each transformer module, multiple processing units are configured in the FPGA chip, and processing subunits corresponding one-to-one with the network modules in each transformer module are configured in the processing unit corresponding to each transformer module, and the following steps 141 to 143 can be executed: Step 141, respectively determine the number of intermediate network modules included in each transformer module.

[0076] Step 142, generate a second chip configuration file based on the hardware resources allocated to each transformer module and the number of intermediate network modules included in each transformer module, and the second chip configuration file is used to describe the state and connection of the internal hardware resources of the FPGA chip.

[0077] Step 143, load the second chip configuration file into the FPGA chip to configure multiple processing units in the FPGA chip, and configure processing subunits corresponding one-to-one with the network modules in each transformer module in the processing unit corresponding to each transformer module.

[0078] In this application, first, the internal structure of each converter module can be analyzed in detail to determine the specific number of network modules it contains. By counting each converter module one by one, a correspondence relationship between the converter module and the number of its internal network modules is formed. Then, based on the number of network modules contained in each obtained converter module and the hardware resources allocated to each converter module, the hardware resource allocation scheme is further refined. Specifically, according to the above information, the hardware structure inside the FPGA chip can be designed and planned, including the layout of each processing unit and its internal processing subunits, as well as the connection relationship between them. Finally, the above hardware resource allocation and structure information are used to generate a second chip configuration file, which is used to comprehensively describe the configuration status and connection situation of the hardware resources inside the FPGA chip.

[0079] Finally, the generated second chip configuration file can be loaded into the FPGA chip. According to this configuration file, the FPGA chip automatically completes the configuration of multiple processing units, and inside the processing unit corresponding to each converter module, the processing subunits corresponding one by one to the network modules are further configured. Through the above process, a multi-level parallel hardware structure is formed inside the FPGA chip, thereby significantly improving the parallel processing ability and overall operation efficiency of the FPGA chip.

[0080] In an embodiment of this application, the training operation task of the deep learning model can be executed based on the above-configured FPGA chip.

[0081] After the configuration of the multi-level processing units and / or processing subunits of the FPGA chip is completed, the powerful parallel processing ability of the FPGA chip can be fully utilized to execute the training operation task of the deep learning model. The training process includes multiple stages such as forward propagation, loss calculation, backpropagation, and parameter update. The processing units at all levels inside the FPGA chip can respectively process different converter modules and the network modules in the converter modules in parallel, thereby realizing the efficient operation of each layer of the model.

[0082] In this embodiment, the programmable advantage of the FPGA chip can be fully utilized. According to the actual operation requirements of different converter modules, the processing units on the FPGA chip are dynamically configured, so that each converter module obtains the hardware resources matching its operation volume. In this way, the FPGA chip can flexibly adjust the internal hardware resources, improve the hardware resource utilization efficiency, avoid resource waste, and then efficiently dynamically adapt the hardware resources required for its operation to the deep learning model, significantly improving the performance and efficiency of the deep learning model training process and reducing the model training cost.

[0083] Taking the training of the DeepSeekR1 model with the FPGA chip configured by the technical solution of this application as an example, the utilization rate of the logic units in the chip is increased from 40% to 75%, the utilization rate of BRAM is increased from 50% to 80%, the training throughput is increased to 1.2 TFLOPS, the reuse rate of general IP cores reaches 90%, and the development cycle is shortened from 6 months to 2 weeks.

[0084] In another embodiment of this application, the inference operation task of the deep learning model can also be executed based on the configured FPGA chip.

[0085] After the configuration of the multi-level processing units and / or processing subunits of the FPGA chip is completed, the FPGA chip can also efficiently execute the inference (inferencing) operation task of the deep learning model. The inference stage mainly includes the forward propagation of input data and the result output. At this time, the FPGA chip has configured the corresponding processing units and processing subunits according to the model structure, and can give full play to the hardware parallelism to batch process or pipeline process the input data, greatly improving the throughput and real-time performance of the inference.

[0086] It can be seen that based on the technical solution proposed in this application, according to the actual operation requirements of the deep learning model, by dynamically configuring the hardware resources in the FPGA chip, whether it is to execute the training task of the deep learning model or the inference task, the FPGA chip can rely on its flexible reconfigurable and highly parallel hardware architecture to provide strong support for the efficient operation of the deep learning model. For example, in the model training stage, more computing resources can be configured to accelerate the model training speed; in the model inference stage, more streamlined resources can be configured to reduce power consumption and costs, thus effectively solving the problems of hardware resource waste and performance bottlenecks existing in the prior art.

[0087] The following introduces the device embodiments of this application, which can be used to execute the configuration method of the FPGA chip in the above embodiments of this application. For the details not disclosed in the device embodiments of this application, please refer to the embodiments of the configuration method of the FPGA chip in the above of this application.

[0088] See Figure 2 , which shows the block diagram of the configuration device of the FPGA chip in the embodiments of this application.

[0089] As Figure 2 shown, the configuration device 200 of the FPGA chip according to the embodiments of this application includes: a first acquisition unit 201, a second acquisition unit 202, and a configuration unit 203.

[0090] Among them, the first acquisition unit 201 is configured to acquire the operation exponents of each transformer module in the deep learning model, where the operation exponents are used to characterize the amount of operations of the transformer module; the second acquisition unit 202 is configured to acquire the total amount of programmable hardware resources in the FPGA chip, and the FPGA chip is used to execute the operation task of the deep learning model; the configuration unit 203 is configured to configure a plurality of processing units in the FPGA chip according to the operation exponents of each transformer module and the total amount of hardware resources, and the processing units are adapted to the amount of operations of the transformer module.

[0091] In some embodiments of the present application, based on the foregoing solution, the configuration unit 203 is configured to: allocate hardware resources to each transformer module according to the operation exponents of each transformer module and the total amount of hardware resources, where the amount of hardware resource allocation for each transformer module is positively correlated with the operation exponent of each transformer module, and the sum of the amounts of hardware resource allocation for each transformer module is less than or equal to the total amount of hardware resources; based on the hardware resources allocated to each transformer module, configure a plurality of processing units in the FPGA chip.

[0092] In some embodiments of the present application, based on the foregoing solution, the configuration unit 203 is configured to: determine the total sum of the operation exponents of each transformer module, and determine the first proportion of the operation exponent of each transformer module in the total sum of the operation exponents; based on the first proportion corresponding to each transformer module, allocate hardware resources to each transformer module, where the absolute value of the difference between the second proportion of the amount of hardware resource allocation for each transformer module in the total amount of hardware resources and the first proportion is less than a preset proportion.

[0093] In some embodiments of the present application, based on the foregoing solution, the configuration unit 203 is configured to: generate a first chip configuration file based on the hardware resources allocated to each transformer module, where the first chip configuration file is used to describe the status and connection of the internal hardware resources of the FPGA chip; load the first chip configuration file into the FPGA chip to configure a plurality of processing units in the FPGA chip.

[0094] In some embodiments of the present application, based on the foregoing solution, each transformer module includes a plurality of network modules, and the configuration unit 203 is further configured to: after configuring a plurality of processing units in the FPGA chip, configure processing subunits corresponding one-to-one to the network modules in each transformer module in the processing unit corresponding to each transformer module, and the processing subunits are used to execute the operation tasks of the network modules.

[0095] In some embodiments of the present application, based on the foregoing solution, the configuration unit 203 is further configured to: respectively determine the number of medium network modules included in each converter module; generate a second chip configuration file based on the hardware resources allocated to each converter module and the number of medium network modules included in each converter module, where the second chip configuration file is used to describe the state and connection of the internal hardware resources of the FPGA chip; load the second chip configuration file into the FPGA chip to configure a plurality of processing units in the FPGA chip, and configure processing subunits corresponding one-to-one to the medium network modules in each converter module in the processing unit corresponding to each converter module.

[0096] In some embodiments of the present application, based on the foregoing solution, the converter module includes a Transformer block, and the plurality of network modules include a plurality of parallel attention network modules and a plurality of parallel expert network modules.

[0097] In some embodiments of the present application, based on the foregoing solution, the device further includes: an operation module, configured to execute a training operation task of the deep learning model through the FPGA chip; or execute an inference operation task of the deep learning model through the FPGA chip.

[0098] Based on the same inventive concept, an embodiment of the present application provides a computer program product, where the computer program product includes computer instructions, and the computer instructions are stored in a computer-readable storage medium and are adapted to be read and executed by a processor so that a computer device having the processor executes to implement the operations performed by the configuration method of the FPGA chip as described above.

[0099] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium, where at least one computer program instruction is stored in the computer-readable storage medium, and the at least one computer program instruction is loaded and executed by a processor to implement the operations performed by the configuration method of the FPGA chip as described above.

[0100] Based on the same inventive concept, an embodiment of the present application further provides an electronic device. Refer to Figure 3 , which shows a schematic structural diagram of the electronic device in the embodiment of the present application. The electronic device includes one or more memories 304, one or more processors 302, and at least one computer program (computer program instruction) stored on the memory 304 and executable on the processor 302. When the processor 302 executes the computer program, it implements the configuration method of the FPGA chip as described above.

[0101] Among them, in Figure 3Among them, a bus architecture (represented by bus 300) may include any number of interconnected buses and bridges. Bus 300 links together various circuits including one or more processors represented by processor 302 and a memory represented by memory 304. Bus 300 may also link together various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and thus will not be further described herein. Bus interface 305 provides an interface between bus 300 and receiver 301 and transmitter 303. Receiver 301 and transmitter 303 may be the same element, i.e., a transceiver, providing a unit for communicating with various other devices over a transmission medium. Processor 302 is responsible for managing bus 300 and general processing, while memory 304 may be used to store data used by processor 302 when performing operations.

[0102] The functions described herein may be implemented in hardware, software executed by a processor, firmware, or any combination thereof. If implemented in software executed by a processor, the functions may be stored on or transmitted via a computer-readable medium as one or more instructions or codes. Other examples and implementations are within the scope and spirit of the present application and the appended claims. For example, due to the nature of software, the functions described above may be implemented using software executed by a processor, hardware, firmware, hardwiring, or any combination thereof. In addition, each functional unit may be integrated in one processing unit, may exist separately physically as individual units, or two or more units may be integrated in one unit.

[0103] In several embodiments provided in the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units can be a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces, and the indirect couplings or communication connections of units or modules can be in electrical or other forms.

[0104] The units described as separate components may or may not be physically separated. The components serving as control devices may or may not be physical units, i.e., they may be located in one place or distributed over multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0105] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs that can store computer program instructions.

[0106] The above are only the embodiments of this application and are not used to limit this application. For those skilled in the art, this application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this application shall be included within the scope of the claims of this application.

Claims

1. A configuration method for an FPGA chip, characterized in that, The method includes: Obtaining the operation exponents of each transformer module in the deep learning model, where the operation exponents are used to characterize the amount of operations of the transformer module; Obtaining the total amount of programmable hardware resources in the FPGA chip, where the FPGA chip is used to execute the operation tasks of the deep learning model; Configuring a plurality of processing units in the FPGA chip according to the operation exponents of each transformer module and the total amount of hardware resources, where the processing units are adapted to the amount of operations of the transformer module.

2. The method according to claim 1, characterized in that, The configuring a plurality of processing units in the FPGA chip according to the operation exponents of each transformer module and the total amount of hardware resources includes: Allocating hardware resources to each transformer module respectively according to the operation exponents of each transformer module and the total amount of hardware resources, where the amount of hardware resource allocation for each transformer module is positively correlated with the operation exponent of each transformer module, and the sum of the amounts of hardware resource allocation for each transformer module is less than or equal to the total amount of hardware resources; Configuring a plurality of processing units in the FPGA chip based on the hardware resources allocated to each transformer module.

3. The method according to claim 2, wherein The allocating hardware resources to each transformer module respectively according to the operation exponents of each transformer module and the total amount of hardware resources includes: Determining the total sum of the operation exponents of each transformer module, and determining the first proportion of the operation exponent of each transformer module in the total sum of the operation exponents; Allocating hardware resources to each transformer module based on the first proportion corresponding to each transformer module, where the absolute value of the difference between the second proportion of the amount of hardware resource allocation for each transformer module in the total amount of hardware resources and the first proportion is less than a preset proportion.

4. The method according to claim 2, wherein The configuring a plurality of processing units in the FPGA chip based on the hardware resources allocated to each transformer module includes: Generating a first chip configuration file based on the hardware resources allocated to each transformer module, where the first chip configuration file is used to describe the status and connection of the internal hardware resources of the FPGA chip; Loading the first chip configuration file into the FPGA chip to configure a plurality of processing units in the FPGA chip.

5. The method according to claim 2, characterized in that Each transformer module includes a plurality of network modules. After configuring a plurality of processing units in the FPGA chip, the method further includes: Configuring processing subunits corresponding one-to-one to the network modules in each transformer module in the processing unit corresponding to each transformer module, where the processing subunits are used to execute the operation tasks of the network modules.

6. The method according to claim 5, characterized in that, The configuring a plurality of processing units in the FPGA chip based on the hardware resources allocated to each transformer module, and configuring processing subunits corresponding one-to-one to the network modules in each transformer module in the processing unit corresponding to each transformer module includes: Respectively determining the number of network modules included in each transformer module; Generate a second chip configuration file based on the hardware resources allocated to each converter module and the number of middle network modules included in each converter module, where the second chip configuration file is used to describe the status and wiring of the internal hardware resources of the FPGA chip; Load the second chip configuration file into the FPGA chip to configure multiple processing units in the FPGA chip and configure processing subunits corresponding one-to-one to the network modules in each converter module in the processing unit corresponding to each converter module.

7. The method according to claim 6, characterized in that, The converter module includes a Transformer block, and the multiple network modules include multiple parallel attention network modules and multiple parallel expert network modules.

8. The method according to any one of claims 1 to 7, characterized in that, The method further includes: Performing a training operation task on the deep learning model through the FPGA chip; or, Performing an inference operation task on the deep learning model through the FPGA chip.

9. A configuration device for an FPGA chip, characterized in that, The device includes: A first acquisition unit, configured to acquire the operation indices of each converter module in the deep learning model, where the operation indices are used to characterize the amount of operations of the converter module; A second acquisition unit, configured to acquire the total amount of programmable hardware resources in the FPGA chip, where the FPGA chip is used to perform the operation task on the deep learning model; A configuration unit, configured to configure multiple processing units in the FPGA chip according to the operation indices of each converter module and the total amount of hardware resources, where the processing units are adapted to the amount of operations of the converter module.

10. A computer program product, characterized in that, The computer program product includes computer instructions, which are stored in a computer-readable storage medium and are suitable for being read and executed by a processor, so that a computer device having the processor executes the method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, At least one program code is stored in the computer-readable storage medium, and the at least one program code is loaded and executed by a processor to implement the operations performed by the method according to any one of claims 1 to 8.

12. An electronic device, characterized in that, The electronic device includes one or more processors and one or more memories, and at least one program code is stored in the one or more memories, and the at least one program code is loaded and executed by the one or more processors to implement the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • FPGA-based secure multi-party computing machine learning resource configuration method and system

    CN118277092A

  • Method and device for optimizing reasoning resources and electronic equipment

    CN118796471A

  • Resource allocation method, device and equipment for large model cluster, storage medium and program product

    CN119597368A

  • Electronic device for performing token pruning in frequency domain and method for operating the same

    US20240214205A1