Artificial intelligence device based on FPGA
By deploying large models using FPGA-based AI devices, the risk of data leakage in fields with strict privacy requirements is addressed, enabling efficient, low-power local AI functionality suitable for mobile platforms.
Patent Information
- Application Number
- CN202421477821.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Utility models(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-25
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2034-06-25
AI Technical Summary
Existing technologies are difficult to deploy large models locally in areas with strict privacy requirements, posing risks of data leakage and issues with device efficiency and power consumption.
The system employs FPGA-based artificial intelligence devices, which implement local AI functions by deploying large-scale FPGA chips. Combined with MCUs and external memory, it provides interfaces and heat dissipation components to ensure data security and efficient operation.
It enables local deployment of large models, avoids data leakage, improves device efficiency and reduces power consumption, making it suitable for use as a mobile platform.
Smart Images

Figure CN223679641U_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The utility model relates to FPGA technical field especially relates to an artificial intelligence device based on FPGA. BACKGROUND
[0002] Large Model, also known as Foundation Model, refers to a machine learning model with a large number of parameters and complex structure, which can process massive data and complete various complex tasks.
[0003] For example, with the development of large model technology, it has been widely used in the field of natural language text generation. Currently, text is usually generated online based on cloud services, but for some fields, privacy requirements are high. For example, law firms have very strict privacy requirements, and client case details need to be strictly confidential, so most law firms require not to use online large models to assist in generating text content such as case descriptions, to prevent potential leakage risks caused by uploading clients' privacy to public networks.
[0004] Therefore, it is urgent to provide an artificial intelligence device that can realize local deployment of large models, while considering the efficiency and power consumption of the device. UTILITY MODEL CONTENT
[0005] In view of the above deficiencies of the prior art, the purpose of the utility model is to provide an artificial intelligence device based on FPGA (Field Programmable Gate Array) to realize corresponding AI functions through locally deployed large models and ensure reasonable power consumption and efficiency.
[0006] In order to achieve the above purpose, the utility model adopts the following technical solutions:
[0007] An artificial intelligence device based on FPGA, comprising a shell and a circuit board arranged in the shell, wherein:
[0008] The circuit board comprises an AI component and an external memory connected to the AI component, and the AI component comprises an FPGA chip deployed with a preset large model.
[0009] The shell is provided with an interface component connected to the AI component and a human-computer interaction component.
[0010] Further, the AI component further comprises an MCU connected to the FPGA chip.
[0011] Alternatively, the FPGA chip is integrated with an MCU.
[0012] Further, the AI component comprises a plurality of the FPGA chips, and one of the FPGA chips is a master FPGA chip, and the rest of the FPGA chips are slave FPGA chips.
[0013] The MCU is connected with the master FPGA chip, or the MCU is integrated in the master FPGA chip.
[0014] Further, the external memory comprises an SD card and / or a DDR memory.
[0015] Further, a card slot for mounting the SD card is further arranged on the circuit board, and the card slot is connected with the pins of the FPGA chip through pin multiplexing.
[0016] Further, the interface component comprises any one or more of the following interfaces: an Ethe interface, a JTAG interface, a UART interface, an SPI interface, an I2C interface, an RJ45 local area network interface and a power supply interface.
[0017] Further, the shell is provided with a heat dissipation hole.
[0018] Further, the shell is further provided with a heat dissipation component for heat dissipation.
[0019] Further, the heat dissipation component comprises an exhaust fan and a heat dissipation fin arranged on the circuit board.
[0020] Further, the human-computer interaction component comprises a switch button and a reset button.
[0021] By adopting the above technical scheme, the artificial intelligence device has the following beneficial effects:
[0022] The artificial intelligence device of the utility model deploys the FPGA chip with the preset large model, and can realize the corresponding AI function in an offline form based on the FPGA chip, without using an online form, thereby ensuring the safety of data; when the large model version is updated, the latest large model can be redeployed in the FPGA chip, without using the online form to update the model, and further avoiding possible data leakage. Meanwhile, the FPGA is high in efficiency, low in power consumption and cost, and convenient to carry, and can be used as a mobile platform. BRIEF DESCRIPTION OF DRAWINGS
[0023] Figure 1 Fig. 1 is a structural schematic view of an artificial intelligence device based on FPGA in the utility model embodiment 1;
[0024] Figure 2 Fig. 1 is a structural schematic view of an artificial intelligence device based on FPGA in the utility model embodiment 1;
[0025] Figure 3 The circuit principle diagram of the artificial intelligence device based on FPGA in the embodiment 2 of the present application;
[0026] Figure 4 The circuit principle diagram of the artificial intelligence device based on FPGA in the embodiment 3 of the present application. DETAILED DESCRIPTION
[0027] In order to make the purpose, technical scheme and advantages of the present application more clear, the present application will be further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of the present application.
[0028] The terms used in the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present disclosure. The singular forms "a", "said" and "the" used in the present disclosure and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein means and includes any or all possible combinations of one or more associated listed items.
[0029] Embodiment 1
[0030] The present embodiment provides an artificial intelligence device based on FPGA, as shown in Figure 1 and 2 , the device comprises a shell 1 and a circuit board (not shown) arranged in the shell 1. Wherein, the circuit board comprises an AI component and an external memory connected with the AI component, and the core component of the AI component is a FPGA chip deployed with a preset large model, which is used to generate an AI processing result corresponding to user input information based on the large model.
[0031] Preferably, the large model in the present embodiment can be a multi-modal large model, which can be used in natural language processing, computer vision, multimedia processing and other scenarios. The FPGA in the present embodiment has higher efficiency compared with deploying a large model using a CPU chip; has a higher and reasonable power ratio compared with GPU, can exist independently without relying on a computer case; and has higher flexibility compared with an ASIC chip, can follow the latest algorithm changes, and can be customized to adapt to various underlying logic. Therefore, deploying a local large model using FPGA as a hardware platform is a suitable technical solution for law firm scenarios and has more upgrade potential. At the same time, it is convenient to carry and can be used as a mobile platform.
[0032] As the core component of a circuit board, FPGA chips can be selected with HBM (High Bandwidth Memory) or without HBM, depending on cost and power consumption considerations.
[0033] This embodiment deploys a trained large model on an FPGA chip, enabling offline implementation of AI functions without the need for online updates, thus ensuring data security. When the large model version is updated, the latest model can be redeployed on the FPGA chip without requiring online updates, further preventing potential data leaks. Furthermore, FPGAs offer high efficiency, low power consumption and cost, and are portable, making them suitable for use as mobile platforms.
[0034] In this embodiment, an MCU can be integrated into the FPGA chip to implement basic process control for the corresponding AI functions. Specifically, an MCU can be instantiated in the FPGA for top-level control.
[0035] See again Figure 2 As shown, the FPGA chip in this embodiment also includes, but is not limited to, a state machine control module, a data processing module, internal memory, and an external memory access module. The state machine control module controls the state of the FPGA chip, and the data processing module performs corresponding data processing based on the deployed large model. Furthermore, due to the limited storage space of the internal memory in the FPGA chip, this embodiment uses external memory to store intermediate data and various parameters of the large model. Therefore, the FPGA needs to access the external memory through the external memory access module to interact with parameters between different layers of the large model.
[0036] In an optional implementation, the external storage may include an SD card or DDR memory. A card slot for installing the SD card may be provided on the circuit board, wherein the SD card uses a 64-card parallel transmission method to maximize the access bandwidth of the SDIO interface; the SD card slot is connected to the FPGA chip pins through pin multiplexing to reduce FPGA chip pin consumption and accelerate data access. The DDR memory serves as a backup data storage medium for storing temporary parameters, such as a KV stack.
[0037] In this embodiment, simply formatting all SD cards and refreshing the internal data of the FPGA chip is sufficient to update to the latest large model, allowing it to keep pace with the latest algorithmic advancements. Online model updates are not used, thus avoiding potential data leaks.
[0038] In this embodiment, as Figure 1As shown, the shell 1 preferably adopts a plastic hollow square box (for example, the size is 30*20*10cm), and the hollow structure constitutes the heat dissipation hole 11, and the shell 1 can also be provided with a heat dissipation assembly (not shown) for heat dissipation, which includes but is not limited to a built-in exhaust fan and a heat sink arranged on the circuit board adjacent to the FPGA chip.
[0039] In addition, the shell 1 of the embodiment is provided with an interface assembly 12 in communication connection with the AI assembly and a human-computer interaction assembly 13. Among them, the interface assembly 12 includes but is not limited to any one or more of the following interfaces: Ethe interface, JTAG interface, UART interface, SPI interface, I2C interface, RJ45 local area network interface and power interface, etc. The human-computer interaction assembly 13 includes but is not limited to any one or more of the following components: switch button, reset button and display screen, etc.
[0040] Again refer to Figure 2 As shown, the device of the embodiment can be connected with the host computer through the interface assembly to input user instructions through the host computer and feed back the AI processing result to the host computer. Among them, the UART interface or the RJ45 local area network interface is preferably used for communication between the device and the host computer, the interface is more bottom layer, the response is faster, and the delay is lower.
[0041] In addition, the device of the embodiment is not configured with the PCIE interface on the general graphics card, and cannot be directly connected with the computer to prevent the device from being inserted into the computer case and data leakage.
[0042] The device of the embodiment can be used not only for assisting lawyers' offices to assist in generating related legal documents, but also can be applied to AI part implementation in private equity quantitative fund strategy, small military equipment requiring offline deployment of large models, intelligent household appliances, wearable devices, and large model deployment in scenes requiring offline or low-power large models, and is convenient to carry, and can be used as a mobile platform.
[0043] Embodiment 2
[0044] Considering that the complexity consumed by the transformer architecture used by the large model is much higher than the complexity required for transmission, the embodiment fully utilizes the advantages of multiple FPGA chips to linearly improve the operation efficiency by increasing the number of chips.
[0045] Specifically as Figure 3 As shown, the embodiment increases the number of FPGA chips in the AI assembly based on embodiment 1, and each FPGA chip is respectively provided with a data processing module for performing corresponding data processing based on the large model, to generate the AI processing result corresponding to the control instruction based on the large model.
[0046] In the embodiment, one of the FPGA chips is a master FPGA chip, and the rest of the FPGA chips are slave FPGA chips. The master FPGA chip and the slave FPGA chips can be connected in communication through a SerDes interface, an Ethernet interface, an optical communication interface, or any other suitable interface, to exchange data between transformer layers, provide the possibility of multi-chip parallel computing, and thus improve the computing capability. The SerDes is a short form of SERializer (serializer) / DESerializer (deserializer).
[0047] In the embodiment as shown in Figure 2 In the embodiment as shown in
[0048] Each FPGA chip includes an internal memory, is connected with an external memory, and includes an external memory access module for accessing the corresponding external memory.
[0049] The specific operation process of the artificial intelligence device of the embodiment is as follows:
[0050] S1. Turn on the power supply and perform hardware reset.
[0051] S2. Burn and write the internal logic bitstream of each FPGA chip to each FPGA chip through the host computer.
[0052] S3. Transfer all parameters of the model to the DDR or SD card and other external memories through the host computer.
[0053] S4. The host computer gives an instruction to the master FPGA chip, instructing all the FPGAs to read the corresponding parameter data from the external memory and cache them into the internal memory BRAM. If the model size is larger than the space of the internal memory, as many contents as possible are read.
[0054] S5. All the preparation work is completed.
[0055] S6. The user transmits user input information to the host computer through a serial port or other interface.
[0056] S7. The host computer performs preprocessing calculation, and codes the user input information such as text, image, video, etc. into a form readable by the model, and transmits it to the main FPGA chip.
[0057] S8. The main FPGA chip allocates the calculation tasks in the transformer layer operator to each FPGA chip according to the model size parameters, including itself, and then each FPGA chip performs calculation layer by layer.
[0058] S9. After all the calculations are completed, the results are returned to the host computer.
[0059] S10. The host computer decodes the ID list output by each FPGA chip into the expected information carrier such as text, image, video, etc. and displays it on the external device.
[0060] Compared with Embodiment 1, the operation efficiency of the large model is improved.
[0061] It should be noted that the number of FPGA chips in this embodiment is not limited, and can be 2, 3 or more, depending on the actual size of the large model. Specifically, the number of FPGA chips is optimal when its internal memory can cache all model parameters as much as possible.
[0062] Embodiment 3
[0063] As shown in Figure 4 The difference between this embodiment and Embodiment 1 is that the MCU in the AI component is not integrated in the FPGA chip, but is implemented separately using a chip (such as a CPU chip, etc.). The chip is connected to the FPGA chip and is used for process control, thereby achieving the same technical effect as Embodiment 1.
[0064] Although the specific embodiments of the present application are described above, those skilled in the art should understand that this is only an example, and the protection scope of the present application is defined by the appended claims. Those skilled in the art can make various changes or modifications to these embodiments without departing from the principles and essence of the present application, and such changes and modifications fall within the protection scope of the present application.
Claims
1. An FPGA-based artificial intelligence device, characterized by comprising: The application relates to a device for deploying a large model, comprising a shell and a circuit board arranged in the shell, wherein: the circuit board comprises an AI component and an external memory connected with the AI component, and the AI component comprises an FPGA chip in which a preset large model is deployed; the AI component further comprises an MCU connected with the FPGA chip; or the FPGA chip is integrated with an MCU; the AI component comprises a plurality of FPGA chips, one of which is a master FPGA chip and the rest are slave FPGA chips; the MCU is connected with the master FPGA chip or the MCU is integrated in the master FPGA chip; the external memory is used for storing intermediate data and parameters of a large model, so that the FPGA chip can call the intermediate data and parameters through an external storage access module to realize parameter interaction between different layers of the large model; the shell is provided with an interface component connected with the AI component and a human-computer interaction component. 2.The artificial intelligence device of claim 1, wherein, The external memory comprises an SD card and / or a DDR memory. 3.The artificial intelligence device of claim 2, wherein, The circuit board is further provided with a card slot for mounting the SD card, and the card slot is connected with pins of the FPGA chip through pin multiplexing. 4.The artificial intelligence device of claim 1, wherein, The interface component comprises any one or more of the following interfaces: an Ethe interface, a JTAG interface, a UART interface, an SPI interface, an I2C interface, an RJ45 local area network interface and a power supply interface. 5.The artificial intelligence device of claim 1, wherein, The shell is provided with a heat dissipation hole. 6.The artificial intelligence device of claim 1, wherein, The shell is further provided with a heat dissipation component for heat dissipation.
7. The artificial intelligence device of claim 6, wherein, The heat dissipation component comprises an exhaust fan and a heat dissipation fin arranged on the circuit board. 8.The artificial intelligence device of claim 1, wherein, The human-computer interaction component comprises a switch button and a reset button.