AI computing subsystem, end-side device and control method of end-side device

By designing an AI computing subsystem on mobile phones, adopting near-access computing architecture and advanced packaging technology, the generation rate and power consumption problems of traditional mobile phone SoC systems when running large models are solved, and efficient and low-power AI computing capabilities are achieved.

CN120045512APending Publication Date: 2025-05-27HANG ZHOU NANO CORE CHIP ELECTRONIC TECH CO LTD
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510126185.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-27
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

When running large models in traditional mobile SoC systems, the generation rate is difficult to meet user needs, and each calculation requires frequent transmission of model weight parameters and intermediate layer parameters, resulting in high power consumption and heat dissipation problems.

Method used

An AI computing subsystem is designed, using a near-memory computing architecture, combining FLASH memory chips and DRAM memory chips, and data reading and computing are performed through logic chips, and bandwidth between chips is increased using three-dimensional stacking and Hybrid Bonding processes.

Benefits of technology

It improves the token generation rate of AI models on mobile phones, reduces the power consumption of data handling, extends the battery life of the device, reduces the difficulty of cooling design, and supports multiple models to run simultaneously and switch quickly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045512A_ABST
    Figure CN120045512A_ABST
Patent Text Reader

Abstract

The invention relates to an AI computing subsystem, end-side equipment and a control method of the end-side equipment, and the AI computing subsystem comprises at least one FLASH storage chip which is used for pre-storing weight parameters of at least one AI network model; the at least one DRAM storage chip is used for storing weight parameters and activation functions of the AI model; and the logic chip is connected with the FLASH storage chip and the DRAM storage chip, responds to an external instruction, reads data of the FLASH storage chip so as to write the data into the DRAM storage chip, and / or reads weight parameters and activation functions stored in the DRAM storage chip and executes a calculation task. The SoC system only needs to send parameters or start commands of the AI network, and the whole network runs in the AI computing subsystem; and the most complex part of the AI network part can be deployed in the AI computing subsystem, and the SoC system only needs to compute the simple operator of the part, so that the SoC system has larger space to complete other responsible tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence chips, and specifically relates to an AI computing subsystem, an edge device, and a control method for the edge device. Background Art

[0002] With the continuous development of artificial intelligence technology, AI technology has penetrated into all aspects of our daily lives; among them, large models, with their powerful language understanding, generation, and complex task processing capabilities, have become the key force driving the intelligent transformation of various industries. Large models such as GPT series, Baidu Wenxin Yiyan, Alibaba Tongyi Qianwen, etc., have demonstrated remarkable performance in many aspects such as natural language processing, image recognition, and intelligent question answering. Smartphones, as the most commonly used and indispensable mobile devices in people's lives, carry more and more diverse functional requirements. From daily social chatting, photo beautification, to mobile office, assisted learning, etc., users expect mobile phones to provide a more intelligent, efficient, and personalized service experience. Traditional mobile phone intelligent applications based on small-scale models or simple algorithms have gradually become difficult to meet users' expectations for complex task processing, accurate content generation, and highly intelligent interaction. Therefore, it is an irresistible trend to run large model networks on mobile phones to provide more intelligent services. Due to the transmission rate limitation between the SoC system and the DRAM storage chip on mobile phones, if a large model is directly run on the mobile phone SoC system, the generation rate of the large model will be difficult to meet the actual needs of users. At the same time, for each calculation, the weight parameters and some intermediate layer parameters of the large model need to be transmitted between the SoC system and the DRAM storage chip, and the power consumption generated will be difficult to meet the customer requirements of the mobile terminal; and the excessive power consumption will bring higher heat dissipation requirements to the limited space on the mobile phone, which will further increase the design requirements for the mobile phone.

[0003] Therefore, it is necessary to improve the existing edge devices. Summary of the Invention

[0004] To solve the above technical problems, the present invention provides an AI computing subsystem, including: at least one FLASH storage chip for pre-storing the operation parameters of at least one AI network model; at least one DRAM storage chip for storing the operation parameters of the AI network model; and a logic chip connected to the FLASH storage chip and the DRAM storage chip, which, in response to an external instruction, reads the data of the FLASH storage chip for writing into the DRAM storage chip, and / or reads the data stored in the DRAM storage chip and executes a calculation task.

[0005] Among them, the operating parameters of the AI network model include: weight parameters, initialization configuration parameters, startup parameters of the system, and other parameters or data that can be applied to the operation of the AI network model.

[0006] Optionally, the DRAM memory chip and the logic chip are packaged by a three-dimensional stacking process and connected by a Hybrid Bonding process.

[0007] Optionally, the FLASH memory chip and the logic chip are connected by a Hybrid Bonding, Wire Bonding, or PCB Wires process.

[0008] Optionally, when the number of DRAM memory chips is greater than one, the DRAM memory chips are packaged by a three-dimensional stacking process and connected by a Hybrid Bonding process.

[0009] To achieve the above invention purpose, the present application provides an edge device, including at least one of the above-mentioned AI computing subsystems and a SoC system connected to each of the AI computing subsystems.

[0010] Optionally, the SoC system and the AI computing subsystem are connected through a DDR interface.

[0011] Optionally, the protocol of the DDR interface can be any one of the sub-protocols of LPDDR, Normal DDR, and GDDR.

[0012] Optionally, the DDR interface is provided on the DRAM memory chip or the logic chip.

[0013] Optionally, when the number of AI computing subsystems is multiple, the AI computing subsystems are configured such that each AI computing subsystem is connected to the SoC system through an independent DDR interface.

[0014] Optionally, when the number of AI computing subsystems is multiple, the AI computing subsystems are configured such that at least some or all of the AI computing subsystems share a DDR interface to be connected to the SoC system.

[0015] Optionally, when the DDR interface is provided on the DRAM memory chip, each DRAM memory chip is provided with one DDR interface and each DDR interface is connected to the SoC system.

[0016] To achieve the above-mentioned invention objectives, the present application provides a control method for an edge device, which is applied to control the edge device described above, and includes the following steps: Initialization step: The SoC system prepares initialization data and sends it to the logic chip of the AI computing subsystem, and the logic chip performs initialization configuration based on the initialization data; Calculation step: The logic chip reads the data required for calculation stored in the DRAM memory chip based on the calculation instruction, and writes the intermediate result data and the final result data of the calculation back to the DRAM memory; and

[0017] Calculation result reading step: The SoC system reads the final result data.

[0018] Optionally, the initialization step further includes: determining whether it is necessary to update the operating parameters; if so, the SoC system sends a command to the AI computing subsystem to update the operating parameters of the DRAM memory chip, and in response to this command, the logic chip can read the required operating parameters from the FLASH memory chip and write them into the DRAM memory chip.

[0019] Optionally, the calculation step further includes: The SoC system writes the activation feature data into the DRAM memory chip through the DDR interface provided in the DRAM memory chip, and determines whether there are calculation parameters. If there are, it writes the calculation parameters into the DRAM memory chip, and the DRAM memory chip sends the calculation parameters to the logic chip.

[0020] Optionally, the calculation step further includes: The SoC system writes the activation feature data into the DRAM memory chip through the DDR interface provided in the logic chip. If the activation feature data is larger than the on-chip storage of the logic chip, the logic chip writes the activation feature data back to the DRAM memory chip.

[0021] Optionally, the control method for the edge device further includes: The SoC system can control one or more of the AI computing subsystems to alternately execute multiple AI network models, and selectively read the calculation results of any one of the AI network models.

[0022] In the AI computing subsystem, due to the use of the near-memory computing architecture, and through the advanced packaging process of three-dimensional stacking, the chips are connected through the Hybrid Bonding process, which greatly improves the bandwidth between the DRAM memory chip and the logic chip. Therefore, for models with limited bandwidth, their performance will also be greatly improved; for example, the Token generation rate of the AI large model in mobile phone applications is several times that of the conventional SoC system, meeting the broad usage requirements of users in daily life.

[0023] In the AI computing subsystem, due to the use of the in-memory computing architecture, the power consumption of data transfer is reduced. Under the same computing volume, the operating power consumption on the edge device can be reduced, and the battery life of the edge device can be improved. In the AI computing subsystem, since a FLASH storage chip is integrated, it can store the parameters of multiple models. Then, multiple models can run on the AI computing subsystem simultaneously. Thanks to the tight integration of the FLASH storage chip, DRAM storage chip, etc. with the logic chip, the AI computing subsystem can quickly switch between different models. In addition, due to the improvement in the transmission efficiency of model data among the FLASH storage chip, logic chip, and DRAM storage chip in this architecture, the overall system power consumption and heat generation are reduced. Therefore, the design difficulty of the heat dissipation of the edge device applying this system can be reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 It is a schematic structural diagram of the edge device and the AI computing subsystem provided by the embodiment of the present invention.

[0025] Figure 2 It is a schematic structural diagram of the edge device and the AI computing subsystem provided by the embodiment of the present invention.

[0026] Figure 3 It is a schematic structural diagram of the edge device and the AI computing subsystem provided by the embodiment of the present invention.

[0027] Figure 4 It is a schematic structural diagram of the edge device and the AI computing subsystem provided by the embodiment of the present invention.

[0028] Figure 5 It is a schematic structural diagram of the edge device provided by the embodiment of the present invention.

[0029] Figure 6 It is a schematic structural diagram of the edge device provided by the embodiment of the present invention.

[0030] Figure 7 It is a schematic structural diagram of the edge device provided by the embodiment of the present invention.

[0031] Figure 8 It is a schematic structural diagram of the edge device provided by the embodiment of the present invention.

[0032] Figure 9 It is a schematic diagram of the steps of the control method of the edge device provided by the embodiment of the present invention.

[0033] Figure 10 It is a schematic diagram of the process of initializing the edge device provided by the embodiment of the present invention.

[0034] Figure 11It is a schematic flowchart of a terminal device executing a computing task provided by an embodiment of the present invention.

[0035] Figure 12 It is a schematic flowchart of a terminal device executing a computing task provided by an embodiment of the present invention.

[0036] Figure 13 It is a schematic flowchart of a terminal device executing start instructions of multiple network models provided by an embodiment of the present invention. Detailed implementation manners

[0037] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0038] As Figure 1 shown, this embodiment provides an AI computing subsystem 10, including at least one DRAM storage chip 100, at least one FLASH storage chip 200, and a logic chip 300. Among them, the DRAM storage chip 100 and the FLASH storage chip 200 are respectively connected to the logic chip 300.

[0039] At least one FLASH storage chip 200 is used to pre-store the operation parameters of at least one AI network model; at least one DRAM storage chip 100 is used to store the operation parameters of the AI network model; and the logic chip 300 is connected to the FLASH storage chip 200 and the DRAM storage chip 100, and in response to an external instruction, reads the data of the FLASH storage chip 200 for writing into the DRAM storage chip, and / or reads the operation parameters stored in the DRAM storage chip 100 and executes a computing task.

[0040] With such a setting, the AI network can run offline in the terminal device without relying on real-time Internet or connections such as WIFI / Bluetooth to the outside before it can run. When the AI network is deployed under this system computing architecture, the SoC system only needs to send the parameters or start commands of the AI network, and the entire network runs in the AI computing subsystem; or the most complex part of the AI network can be deployed in the AI computing subsystem, and the SoC system only needs to calculate some simple operators, enabling the SoC system to have more space to complete other responsible tasks.

[0041] In the AI computing subsystem, due to the use of the near-memory computing architecture and the advanced packaging process of three-dimensional stacking, the chips are connected through the Hybrid Bonding process, which greatly improves the bandwidth between the DRAM memory chip and the logic chip. Therefore, for bandwidth-limited models, their performance will also be greatly improved. For example, the token generation rate of the AI large model in mobile applications is several times that of the conventional SoC system, meeting the broad daily usage needs of users.

[0042] In the AI computing subsystem, due to the use of the near-memory computing architecture, the power consumption of data transfer is reduced. Under the same computing volume, the operating power consumption on the edge device can be reduced, and the battery life of the edge device can be improved. In the AI computing subsystem, because the FLASH memory chip is integrated, it can store the parameters of multiple models. Then, multiple models can run on the AI computing subsystem simultaneously. Thanks to the close integration of the FLASH memory chip, DRAM memory chip, etc. with the logic chip, the AI computing subsystem can quickly switch between different models. In addition, due to the improvement of the data transfer efficiency between the model data in the FLASH memory chip, logic chip, and DRAM memory chip in this architecture, the overall system power consumption and heat generation are reduced. Therefore, the design difficulty of the heat dissipation of the edge device applying this system can be reduced.

[0043] Among them, the DRAM memory chip 100 and the logic chip 300 are packaged by a three-dimensional stacking process; more specifically, the memory chip 100 and the logic chip 300 are connected through the Hybrid bonding process. Since the Hybrid bonding process can set more interfaces between the two chips, and thus can make the data transfer bandwidth between the two chips larger. Therefore, a smaller data buffer can be designed inside the logic chip 300 to reduce the chip size under the same computing power or improve the computing power under the same chip size.

[0044] Optionally, as Figures 1 - 3 shown, the connection method between the FLASH memory chip 200 and the logic chip 300 can be: packaged by a three-dimensional stacking process and connected through the Hybrid Bonding or Wire Bonding process; or, packaged by a two-dimensional process in the way of PCB Wires.

[0045] Optionally, as Figure 4 shown, when the number of DRAM memory chips 100 is greater than one, the DRAM memory chips 100 are packaged by a three-dimensional stacking process and connected through the Hybrid Bonding process.

[0046] As Figure 1As shown, this embodiment provides an edge device, including at least one AI computing subsystem 10 and an SoC system 20 connected to each AI computing subsystem 10.

[0047] Optionally, the SoC system 20 and the AI computing subsystem 10 are connected through a DDR interface.

[0048] Optionally, the protocol of the DDR interface can be any one of the sub-protocols of LPDDR, Normal DDR, and GDDR.

[0049] LPDDR (Low Power Double Data Rate) has low power consumption and is especially suitable for power-sensitive mobile devices (such as smartphones, tablets, and embedded devices). It adopts a low operating voltage (such as 1.1V, 1.2V, etc., lower than 1.5V of DDR3), reducing energy consumption. It has a "Deep Sleep Mode" and a "Partial Array Self-Refresh" to further reduce static power consumption. It has a compact design: the chip package size is small, supports multi-chip packaging (PoP, Package on Package), and reduces the occupied space. It is suitable for mobile scenarios: it can maintain high efficiency at low frequencies and is suitable for scenarios with medium requirements for latency and bandwidth but with priority on battery life.

[0050] Normal DDR (Standard DDR) has high compatibility: Standard DDR protocols (such as DDR3, DDR4, DDR5) are widely used in servers, PCs, and embedded systems. DDR supports the memory interface of general platforms and has high ecosystem support and compatibility. With the progress of generations, the frequency and data transfer rate of DDR have increased significantly. For example, DDR5 supports higher bandwidths (6400MT / s and above). Standard DDR production is mature, the market demand is large, and the cost is relatively low. It supports multi-channel and multi-DIMM configurations and is suitable for scenarios that require large-capacity storage and high memory bandwidth, such as cloud computing, AI inference servers, etc.

[0051] GDDR (Graphics Double Data Rate) has extremely high bandwidth: It is optimized for graphics processing and provides a higher data transfer rate than Normal DDR and LPDDR. The latest GDDR6X can reach 20+Gbps, meeting the needs of GPUs and high-performance computing. Its design is optimized for the parallel computing architecture of GPUs to ensure lower data transfer latency between the video memory and the processor.

[0052] Suitable for application scenarios that need to process a large amount of parallel data, such as high-resolution graphics processing, deep learning training, game development, and real-time rendering. It supports larger access lines and faster page switching rates, making it an ideal choice for high-throughput scenarios (such as graphics rendering and AI training).

[0053] Designers can select appropriate protocols based on the characteristics and requirements of each protocol.

[0054] Optionally, the DDR interface is set on the DRAM storage chip 100 or the logic chip 300.

[0055] Optionally, as Figure 5 shown, when the number of AI computing subsystems 10 is multiple, the AI computing subsystems 10 are configured such that each AI computing subsystem 10 is connected to the SoC system 20 through an independent DDR interface.

[0056] Optionally, as Figure 6 shown, when the number of AI computing subsystems 10 is multiple, the AI computing subsystems 10 are configured such that at least some or all of the AI computing subsystems 10 share a DDR interface to be connected to the SoC system 20.

[0057] In this embodiment, multiple AI computing subsystems 10 are connected to the SoC system through the same DDR interface, enabling multiple AI computing subsystems 10 to execute the same computing tasks. When multiple AI computing subsystems are connected to the SoC system through different DDR interfaces, multiple AI computing subsystems 10 can execute the same or different computing tasks, and the SoC system and the AI computing subsystems can have a distributed free topology, thus being able to flexibly adapt to different application scenarios.

[0058] Optionally, as Figure 7 shown, when the DDR interface is set on the DRAM storage chip 100, each DRAM storage chip 100 is provided with a DDR interface and each DDR interface is connected to the SoC system 20.

[0059] Optionally, as Figure 8 shown, the DDR interface can also be set on the logic chip 300, and the SoC system 20 interacts with the AI computing subsystem 10 through the logic chip 300.

[0060] Optionally, as Figure 9 shown, this embodiment provides a control method for an edge device, which is applied to control the edge device described above, including the steps of:

[0061] Initialization step: The SoC system 20 prepares initialization data and sends it to the logic chip 300 of the AI computing subsystem 10, and the logic chip 300 performs initialization configuration based on the initialization data;

[0062] Calculation steps: Based on the calculation instructions, the logic chip 300 reads the data required for calculation stored in the DRAM memory chip 100, and writes the intermediate result data and the final result data of the calculation back to the DRAM memory 100; and

[0063] Calculation result reading step: The SoC system 20 reads the final result data.

[0064] Optionally, the initialization step further includes: determining whether it is necessary to update the operating parameters; if so, the SoC system 20 sends a command to the AI calculation subsystem 10 to update the operating parameters of the DRAM memory chip 100, and in response to this command, the logic chip 300 can read the required operating parameters from the FLASH memory chip 200 and write them into the DRAM memory chip 100.

[0065] Wherein, the operating parameter can be a weight parameter.

[0066] Optionally, the calculation step further includes: The SoC system 200 writes activation feature data to the DRAM memory chip 100 through the DDR interface provided on the DRAM memory chip 100, and determines whether there are calculation parameters. If so, it writes the calculation parameters to the DRAM memory chip 100, and the DRAM memory chip 100 sends the calculation parameters to the logic chip 300.

[0067] Optionally, the calculation step further includes: The SoC system 20 writes activation feature data to the DRAM memory chip 100 through the DDR interface provided on the logic chip 300. If the activation feature data is greater than the on-chip storage of the logic chip, the logic chip 300 writes the activation feature data back to the DRAM memory chip 100.

[0068] The following combines Figures 10 - 12 to elaborate in detail on the operation process of the edge device:

[0069] Figure 10The flowchart of the initialization of the edge device is shown. First, the SoC system 20 prepares the initialization data and determines whether weight parameters need to be updated. If so, the SoC system 20 sends a command to update the weight parameters of the DRAM storage chip 100 to the AI computing subsystem 10 through the DDR interface. Based on this command, the logic chip 300 reads the corresponding weight parameter command stored in the FLASH storage chip 200 and writes it into the DRAM storage chip 100 until all the weight parameters are written. This result may require multiple operations to achieve. The SoC system 20 sends the initialization configuration command and initialization data of the logic chip 300 to the AI computing subsystem 10 through the DDR interface. After receiving the initialization data, the logic chip 300 performs the initialization configuration.

[0070] Figure 11 The flowchart of the edge device executing a computing task is shown. Taking the DDR interface being set in the DRAM storage chip 100 as an example, first, the SoC system 20 prepares the input activation feature data and determines whether there are computing parameters. If so, the SoC system 20 writes the computing parameters of the logic chip 300 into the DRAM storage chip 100 through the DDR interface. The DRAM storage chip 100 sends the computing parameters to the logic chip 300. The logic chip 300 receives the computing parameters. The SoC system 20 writes a start computing instruction into the DRAM storage chip 100 through the DDR interface. The DRAM storage chip 100 sends the start computing instruction to the logic chip 300. The logic chip 300 obtains the data required for computing from the DRAM storage chip 100 based on the computing instruction. The logic chip 300 writes some intermediate result data back to the DRAM storage chip 100. The logic chip 300 writes the final result data back to the DRAM storage chip 100. The SoC system 20 sends a final result data reading command through the DDR interface and reads the final result data based on this command, and the computing is completed.

[0071] Figure 12 The flowchart of the edge device executing a computing task is shown. Taking the DDR interface being set in the logic circuit 300 as an example, the main difference from the Figure 11 flowchart shown is that the SoC system 20 writes the activation feature data into the logic chip 300 through the DDR interface and determines whether the activation feature data is greater than the on-chip storage of the logic chip 300. If so, the activation feature data is written back to the DRAM storage chip 100.

[0072] Figure 13A schematic diagram showing the SoC system 20 interleaving and operating multiple AI computing subsystems 10 is shown. Specifically, when an edge device is configured with multiple AI computing subsystems 10, taking the example that the multiple AI computing subsystems 10 include a first AI computing subsystem and a second AI computing subsystem, for the first AI computing subsystem, wherein the first AI computing subsystem runs the AI network model A, and the second AI computing subsystem runs the AI network model B, the control method of the edge device further includes: the SoC system respectively sends an AI network model A start command and an AI network model B start command to the first AI computing subsystem and the second AI computing subsystem, and selectively reads the calculation result of the AI network model A or the calculation result of the AI network model B.

[0073] In addition, multiple network models can also be run within the same AI computing subsystem 10, and the interleaving operation method shown in Figure 13 is adopted.

[0074] In the AI computing subsystem, the SoC system can control the AI computing subsystem to perform an interleaving operation for different model calculations. For example, when there are two models A and B that need to be calculated, the start command of the network model A can be sent first, and the start command of the network model B can be sent without waiting for the network model A to complete the calculation; when reading the calculation result, the result of the network model B can be read first and then the result of the network model A can be read, or the result of the network model that is calculated first can be read. With such a design, the computing power of the AI computing subsystem and the flexibility of computing task execution can be further improved.

[0075] So far, the technical solution of the present invention has been described with reference to the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to the above specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present invention.

Claims

1. An AI computing subsystem, characterized in that: include: At least one FLASH storage chip, used to pre-store operating parameters of at least one AI network model; At least one DRAM memory chip, used to store the operating parameters of the AI ​​network model; as well as A logic chip is connected to the FLASH memory chip and the DRAM memory chip, and responds to external instructions to read data from the FLASH memory chip for writing into the DRAM memory chip, and / or reads data stored in the DRAM memory chip and performs computing tasks.

2. The AI ​​computing subsystem according to claim 1, characterized in that: The DRAM memory chip and the logic chip are packaged in a three-dimensional stacking process and connected through a Hybrid Bonding process.

3. The AI ​​computing subsystem according to claim 2, characterized in that: The FLASH memory chip and the logic chip are connected by Hybrid Bonding, Wire Bonding or PCB Wires technology.

4. The AI ​​computing subsystem according to any one of claims 1 to 3, characterized in that: When the number of the DRAM memory chips is greater than one, the DRAM memory chips are packaged using a three-dimensional stacking process and connected using a HybridBonding process.

5. A terminal side device, characterized in that: The method comprises at least one AI computing subsystem according to any one of claims 1 to 4 and a SoC system connected to each of the AI ​​computing subsystems.

6. The terminal side device according to claim 5, characterized in that: The SoC system and the AI ​​computing subsystem are connected via a DDR interface.

7. The terminal side device according to claim 6, characterized in that: The protocol of the DDR interface may be any sub-protocol of LPDDR, NormalDDR, or GDDR.

8. The terminal side device according to any one of claims 5 to 7, characterized in that: The DDR interface is arranged on the DRAM memory chip or the logic chip.

9. The terminal side device according to any one of claims 5 to 7, characterized in that: When there are multiple AI computing subsystems, the AI ​​computing subsystems are configured as follows: Each of the AI ​​computing subsystems is connected to the SoC system via an independent DDR interface.

10. The terminal side device according to any one of claims 5 to 7, characterized in that: When there are multiple AI computing subsystems, the AI ​​computing subsystems are configured as follows: At least part or all of the AI ​​computing subsystems share a DDR interface to connect to the SoC system.

11. The terminal side device according to claim 6, characterized in that: When the DDR interface is provided in the DRAM memory chip, each of the DRAM memory chips is provided with a DDR interface and each of the DDR interfaces is connected to the SoC system.

12. A method for controlling a terminal side device, applied to control the terminal side device according to any one of claims 5 to 11, characterized in that: Includes steps: Initialization step: the SoC system prepares initialization data and sends it to the logic chip of the AI ​​computing subsystem, and the logic chip performs initialization configuration based on the initialization data; Calculation step: the logic chip reads the data required for the calculation stored in the DRAM memory chip based on the calculation instruction, and writes the intermediate result data and the final result data of the calculation back to the DRAM memory; as well as Calculation result reading step: the SoC system reads the final result data.

13. The control method of the terminal side device according to claim 12, characterized in that: The initialization step also includes: Determine whether the operating parameters need to be updated; If so, the SoC system sends a command to the AI ​​computing subsystem to update the operating parameters of the DRAM memory chip. In response to the command, the logic chip can read the required operating parameters from the FLASH memory chip and write them into the DRAM memory chip.

14. The method for controlling a terminal device according to claim 12 or 13, characterized in that: The calculation step also includes: The SoC system writes activation feature data to the DRAM memory chip through a DDR interface set on the DRAM memory chip, and determines whether there are calculation parameters. If yes, the calculation parameters are written to the DRAM memory chip, and the DRAM memory chip sends the calculation parameters to the logic chip.

15. The method for controlling a terminal device according to claim 12 or 13, characterized in that: The calculation step also includes: The SoC system writes activation feature data to the DRAM memory chip via a DDR interface provided on the logic chip. If the activation feature data is larger than the on-chip storage of the logic chip, the logic chip writes the activation feature data back to the DRAM memory chip.

16. The method for controlling a terminal device according to claim 12 or 13, characterized in that: include: The SoC system is capable of controlling one or more of the AI ​​computing subsystems to alternately execute multiple AI network models, and selectively read the calculation results of any of the AI ​​network models.

Citation Information

Cited By

  • Storage and calculation integrated neural network processor with three-dimensional integration and substrate interconnection

    CN121998007A

  • Three-dimensional integrated and substrate interconnected in-memory neural network processor

    CN121998007B

  • Storage and calculation integrated neural network processor of three-dimensional integrated heterogeneous storage medium

    CN121998008A

  • Monolithic 3d integrated heterogeneous memory media in-memory processing neural network processor

    CN121998008B

  • AI chip, AI system and data scheduling method

    CN122432108A