Memory chip, data processing system, method and equipment and storage medium
By using 3D stacked electrical connection and internal data channels in the computing unit and memory, the problem of memory wall effect is solved, data processing efficiency is improved, and user needs are met.
Patent Information
- Application Number
- CN202510279323.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-27
AI Technical Summary
In the prior art, there is a memory wall effect when the computing unit interacts with the memory, resulting in a reduced data processing efficiency and unable to meet user needs.
The memory module and the base module based on 3D stacked electrical connection are adopted. Through the electrical connection between the internal control unit and the internal computing unit, the data interaction bandwidth between the memory module and the base module is optimized to form an internal data channel to improve data access efficiency.
By increasing the data interaction bandwidth between the memory module and the base module, the memory wall effect is solved, the data processing efficiency is improved, and the user's needs are met.
Smart Images

Figure CN120216408A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a memory chip, a data processing system, a method, a device, and a storage medium. Background Art
[0002] Currently, the application of artificial intelligence at the edge is becoming increasingly common. For example, the application of language models in smartphones, automobiles, and edge computing.
[0003] In the related art, the computing unit for model calculation and the memory are combined through packaging. Among them, the memory is implemented based on Low Power Double Data Rate Memory (LPDDR). The computing unit is implemented based on a logic process. Thus, the computing unit and the memory perform data interaction through the LPDDR interface.
[0004] However, there will be a memory wall effect when the above-mentioned computing unit and memory perform data interaction, which greatly reduces the data processing efficiency and thus cannot meet the user's needs. Summary of the Invention
[0005] Embodiments of this application provide a memory chip, a data processing system, a method, a device, and a storage medium, which can solve the memory wall effect, thereby improving the data processing efficiency and meeting the user's needs. The technical solution is as follows:
[0006] According to the first aspect of the embodiments of this application, a memory chip is provided. The chip includes:
[0007] A memory module and a base module that are electrically connected based on 3D stacking; the memory module is used to store input data and model parameters;
[0008] The base module includes an internally controlled unit and an internal computing unit that are electrically connected; both the internally controlled unit and the internal computing unit are electrically connected to the base module;
[0009] The internally controlled unit is used to control the internal computing unit to read the model parameters and the input data from the memory module;
[0010] The internal computing unit is used to process the input data based on the model parameters to obtain a processing result.
[0011] In a possible implementation, the memory module is provided with a first memory interface; the base module is provided with a second memory interface;
[0012] The first memory interface and the second memory interface are electrically connected based on the 3D stacking, forming an internal data channel between the first memory interface and the second memory interface.
[0013] In a possible implementation, the base module is provided with an internal LPDDR interface; the base module exchanges data with an external chip based on the internal LPDDR interface.
[0014] In a possible implementation, the base module further includes a cache module;
[0015] The memory module, the cache module, and the internal computing unit are electrically connected in sequence.
[0016] The cache module is used to store scheduling instructions;
[0017] The memory module is used to schedule the internal computing unit in response to the scheduling instructions.
[0018] In a possible implementation, the model parameters include neural network parameters or machine learning model parameters.
[0019] In a possible implementation, the internal computing unit is an NPU.
[0020] According to the second aspect of the embodiments of the present application, a data processing system is provided, including:
[0021] An electrically connected memory chip and an external chip;
[0022] The external chip includes an external computing unit and an external control unit;
[0023] The external control unit is used to write input data and model parameters into the memory module;
[0024] When a target model needs to be run, in the prefill stage:
[0025] The external computing unit is used to read the model parameters and the input data from the memory module; process the input data based on the model parameters to obtain intermediate data; write the intermediate data into the memory module;
[0026] In the decode stage:
[0027] The external control unit is further used to send a scheduling instruction to the internal control unit;
[0028] The internal control unit is used to schedule the internal computing unit in response to the scheduling instruction;
[0029] An internal computing unit for reading the intermediate data and the model parameters from the memory module, processing the intermediate data based on the model parameters to obtain a processing result, and writing the processing result into the memory module based on the internal data channel.
[0030] In a possible implementation, the internal computing unit is further configured to read the intermediate data and the model parameters from the memory module based on the internal data channel.
[0031] In a possible implementation, after obtaining the processing result, the internal computing unit generates a status value; and sends the status value to the internal control unit;
[0032] The internal control unit is further configured to write the status information into the cache module; the status information is generated based on the status value;
[0033] The external control unit is further configured to read the status information of the cache module; and determine the processing status of the input data based on the status information.
[0034] In a possible implementation, the external computing unit is an NPU.
[0035] In a possible implementation, the system is further configured to:
[0036] When the target model does not need to be run, the external control unit is further configured to write the data to be processed into the memory module based on the external data channel and the internal data channel.
[0037] In a possible implementation, the external chip further includes an external LPDDR interface, the external LPDDR is electrically connected to the internal LPDDR, and an external data channel is formed between the external LPDDR and the internal LPDDR interface.
[0038] In a possible implementation, the external control unit is further configured to write the input data and the model parameters into the memory module based on the external data channel and the internal data channel.
[0039] In a possible implementation, the external control unit is further configured to send a scheduling instruction to the internal control unit based on the external data channel.
[0040] In a possible implementation, the external control unit is further configured to send the scheduling instruction to the cache module based on the external data channel;
[0041] The internal control unit is further configured to read the scheduling instruction from the cache module.
[0042] According to a third aspect of the embodiments of the present application, a data processing method is provided, including:
[0043] When it is necessary to run a target model, write input data and model parameters into a memory module;
[0044] In the prefill stage:
[0045] Read the model parameters and the input data from the memory module; process the input data based on the model parameters to obtain intermediate data; write the intermediate data into the memory module;
[0046] In the decode stage:
[0047] Send a scheduling instruction to an internal control unit;
[0048] In response to the scheduling instruction, schedule an internal computing unit;
[0049] Read the intermediate data and the model parameters from the memory module, process the intermediate data based on the model parameters to obtain a processing result, and write the processing result into the memory module based on the internal data channel.
[0050] According to a fourth aspect of the embodiments of the present application, a computer device is provided. The computer device includes a processor and a memory. The memory is used to store at least one segment of program, and the at least one segment of program is loaded and executed by the processor to perform the data processing method.
[0051] According to a fifth aspect of the embodiments of the present application, a computer-readable storage medium is provided. At least one segment of program is stored in the computer-readable storage medium, and the at least one segment of program is loaded and executed by a processor to implement the data processing method.
[0052] In the embodiments of the present application, a memory chip is provided, including a memory module and a base module that are electrically connected based on 3D stacking; the base module includes an internally controlled unit and an internal computing unit that are electrically connected; both the internally controlled unit and the internal computing unit are electrically connected to the base module; the internally controlled unit is used to control the internal computing unit to read model parameters and input data from the memory module; the internal computing unit is used to process the input data based on the model parameters to obtain a processing result. The above technical solution improves the data interaction bandwidth between the memory module and the base module, so that the internal computing unit can quickly access the data in the memory module, can solve the memory wall effect, and further improves the data processing efficiency to meet the needs of users. Description of the Drawings
[0053] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0054] Figure 1 It is a schematic diagram of an implementation environment provided according to an embodiment of the present application;
[0055] Figure 2 It is a schematic diagram of the structure of a memory chip provided according to an embodiment of the present application;
[0056] Figure 3 It is a schematic diagram of an example structure of a memory chip provided according to an embodiment of the present application;
[0057] Figure 4 It is a schematic diagram of the structure of a data processing system provided according to an embodiment of the present application;
[0058] Figure 5 It is a schematic diagram of an example structure of a data processing system provided according to an embodiment of the present application;
[0059] Figure 6 It is a schematic diagram of the flowchart of a data processing method provided according to an embodiment of the present application;
[0060] Figure 7 It is a schematic diagram of the structure of a terminal provided according to an embodiment of the present application;
[0061] Figure 8 It is a schematic diagram of the structure of a server provided according to an embodiment of the present application. Detailed implementation manners
[0062] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail in conjunction with the accompanying drawings.
[0063] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application.
[0064] In this application, terms such as "first" and "second" are used to distinguish identical or similar items with basically the same functions. It should be understood that there is no logical or chronological dependency between "first", "second", and "nth", nor are the quantity and execution order limited. It should also be understood that although the following description uses terms such as first and second to describe various elements, these elements should not be limited by the terms.
[0065] These terms are only used to distinguish one element from another. For example, without departing from the scope of various examples, the first action can be referred to as the second action, and similarly, the second action can also be referred to as the first action. Both the first action and the second action can be actions, and in some cases, they can be separate and different actions.
[0066] Among them, at least one means one or more than one. For example, at least one action can be one action, two actions, three actions, etc., any integer greater than or equal to one. And multiple means two or more than two. For example, multiple actions can be two actions, three actions, etc., any integer greater than or equal to two.
[0067] Figure 1 It is a schematic diagram of an implementation environment provided according to an embodiment of this application. This implementation environment may include a terminal 101 and a server 102.
[0068] In the terminal 101, a data processing system is provided; among them, the data processing system includes a memory chip and an external chip.
[0069] The terminal 101 can be a smart phone with a data processing system, a wearable device, a personal computer, a laptop computer, a tablet computer, a smart TV, a vehicle-mounted terminal, etc.
[0070] The server 102 can be a single server, a server cluster composed of multiple servers, or, alternatively, a cloud processing center.
[0071] The terminal 101 is connected to the server 102 through a wired or wireless network.
[0072] In some embodiments, a wireless network or a wired network uses standard communication technologies and / or protocols. The network is typically the Internet, but can also be any network, including but not limited to any combination of a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network, or a virtual private network. In some embodiments, technologies and / or formats including HyperText Mark-up Language (HTML), Extensible Markup Language (XML), etc. are used to represent data exchanged through the network. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec), etc. can be used to encrypt all or some of the links. In other embodiments, custom and / or proprietary data communication technologies can also be used to replace or supplement the above data communication technologies.
[0073] In the related art, the model inference process includes two stages, namely the prefill stage and the decode stage. The prefill stage can be understood as loading data. The decode stage can be understood as decoding data. Among them, in the prefill stage, the input data is transferred from the memory to the computing unit. The trained model parameters (weights, biases) are loaded from the memory to the computing unit. In the decode stage, the computing unit processes the input data based on the model parameters to obtain a processing result, and writes the processing result into the memory.
[0074] Generally, in the prefill stage, there are more multiply-accumulate calculations, which require a larger computing bandwidth and a smaller memory access bandwidth, and the value of the computing bandwidth and / or memory access bandwidth is relatively large. For example, 11 to 10000. In the decode stage, the computing bandwidth requirement is smaller, the memory access bandwidth requirement is larger, and the value of the computing bandwidth / memory access bandwidth is relatively small. For example, less than 10. Among them, the computing bandwidth represents the amount of calculations that the computing unit can perform per unit time. The memory access bandwidth represents the amount of data read from or written to the memory per unit time.
[0075] That is to say, in the model inference process of the related art, the computing bandwidth and the memory access bandwidth are unbalanced, which will lead to the memory wall effect, thus greatly reducing the data processing efficiency and failing to meet the user's needs. If both the prefill stage and the decode stage are optimized simultaneously, the overhead will be very large.
[0076] Figure 2 FIG. 4 is a schematic structural diagram of a memory chip 200 according to an embodiment of the present application. The chip includes:
[0077] A memory module 201 and a base module 202 based on 3D stacked electrical connection; the memory module 201 is used to store input data and model parameters.
[0078] Among them, the memory module 201 can also be represented as a DRAM (Dynamic Random Access Memory) DIE. The base module 202 can also be represented as a base DIE.
[0079] In some embodiments, a DRAM controller is provided in the memory module 201. The DRAM controller is used to control the writing and reading of data in the memory module 201.
[0080] In some embodiments, the memory module 201 includes a plurality of memory arrays and a sensing circuit; the plurality of memory arrays and the sensing circuit are electrically connected by 3D stacking. It should be understood that 3D stacking stacks multiple layers of semiconductor chips or storage media vertically. In the case of limited chip area, the vertical capacity is increased by stacking to achieve exponential growth in capacity. Vertical interconnection reduces latency, increases bandwidth, and reduces unit cost: In addition, in mass production, the unit area cost of the stacking process is lower than that of planar expansion. That is, 3D stacking can achieve high memory access bandwidth. Optionally, 3D stacking is implemented through TSV (Through-Silicon Via technology) and bonding technology.
[0081] Figure 3 FIG. 5 is an exemplary structural diagram of a memory chip 200 according to an embodiment of the present application.
[0082] The following Figure 3 is an exemplary illustration of the memory chip 200.
[0083] In some embodiments, the base module 202 includes an internally connected control unit 2021 and an internal computing unit 2022; both the internally connected control unit 2021 and the internal computing unit 2022 are electrically connected to the base module 202. The internally connected control unit 2021 is used to control the internal computing unit 2022 to read model parameters and input data from the memory module 201. The internal computing unit 2022 is used to process the input data based on the model parameters to obtain a processing result.
[0084] In one example, the base module 202 is provided with an internal LPDDR interface 2024; the base module 202 exchanges data with an external chip based on the internal LPDDR interface 2024. For example, the base module 202 is provided with 4 x16 internal LPDDR interfaces 2024. Or the base module 202 is provided with 2 x32 internal LPDDR interfaces 2024. x16 means 16 channels.
[0085] In one example, the base module 202 is provided with a second memory interface 2025; the memory module 201 is provided with a first memory interface 2011. The first memory interface 2011 and the second memory interface 2025 are electrically connected based on 3D stacking, so as to form an internal data channel between the first memory interface 2011 and the second memory interface 2025. For example, both the first memory interface 2011 and the second memory interface 2025 are DRAM interfaces. The DRAM interface is implemented based on a High-Speed Physical Layer (HB PHY). The bandwidth for the internal data channel to transmit data reaches 2TB / s to 8TB / s.
[0086] In one example, the input data is any one of an image, text, and audio.
[0087] In one example, for the implementation of the internal control unit 2021 and the internal computing unit 2022, for example: the internal control unit 2021 is a Central Processing Unit (CPU); the internal computing unit 2022 is a Neural Processing Unit (NPU).
[0088] In one example, the model parameters include neural network parameters or machine learning model parameters. For example, the neural network parameters include weights, biases, step sizes, and attention mechanism parameters, etc. The machine learning parameters include weights, biases, and activation functions, etc. The embodiments of the present application are applicable to any model and have a wide range of applications.
[0089] In some embodiments, the base module 202 further includes a cache module 2023; the memory module 201, the cache module 2023, and the internal computing unit 2022 are electrically connected in sequence. The cache module 2023 is used to store scheduling instructions; the memory module 201 is used to schedule the internal computing unit 2022 in response to the scheduling instructions.
[0090] In one example, the cache module 2023 is set in the extended address space of the memory chip 200. Optionally, the low-order addresses are mapped to the memory module 201, and the high-order addresses are mapped to the cache module 2023. The addresses of the cache module 2023 are continuous with the addresses of the memory module 201, and the addresses of the cache module 2023 and the memory module 201 do not overlap. For example, the memory module 201 corresponds to the address space of 0 to 2.5 GB; the cache module 2023 corresponds to the address space of 2.5 GB to 3 GB. Such a setting is that when the address of the data transmitted by the internal LPDDR interface 2024 corresponds to the cache module 2023, the internal LPDDR interface 2024 sends the data to the cache module 2023. Optionally, the data of the internal LPDDR interface 2024 is sent to the cache module 2023 after being processed by Protocol Change. For example, the cache module 2023 is implemented based on Static Random Access Memory (SRAM).
[0091] In one example, the base module 202 is provided with registers for storing scheduling instructions.
[0092] In one example, the base module 202 is provided with Direct Memory Access (DMA). The internal computing unit 2022 exchanges data with the memory module 201 through the DMA and the bus.
[0093] In one example, the internal control unit 2021 exchanges data with the memory module 201 through the bus.
[0094] In one example, the memory chip 200 is encapsulated through a substrate into a standard package chip to obtain an improved LPDDR. For example, the memory chip 200 is encapsulated through PoP (Package on Package). PoP is a packaging technology that vertically stacks chips with different functions. Optionally, the memory chip 200 and the LPDDR in the related technology are encapsulated together. The obtained package chip has both the functions of the improved LPDDR and the LPDDR in the related technology, and has a wide range of applications. Optionally, the memory chip 200 and the LPDDR in the related technology are distinguished. For example, for the convenience of representation and distinction, "1" is used to represent the memory chip 200, and "0" is used to represent the LPDDR in the related technology.
[0095] Embodiments of the present application are based on a memory module and a base module with 3D stacked electrical connections; the base module includes an internally connected control unit and an internal computing unit; both the internal control unit and the internal computing unit are electrically connected to the base module; the internal control unit is used to control the internal computing unit to read model parameters and input data from the memory module; the internal computing unit is used to process the input data based on the model parameters to obtain a processing result. The above technical solution improves the data interaction bandwidth between the memory module and the base module, so that the internal computing unit can quickly access the data in the memory module, can solve the memory wall effect, and further improves the data processing efficiency to meet the needs of users.
[0096] Figure 4 FIG. 4 is a schematic structural diagram of a data processing system according to an embodiment of the present application. The system includes: a memory chip 200 and an external chip 300 that are electrically connected.
[0097] Figure 5 FIG. 5 is a schematic example structural diagram of a data processing system according to an embodiment of the present application.
[0098] Next, in conjunction with Figure 5 an exemplary description of the data processing system will be given.
[0099] In some embodiments, the external chip 300 includes an external computing unit 302 and an external control unit 301. Among them, since the external chip 300 can be applied to terminals such as smart phones and automobiles. Then, the data processing system based on the electrically connected memory chip 200 and the external chip 300 can also be applied to terminals such as smart phones and automobiles. Among them, the electrical connection between the memory chip 200 and the external chip 300 can be achieved in various ways. For example, the electrical connection between the memory chip 200 and the external chip 300 is achieved through PoP. The external chip 300 is a system on chip (SoC).
[0100] In some embodiments, the external chip 300 further includes an external LPDDR interface 303, and the external LPDDR is electrically connected to the internal LPDDR, forming an external data channel between the external LPDDR interface 303 and the internal LPDDR interface 2024. Such a setting enables the external chip 300 to be compatible with the LPDDR protocol. Through the external data channel, data interaction between the memory chip 200 and the external chip 300 can be achieved without changing the structure of the external chip 300, which is simple to operate and cost-saving. Moreover, such a setting also makes the bandwidth of the external data channel smaller while the bandwidth of the internal data channel is larger, thereby realizing the separate optimization of the prefill stage and the decode stage. For example, the memory access bandwidth corresponding to the external data channel is 76.8 GB / s, and the bandwidth for the internal data channel to transmit data reaches 4 TB / s. It can be understood that for any external chip 300, only by setting the external LPDDR interface 303 on the external chip 300 to form an external data channel with the internal LPDDR interface 2024 on the base module 202, data interaction between the internal chip and the external chip 300 can be achieved, which is simple to operate and convenient for hardware implementation.
[0101] In one example, since the cache module 2023 is provided in the base module 202, the scheduling instructions sent by the external control unit 301 can be sent to the cache module 2023 through the external data channel, thereby realizing data interaction between the internal control unit 2021 and the external control unit 301, that is, solving the adaptation problem between the memory chip 200 and the external chip 300.
[0102] In one example, a DMA is also provided in the external chip 300, and the DMA is electrically connected to the external computing unit 302, so that the external computing unit 302 interacts with the memory module 201 through the DMA.
[0103] In one example, the DMA is electrically connected to the external LPDDR interface 303 through a bus, so that the external computing unit 302 obtains the input data and model parameters transmitted through the external data channel through the DMA, then sends the obtained intermediate data to the external data channel through the DMA, and then writes the intermediate data into the memory module 201 through the external data transmission channel.
[0104] To solve the problem of excessive overhead in simultaneously optimizing the prefill stage and the decode stage, the embodiments of the present application separately optimize the prefill stage and the decode stage. The following will specifically illustrate how to separately optimize the prefill stage and the decode stage in combination with some embodiments as follows.
[0105] In some embodiments, an external control unit 301 is configured to write input data and model parameters into a memory module 201;
[0106] When the target model needs to be run, in the prefill stage:
[0107] An external computing unit 302 is configured to read model parameters and input data from the memory module 201; process the input data based on the model parameters to obtain intermediate data; and write the intermediate data into the memory module 201;
[0108] In the decode stage:
[0109] The external control unit 301 is further configured to send a scheduling instruction to an internal control unit 2021;
[0110] The internal control unit 2021 is configured to schedule an internal computing unit 2022 in response to the scheduling instruction;
[0111] The internal computing unit 2022 is configured to read the intermediate data and model parameters from the memory module 201, process the intermediate data based on the model parameters to obtain a processing result, and write the processing result into the memory module 201 based on an internal data channel.
[0112] In one example, the internal computing unit 2022 is further configured to read the intermediate data and model parameters from the memory module 201 based on the internal data channel.
[0113] In one example, the external control unit 301 is further configured to write the input data and model parameters into the memory module 201 based on an external data channel.
[0114] In one example, the external control unit 301 is further configured to send a scheduling instruction to the internal control unit 2021 based on the external data channel.
[0115] Combining the above analysis, in the prefill stage, the data flow sequentially passes through: the memory module 201, the internal data channel, the external data channel, the DMA, and the external computing unit 302. It should be noted that although the memory access bandwidth of the external data channel is much smaller than that of the internal data channel, the memory access bandwidth of the external data channel is used as the memory access bandwidth in the prefill stage.
[0116] In the decode stage, the scheduling instruction sequentially passes through: the external control unit 301, the external data channel, the cache module 2023, the internal control unit 2021, and the internal computing unit 2022.
[0117] The data stream sequentially passes through: the memory module 201, the internal data channel, the bus, the DMA, and the internal computing unit 2022. Since the data stream in the decode stage does not need to pass through the external data channel, the memory access bandwidth of the internal data channel is used as the memory access bandwidth in the decode stage. Therefore, the memory access bandwidth in the decode stage is much larger than that in the prefill stage.
[0118] It should be understood that the embodiments of the present application achieve separate optimization of the prefill stage and the decode stage by coordinating the data channel and the computing unit. Specifically, in the prefill stage, the external computing unit 302 is scheduled to process data, and the external data channel is scheduled to achieve data memory access. That is to say, in this process, the internal control unit 2021 and the internal computing unit 2022 do not need to participate. In the decode stage, the internal computing unit 2022 is scheduled to process data, and the internal data channel is scheduled to achieve data memory access. That is to say, in this process, the external computing unit 302 does not need to participate.
[0119] That is to say, in the embodiments of the present application, different computing units are respectively run in the prefill stage and the decode stage, and different data channels are respectively selected. The reason for such a setting is that the bandwidth of the external data channel is small, while the bandwidth of the internal data channel is large. Therefore, considering that the prefill stage has a relatively small requirement for memory access bandwidth and a relatively high requirement for computing bandwidth, the external data channel with a smaller bandwidth and the external computing unit 302 with a larger computing bandwidth are selected in the prefill stage; considering that the decode stage has a relatively large requirement for memory access bandwidth and a relatively small requirement for computing bandwidth, the internal data channel with a larger bandwidth and the internal computing unit 2022 with a smaller computing bandwidth are selected in the decode stage, so as to achieve separate optimization of the prefill stage and the decode stage, and greatly reduce the overhead compared with the related art.
[0120] In addition, for the computations involved in the prefill stage and the decode stage during the inference process of the target model, since the prefill stage requires a large computational bandwidth, the external computing unit 302 in the external chip 300 used for the computations in the prefill stage needs to have a large computational bandwidth. In actual application scenarios, the external computing unit 302 can be selected according to actual needs to meet the computational bandwidth of the prefill stage. Correspondingly, since the decode stage requires a small computational bandwidth, the internal computing unit 2022 used for the computations in the decode stage does not need to have a large computational bandwidth. In actual application scenarios, the internal computing unit 2022 can be selected according to actual needs. According to the above analysis, the internal computing unit 2022 and the external computing unit 302 required in actuality can be directly obtained from related technologies. Therefore, the technical solution of this application is simple to operate and easy to implement in hardware.
[0121] Combining the above analysis, the embodiments of this application provide the following technical solutions. Specifically:
[0122] In some embodiments, the external control unit 301 is further configured to send a scheduling instruction to the cache module 2023 based on an external data channel; the internal control unit 2021 is further configured to read the scheduling instruction from the cache module 2023.
[0123] In one example, the external control unit 301 is further configured to send a scheduling instruction to a register based on an external data channel; the internal control unit 2021 is further configured to read the scheduling instruction from the register.
[0124] In some embodiments, when it is necessary to run the target model, that is, when it is necessary to process data using the target model. For example, in an autonomous driving scenario, the external control unit 301 writes the input data and model parameters into the memory module 201 through an external data channel to complete the initialization of the data.
[0125] In the prefill stage, the external computing unit 302 reads the input data and model parameters from the memory module 201 through DMA, an external data channel, and an internal data channel, and processes the input data based on the model parameters to obtain intermediate data. Thus, data can be accessed at a speed of 76.8 GB / s in the prefill stage. Optionally, for the input data, the external computing unit 302 obtains a part of the input data each time and repeats the process of obtaining the input data multiple times until all the input data is read. Among them, the specific characteristics of the intermediate data are determined according to the model parameters, and the embodiments of this application do not limit this.
[0126] In the decode stage, the external control unit 301 writes the scheduling instructions into the cache module 2023 through the bus and the external data channel. The scheduling instructions are generated by the external control unit 301. The internal control unit 2021 reads the scheduling instructions by looking up the cache module 2023, and in response to the scheduling instructions, schedules the internal computing unit 2022; the internal computing unit 2022 reads the intermediate data and model parameters in the memory module 201 through the bus, the internal data channel, and the DMA; processes the intermediate data based on the model parameters to obtain a processing result; writes the processing result into the memory module 201 based on the DMA, the bus, and the DMA, thereby achieving data access at a speed of 2TB / s to 8TB / s, greatly improving the data access speed, and further improving the efficiency of data processing. Optionally, the internal computing unit 2022 obtains a part of the intermediate data each time, and repeats obtaining the intermediate data multiple times until all the intermediate data is read. It should be noted that the internal control unit 2021 is controlled to look up the cache module 2023 through a preset trigger mechanism. The preset trigger mechanism can be directly obtained from related technologies, and will not be elaborated in the embodiments of the present application.
[0127] In some embodiments, the system is further configured to: when the target model does not need to be run, the external control unit 301 is further configured to write the data to be processed into the memory module 201 based on the external data channel and the internal data channel.
[0128] In one example, when the target model does not need to be run, the external control unit 301 writes the data to be processed into the memory module 201 based on the external data channel and the internal data channel, and reads the data to be processed from the memory module 201, thereby achieving data access in a scenario where the target model does not need to be run. For example, in the initialization scenario of a smart phone. In this case, since there is no complex calculation involved and high bandwidth is not required, the internal control unit 2021, the internal computing unit 2022, and the external computing unit 302 are not required to participate, thus saving resources.
[0129] In one example, after the internal computing unit 2022 obtains the processing result, it generates a status value; sends the status value to the internal control unit 2021; the internal control unit 2021 is further configured to write the status information into the cache module 2023; the status information is generated based on the status value; the external control unit 301 is further configured to read the status information of the cache module 2023; and determine the processing status of the input data based on the status information. For the implementation of the external computing unit 302 and the external control unit 301, for example, the external computing unit 302 is an NPU. The external control unit 301 is a CPU.
[0130] Optionally, after the external control unit 301 determines, based on the status information, that the processing of the input data is completed, the handshake is completed, and the processing result is sent to the client for the user to view the processing result, thereby meeting the user's needs and greatly improving the user's satisfaction.
[0131] It should be noted that the external chip 300 may further include other functional modules, and the embodiments of the present application do not limit this.
[0132] In the embodiments of the present application, the system includes a memory chip and an external chip that are electrically connected; the external chip includes an external computing unit and an external control unit. The external control unit is configured to write the input data and the model parameters into the memory module; when it is necessary to run the target model, in the prefill stage: the external computing unit is configured to read the model parameters and the input data from the memory module; process the input data based on the model parameters to obtain intermediate data; write the intermediate data into the memory module; in the decode stage: the external control unit is further configured to send a scheduling instruction to the internal control unit; the internal control unit is configured to, in response to the scheduling instruction, schedule the internal computing unit; the internal computing unit is configured to read the intermediate data and the model parameters from the memory module, process the intermediate data based on the model parameters to obtain a processing result, and write the processing result into the memory module based on the internal data channel. Combining the above analysis, it can be seen that the technical solution of the present application selects an external data channel with a smaller bandwidth and an external computing unit with a larger computing bandwidth in the prefill stage, so as to access data with a small bandwidth and process data with a large computing bandwidth, while in the decode stage, it selects an internal data channel with a larger bandwidth and an internal computing unit with a smaller computing bandwidth, so as to access data with a large bandwidth and process data with a small computing bandwidth, thereby realizing the optimization of the prefill stage and the decode stage respectively. Compared with the related art, the overhead is greatly reduced.
[0133] Figure 6 It is a schematic flowchart of a data processing method provided by an embodiment of the present application. As Figure 6 shown, in the embodiments of the present application, it is described by taking an application to a terminal with a data processing system as an example. The method includes the following steps:
[0134] In step 601, when it is necessary to run the target model, the input data and the model parameters are written into the memory module.
[0135] In step 602, in the prefill stage: read the model parameters and the input data from the memory module; process the input data based on the model parameters to obtain intermediate data; write the intermediate data into the memory module;
[0136] In step 603, in the decode stage: send a scheduling instruction to the internal control unit;
[0137] In step 604, in response to a scheduling instruction, schedule an internal computing unit;
[0138] In step 605, read intermediate data and model parameters from a memory module, process the intermediate data based on the model parameters to obtain a processing result, and write the processing result into the memory module based on an internal data channel.
[0139] In some embodiments, the method further includes: reading intermediate data and model parameters from a memory module based on an internal data channel.
[0140] In some embodiments, the method further includes: after obtaining the processing result, the internal computing unit generates a status value; the status value is sent to an internal control unit;
[0141] The internal control unit is further configured to write status information into a cache module; the status information is generated based on the status value;
[0142] The external control unit is further configured to read the status information of the cache module; determine the processing status of input data based on the status information.
[0143] In some embodiments, the external computing unit is an NPU.
[0144] In some embodiments, the method further includes: when the target model does not need to be run, the external control unit is further configured to write data to be processed into the memory module based on an external data channel and an internal data channel.
[0145] In some embodiments, the external chip further includes an external LPDDR interface, the external LPDDR is electrically connected to the internal LPDDR, and an external data channel is formed between the external LPDDR and the internal LPDDR interface.
[0146] In some embodiments, the method further includes: writing input data and model parameters into the memory module based on an external data channel and an internal data channel.
[0147] In some embodiments, the method further includes: sending a scheduling instruction to the internal control unit based on an external data channel.
[0148] In some embodiments, the method further includes: sending the scheduling instruction to the cache module based on an external data channel;
[0149] The internal control unit is further configured to read the scheduling instruction from the cache module.
[0150] It should be noted that: when the data processing method provided in the above embodiments executes the corresponding steps, only the division of the above functional modules is used for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the data processing system provided in the above embodiments and the embodiments of the data processing method belong to the same concept. For the specific implementation process, please refer to the system embodiments and will not be elaborated here.
[0151] In an embodiment of the present application, the system includes a memory chip and an external chip that are electrically connected; the external chip includes an external computing unit and an external control unit. The external control unit is configured to write input data and model parameters into the memory module; when it is necessary to run the target model, in the prefill stage: the external computing unit is configured to read the model parameters and input data from the memory module; process the input data based on the model parameters to obtain intermediate data; write the intermediate data into the memory module; in the decode stage: the external control unit is further configured to send a scheduling instruction to the internal control unit; the internal control unit is configured to schedule the internal computing unit in response to the scheduling instruction; the internal computing unit is configured to read the intermediate data and model parameters from the memory module, process the intermediate data based on the model parameters to obtain a processing result, and write the processing result into the memory module based on the internal data channel. Combining the above analysis, it can be seen that the technical solution of the present application selects an external data channel with a smaller bandwidth and an external computing unit with a larger computing bandwidth in the prefill stage, so as to access data with a small bandwidth and process data with a large computing bandwidth, while in the decode stage, it selects an internal data channel with a larger bandwidth and an internal computing unit with a smaller computing bandwidth, so as to access data with a large bandwidth and process data with a small computing bandwidth, thereby realizing the optimization of the prefill stage and the decode stage respectively. Compared with the related art, the overhead is greatly reduced.
[0152] An embodiment of the present application further provides a computer device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the above method is implemented.
[0153] Taking the computer device as a terminal as an example, Figure 7 is a schematic structural diagram of a terminal provided by an embodiment of the present application. Refer to Figure 7, the terminal 700 may be: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer or a desktop computer. The terminal 700 may also be referred to by other names such as a user equipment, a portable terminal, a laptop terminal, a desktop terminal, etc.
[0154] Generally, the terminal 700 includes: a processor 701 and a memory 702.
[0155] The processor 701 may include one or more processing cores, such as a 4-core processor, a 5-core processor, etc. The processor 701 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), PLA (Programmable Logic Array). The processor 701 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 701 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 701 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process the computational operations related to machine learning.
[0156] The memory 702 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 702 may further include a high-speed random access memory and a non-volatile memory, such as one or more disk storage devices, flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 702 is used to store at least one program code, and the at least one program code is used to be executed by the processor 701 to implement the process executed by the terminal in the method provided in the method embodiments of the present application.
[0157] In some embodiments, the terminal 700 may further optionally include: a peripheral device interface 703 and at least one peripheral device. The processor 701, the memory 702, and the peripheral device interface 703 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 703 through a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of a display screen 704, a camera assembly 705, an audio circuit 706, and a power supply 707.
[0158] The peripheral device interface 703 can be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 701 and the memory 702. In some embodiments, the processor 701, the memory 702, and the peripheral device interface 703 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 701, the memory 702, and the peripheral device interface 703 can be implemented on a separate chip or circuit board, and the embodiments of the present application do not limit this.
[0159] The display screen 704 is used to display a UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 704 is a touch display screen, the display screen 704 also has the ability to collect touch signals on or above the surface of the display screen 704. The touch signals can be input to the processor 701 as control signals for processing. At this time, the display screen 704 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 704, which is disposed on the front panel of the terminal 700; in some other embodiments, there may be at least two display screens 704, which are respectively disposed on different surfaces of the terminal 700 or are in a foldable design; in some other embodiments, the display screen 704 may be a flexible display screen, which is disposed on the curved surface or the folding surface of the terminal 700. Even, the display screen 704 can be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 704 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0160] The camera component 705 is used to collect images or videos. In some embodiments, the camera component 705 includes a front camera and a rear camera. Generally, the front camera is disposed on the front panel of the terminal, and the rear camera is disposed on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth camera, a wide-angle camera, and a telephoto camera, so as to implement the function of background blurring by fusing the main camera and the depth camera, panoramic shooting by fusing the main camera and the wide-angle camera, and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera component 705 may further include a flash. The flash may be a single-color temperature flash or a dual-color temperature flash. The dual-color temperature flash refers to the combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.
[0161] The audio circuit 706 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals and input them to the processor 701 for processing. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively disposed at different parts of the terminal 700. The microphone may also be an array microphone or an omnidirectional collection type microphone. The speaker is used to convert the electrical signal from the processor 701 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 706 may further include a headphone jack.
[0162] The power supply 707 is used to supply power to each component in the terminal 700. The power supply 707 may be alternating current, direct current, a disposable battery or a rechargeable battery. When the power supply 707 includes a rechargeable battery, the rechargeable battery may support wired charging or wireless charging. The rechargeable battery may also be used to support fast charging technology.
[0163] Those skilled in the art can understand that Figure 7 the structure shown in
[0164] does not constitute a limitation on the terminal 700, and may include more or fewer components than shown in the figure, or combine certain components, or adopt different component arrangements. Figure 8It is a schematic structural diagram of a server provided by an embodiment of the present application. The server 800 may vary greatly due to different configurations or performances, and may include one or more processors (Central Processing Units, CPUs) 801 and one or more memories 802. Among them, at least one computer program is stored in the one or more memories 802, and the at least one computer program is loaded and executed by the one or more processors 801 to implement the above data processing method. Of course, the server 800 may also have components such as wired or wireless network interfaces, keyboards, and input / output interfaces for input and output. The server 800 may also include other components for implementing device functions, which will not be elaborated here.
[0165] An embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium includes a stored computer program. Among them, when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the method as described above. Optionally, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.
[0166] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware or by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk, or an optical disc, etc.
[0167] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A memory chip, characterized in that: include: Memory modules and base modules based on 3D stacking electrical connections; The memory module is used to store input data and model parameters; The base module comprises an internal control unit and an internal computing unit which are electrically connected; the internal control unit and the internal computing unit are both electrically connected to the base module; The internal control unit is used to control the internal computing unit to read the model parameters and the input data from the memory module; The internal computing unit is used to process the input data based on the model parameters to obtain a processing result.
2. The chip according to claim 1, characterized in that: The memory module is provided with a first memory interface; the base module is provided with a second memory interface; The first memory interface and the second memory interface are electrically connected based on the 3D stacking, and an internal data channel is formed between the first memory interface and the second memory interface.
3. The chip according to claim 1, characterized in that: The base module is provided with an internal LPDDR interface; the base module exchanges data with an external chip based on the internal LPDDR interface.
4. The chip according to claim 1, characterized in that: The base module also includes a cache module; The memory module, the cache module and the internal computing unit are electrically connected in sequence; The cache module is used to store scheduling instructions; The memory module is used to schedule the internal computing unit in response to the scheduling instruction.
5. The chip according to claim 1, characterized in that: The model parameters include neural network parameters or machine learning model parameters.
6. The chip according to claim 1, characterized in that: The internal computing unit is an NPU.
7. A data processing system, characterized in that: include: The memory chip and the external chip electrically connected as claimed in any one of claims 1 to 6; The external chip includes an external computing unit and an external control unit; The external control unit is used to write input data and model parameters into the memory module.
8. The system according to claim 7, characterized in that The system also includes, when it is necessary to run the target model, in the prefill phase: The external computing unit is used to read the model parameters and the input data from the memory module; Process the input data based on the model parameters to obtain intermediate data; write the intermediate data into the memory module; In the decoding stage: The external control unit is further used to send a scheduling instruction to the internal control unit; The internal control unit is used to schedule the internal computing unit in response to the scheduling instruction; The internal computing unit is used to read the intermediate data and the model parameters from the memory module, process the intermediate data based on the model parameters to obtain a processing result, and write the processing result into the memory module based on the internal data channel.
9. The system according to claim 8, characterized in that The internal computing unit is further used to read the intermediate data and the model parameters from the memory module based on the internal data channel.
10. The system according to claim 8, characterized in that After obtaining the processing result, the internal calculation unit generates a state value; and sends the state value to the internal control unit; The internal control unit is further used to write the state information into a cache module; the state information is generated based on the state value; The external control unit is further used to read the status information of the cache module; and determine the processing status of the input data based on the status information.
11. The system according to claim 8, characterized in that The external computing unit is an NPU.
12. The system according to claim 8, characterized in that The system is also used to: When the target model does not need to be run, the external control unit is further used to write the data to be processed into the memory module based on the external data channel and the internal data channel.
13. The system according to claim 8, characterized in that The external chip further includes an external LPDDR interface, the external LPDDR is electrically connected to the internal LPDDR, and an external data channel is formed between the external LPDDR and the internal LPDDR interface.
14. The system according to claim 9, characterized in that The external control unit is further used to write the input data and model parameters into the memory module based on the external data channel and the internal data channel.
15. The system according to claim 9, characterized in that The external control unit is further used to send a scheduling instruction to the internal control unit based on the external data channel.
16. The system according to claim 9, characterized in that The external control unit is further used to send the scheduling instruction to the cache module based on the external data channel; The internal control unit is further configured to read the scheduling instruction from the cache module.
17. A data processing method, characterized in that: include: When the target model needs to be run, the input data and model parameters are written into the memory module; In the prefill phase: Read the model parameters and the input data from the memory module; Process the input data based on the model parameters to obtain intermediate data; write the intermediate data into the memory module; In the decoding stage: Send dispatch instructions to the internal control unit; In response to the scheduling instruction, scheduling an internal computing unit; The intermediate data and the model parameters are read from the memory module, the intermediate data are processed based on the model parameters to obtain a processing result, and the processing result is written into the memory module based on the internal data channel.
18. A computer device, characterized in that: The computer device includes a processor and a memory, wherein the memory is used to store at least one program, and the at least one program is loaded by the processor and executes the data processing method as claimed in claim 17.
19. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one program, and the at least one program is loaded and executed by the processor to implement the data processing method as claimed in claim 17.