Embedded Hybrid Computing Semiconductor Memory
By introducing a computing storage fusion engine and heterogeneous resource management module in embedded hybrid computing semiconductor memory, dynamically adjusting the data transmission path and computing unit use, the performance bottleneck of the traditional computing and storage separation architecture is solved, and the effect of efficient computing, optimization of energy consumption and reducing latency is achieved.
Patent Information
- Application Number
- CN202510179446.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-02-19
AI Technical Summary
The traditional computing and storage separation architecture faces performance bottlenecks, including large data handling overhead, limited interface bandwidth, low computing architecture efficiency and high programming model complexity, which affect system performance and scalability.
Design an embedded hybrid computing semiconductor memory, including a computing storage fusion engine and a heterogeneous resource management module. The computing storage convergence engine dynamically changes data transmission paths through reconfigurable interconnect networks to reduce data movement. The heterogeneous resource management module dynamically adjusts the use of computing units based on reinforcement learning and optimizes resource utilization.
By reducing data movement, improving computing efficiency, optimizing energy consumption ratio, reducing system delays, and achieving flexible computing unit allocation and intelligent task scheduling, improving the efficiency of resource collaborative optimization and unified management framework.
Smart Images

Figure CN119668524B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of memory technology, and in particular to embedded hybrid computing semiconductor memory. Background Art
[0002] With the rapid growth of data scale and the continuous improvement of computing needs, the traditional architecture of separation of computing and storage faces serious performance bottlenecks. The main problems include high data handling overhead, limited performance improvement due to traditional interface bandwidth, low efficiency of general computing architecture, and high complexity of programming model. These limitations affect the overall performance and scalability of the system. Embedded semiconductor memory (such as embedded eMMC, UFS, etc.) excels in improving system integration and performance, providing a possibility to solve these problems. Summary of the invention
[0003] The embedded hybrid computing semiconductor memory provided in the present application can reduce data movement to improve computing efficiency, optimize energy consumption ratio, and reduce system latency.
[0004] The present application provides an embedded hybrid computing semiconductor memory, comprising: a computing-storage fusion engine, comprising several types of computing units; wherein a reconfigurable interconnection network is formed between the several computing units and the memory, allowing dynamic change of the data transmission path; and a heterogeneous resource management module, for dynamically adjusting the use of computing units according to real-time loads based on reinforcement learning.
[0005] Among them, several types of computing units include deep learning neural network computing units, encryption algorithm computing units, FPGA and GPU, and pulse neural network computing units.
[0006] Among them, the neural network computing unit of deep learning includes a shallow embedding sub-model and a deep decision sub-model. The shallow embedding sub-model is used to process low-dimensional features, and the deep decision sub-model is used to process high-dimensional features. Different computing resources are allocated to the shallow embedding sub-model and the deep decision sub-model. The shallow embedding sub-model is allocated for global collaborative learning, and the deep decision sub-model is allocated for intra-cluster collaborative learning.
[0007] Among them, the interconnection network includes: an optical path switching regional interconnection architecture, which is used to use optical path switching to realize the interconnection between computing units and storage; a dynamically adjusted interconnection network module, which is used to use a traffic monitor to track regional network needs, and use a greedy algorithm to generate OCS topology and dynamically adjust optical path connections.
[0008] Among them, the computing and storage fusion engine also includes an on-chip data compression and decompression module, which is used to compress and / or decompress data during data transmission, and adopts hardware acceleration when compressing and / or decompressing data.
[0009] Among them, the heterogeneous resource management module includes: a hybrid parallel module, which is used to use feature-level parallelism for neighbor aggregation and node-level parallelism for node update; a hybrid computing unit, including a sparse computing unit and a dense computing unit, the sparse computing unit is used for sampling sparse matrix multiplication, and the dense computing unit is used for node update.
[0010] The sparse computing unit is used to select the required neighbor features using a hybrid sparse indexing module and to pipeline data processing using local double buffering.
[0011] The heterogeneous resource management module is also used to reduce the frequency of the computing unit or select computing resources with power consumption lower than a preset power consumption for tasks with a delay requirement lower than a threshold.
[0012] Among them, the embedded hybrid computing semiconductor memory also includes: a development framework support module for providing a programming abstraction level higher than a preset level, providing automated code optimization tools and formal verification tools.
[0013] Among them, the automated code optimization tool is used for automated code optimization, equivalence verification using instruction-level equivalence detection tools, and verification of optimization features for embedded systems.
[0014] The beneficial effects of the present application are as follows: Different from the prior art, the embedded hybrid computing semiconductor memory provided by the present application includes: a computing storage fusion engine, including several types of computing units; wherein a reconfigurable interconnection network is formed between several computing units and memories, allowing dynamic changes in data transmission paths; a heterogeneous resource management module, which is used to dynamically adjust the use of computing units according to real-time loads based on reinforcement learning, and can reduce data movement by allowing dynamic changes in data transmission paths to improve computing efficiency, optimize energy consumption ratio, and reduce system latency. And the heterogeneous resource management module is used to perform flexible computing unit allocation, intelligent task scheduling, resource collaborative optimization, and a unified management framework. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. Among them:
[0016] Figure 1 It is a structural schematic diagram of an embodiment of an embedded hybrid computing semiconductor memory provided by the present application;
[0017] Figure 2 It is a structural diagram of an embodiment of an interconnection network provided by the present application;
[0018] Figure 3 It is a structural diagram of an embodiment of a heterogeneous resource management module provided by the present application. DETAILED DESCRIPTION
[0019] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It will be understood that the specific embodiments described herein are only used to explain the present application, rather than to limit the present application. It should also be noted that, for ease of description, only some but not all structures related to the present application are shown in the drawings. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.
[0020] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0021] With the rapid growth of data scale and the continuous improvement of computing needs, the traditional architecture of separation of computing and storage faces serious performance bottlenecks. The main problems include high data handling overhead, limited performance improvement due to traditional interface bandwidth, low efficiency of general computing architecture, and high complexity of programming model. These limitations affect the overall performance and scalability of the system. Embedded semiconductor memory (such as embedded eMMC, UFS, etc.) excels in improving system integration and performance, providing a possibility to solve these problems.
[0022] Based on this, the embedded hybrid computing semiconductor memory proposed in this application includes: a computing storage fusion engine, including several types of computing units; wherein a reconfigurable interconnection network is formed between several computing units and memories, allowing dynamic changes in data transmission paths; a heterogeneous resource management module, which is used to dynamically adjust the use of computing units according to real-time loads based on reinforcement learning, and can reduce data movement by allowing dynamic changes in data transmission paths to improve computing efficiency, optimize energy consumption ratios, and reduce system delays. And use heterogeneous resource management modules to perform flexible computing unit allocation, intelligent task scheduling, resource collaborative optimization, and a unified management framework. For details, please refer to any of the following embodiments or a combination of any of the embodiments.
[0023] See also Figure 1 , Figure 11 is a schematic diagram of the structure of an embodiment of an embedded hybrid computing semiconductor memory provided by the present application. The embedded hybrid computing semiconductor memory 100 comprises: a computing and storage fusion engine 10 and a heterogeneous resource management module 20.
[0024] The computing storage fusion engine 10 includes several types of computing units; wherein, a reconfigurable interconnection network is formed between the computing units and the memory, allowing the data transmission path to be dynamically changed. Among them, several types of computing units include deep learning neural network computing units, encryption algorithm computing units, FPGA and GPU, and pulse neural network computing units. These units can be dynamically activated or adjusted according to the characteristics of the task to further improve computing efficiency.
[0025] Traditional computing storage fusion architectures may rely mainly on general-purpose processors (such as CPUs) or programmable logic (FPGAs). In order to improve the performance of specific applications, specialized hardware computing units can be integrated, such as neural network computing units (NNAs) for deep learning or encryption algorithm computing units for cryptography. These computing units are usually designed using ASIC (Application-Specific Integrated Circuit) and optimized for specific algorithms to achieve higher energy efficiency and performance. First, analyze the computing characteristics of the target application, such as a large number of convolution operations and activation function calculations for convolutional neural networks (CNNs), or modular operations and bit operations in encryption algorithms. Then, design specialized hardware circuits to implement these operations, such as using a systolic array (a hardware architecture specifically for parallel computing, which consists of a grid-like simple processing units (also called processing elements, PEs or cells). These processing units are tightly coupled together to complete computing tasks by rhythmically passing data. Think of it as a data "pipeline" where data flows through the system like blood and is calculated in each processing unit to get the final result) to accelerate matrix multiplication, or use pipelines to increase data throughput. Finally, these hardware computing units are integrated into the computing and storage fusion engine 10, and the corresponding control logic is designed to achieve flexible resource scheduling and data transmission.
[0026] In some embodiments, the neural network computing unit of deep learning includes a shallow embedding sub-model and a deep decision sub-model, the shallow embedding sub-model is used to process low-dimensional features, and the deep decision sub-model is used to process high-dimensional features, and the shallow embedding sub-model and the deep decision sub-model are allocated different computing resources, the shallow embedding sub-model is allocated for global collaborative learning, and the deep decision sub-model is allocated for intra-cluster collaborative learning.
[0027] In deep learning models, shallow and deep parts often have different computing characteristics and data requirements. For example, shallow layers usually process low-dimensional features, while deep layers perform high-dimensional feature extraction and decision-making. The model can be split into a shallow embedding sub-model (processing low-dimensional features) and a deep decision sub-model (processing high-dimensional features), and different computing resources can be allocated. The shallow embedding sub-model is insensitive to data heterogeneity and is suitable for global collaborative learning, while the deep decision sub-model is sensitive to data distribution and is suitable for intra-cluster collaborative learning. Analyze the model structure and determine the nodes suitable for splitting. For example, in the Transformer model, the embedding layer can be regarded as the shallow part, and the Transformer block as the deep part. For the split sub-model, select the appropriate computing unit. For example, the shallow part can use a general-purpose CPU or FPGA, while the deep part can use an NNA. According to the characteristics of the sub-model, design different aggregation strategies, such as global aggregation of the shallow part and intra-cluster aggregation of the deep part.
[0028] In some embodiments, spiking neural networks (SNNs) are used as computing units. SNNs imitate the working mode of biological neurons and use spikes to encode and transmit information. Compared with traditional ANNs, SNNs have higher energy efficiency and stronger robustness. Convert the ANN model to an SNN model. Use the pulse characteristics of SNN for calculation, such as using the leaky-integrate-and-fire (LIF) model to simulate the membrane potential changes of neurons. Train SNNs using methods such as temporal backpropagation or surrogate gradient. Considering the sparse activation characteristics of SNNs, it can be combined with compression techniques such as Top-κ sparsification to further improve energy efficiency.
[0029] See also Figure 2 , Figure 2 1 is a schematic diagram of the structure of an interconnection network embodiment provided by the present application. The interconnection network includes: an optical path switching regional interconnection architecture 31 and a dynamic adjustment interconnection network module 32.
[0030] The optical path switching regional interconnection architecture 31 is used to realize the interconnection between the computing unit and the memory by using optical path switching.
[0031] Optical Path Switching Regional Interconnect Architecture 31 is a new, cost-effective, reconfigurable network architecture designed for large-scale Mixed Experts (MoE) training. Optical Path Switching (OCS) is used to achieve fast interconnection between computing units and storage. OCS can dynamically change the transmission path of optical signals in milliseconds to achieve high-bandwidth, low-latency data transmission. The entire system is divided into multiple regions, and an OCS network is deployed in each region. The system is divided into multiple regions based on the distribution of computing units and storage. The OCS network is deployed in each region, allowing fast interconnection between computing units and storage within the region. The electrical signal network (EPS) is used to connect OCS networks in different regions to achieve global data transmission. According to real-time communication needs, the connection of the OCS network is dynamically adjusted to optimize the data transmission path.
[0032] OCS (Optical Circuit Switching) is an optical circuit switching technology that is mainly used to achieve fast switching and routing of optical signals in optical fiber networks. OCS technology achieves connection by physically moving the optical signal path and is suitable for application scenarios that require high bandwidth and low latency.
[0033] OCS technology can be divided into the following categories according to different implementation technologies:
[0034] 3D MEMS (Micro-Electro-Mechanical System) technology: realizes the switching of optical paths through micro-mechanical structures.
[0035] Digital Liquid Crystal (DLC) technology: uses liquid crystal materials to control the switching of light paths.
[0036] Direct Light Beam Steering DLBS (Direct Light Beam Steering) technology: directly controls the deflection of the light beam electronically.
[0037] OCS can be used in data centers and supercomputers.
[0038] Furthermore, parallel strategies and interconnection methods can be adopted. Such as (data parallelism (DP), model parallelism (MP), expert parallelism (EP), pipeline parallelism (PP)). For example, tensor parallelism (TP) usually requires high-speed interconnection within the server, EP requires high-speed interconnection between different servers, and DP and PP require lower bandwidth. Choose different interconnection technologies for different parallel strategies. Analyze the communication requirements of different parallel strategies. For example, TP usually only requires a GPU switch inside the server, while EP requires broadcast communication to everyone between different servers. Select a high-speed interconnection inside the server, such as a GPU switch, for TP, select OCS for EP, and select an electrical signal network (EPS) for DP and PP. Dynamically adjust the topology of the interconnection network according to the dynamic changes of the parallel strategy.
[0039] The dynamic adjustment interconnection network module 32 is used to track the regional network demand using a traffic monitor, and generate an OCS topology using a greedy algorithm to dynamically adjust the optical path connection.
[0040] The broadcast communication to everyone in the MoE model is semi-predictable, and this feature can be used to dynamically adjust the interconnection network. The reconfigurable network architecture uses traffic monitors to track regional network requirements and uses a greedy algorithm to generate OCS topology, dynamically adjusting optical path connections to adapt to changes in data transmission requirements. Use traffic monitors to collect real-time communication requirements between different computing units and memories. Based on the collected communication requirements, a greedy algorithm is used to generate OCS topology. For example, OCS direct links are preferentially allocated to communication-intensive tasks. The OCS topology is dynamically adjusted to adapt to changes in communication requirements.
[0041] Among them, the computing storage fusion engine 10 also includes an on-chip data compression and decompression module, which is used to compress and / or decompress data during data transmission, and adopts hardware acceleration when compressing and / or decompressing data.
[0042] In some embodiments, for data that needs to be read and written frequently, compression can be performed first to reduce storage occupancy and transmission bandwidth requirements, and reduce latency. For example, the Top-κ sparsification method can be used to reduce the amount of data transmitted. This method only retains the top κ parameters with the largest absolute value in the gradient and sets other parameters to zero. The dynamic-κ reduction strategy refers to gradually reducing the value of κ during model training, that is, gradually reducing the amount of parameters transmitted. Calculate the gradient value of each parameter. Select the top κ parameters with the largest absolute value of the gradient and set other parameters to zero. During model training, dynamically adjust the value of κ according to a predefined strategy, such as using a linear or exponential method to reduce the κ value. The compression operation can be performed at multiple levels. For example, data is compressed before being transmitted to the Internet and compressed before being stored in the memory. It can be compressed before the client uploads the parameters and decompressed before the server aggregates the model. Analyze the characteristics of the data at different levels, such as data size, data type, etc. Select a suitable compression algorithm, such as lossless compression or lossy compression. Integrate the compression and decompression modules into the data transmission path.
[0043] The heterogeneous resource management module 20 is used to dynamically adjust the use of computing units according to real-time load based on reinforcement learning. In some embodiments, in addition to using machine learning algorithms for task scheduling, the heterogeneous resource management module 20 can introduce reinforcement learning algorithms. Through continuous trial and error and feedback, the resource scheduling strategy is optimized to adapt to more complex and dynamic application scenarios.
[0044] Dynamic computing unit allocation for reinforcement learning dynamic scheduling can be as follows: Different computing units are used according to the type of computation (sparse or dense). Reinforcement learning is applied to computing unit allocation so that it can dynamically adjust the use of computing units according to real-time load instead of static allocation.
[0045] Traditional static computing unit allocation may not be able to adapt to changing computing needs. Reinforcement learning can learn the optimal computing unit allocation strategy through trial and error and feedback, improve resource utilization and reduce latency.
[0046] It can also improve the performance and energy efficiency of memory under different loads. For example, when there are more sparse calculations, reinforcement learning can dynamically increase the use of sparse computing units; conversely, it can increase the use of dense computing units.
[0047] In some embodiments, reinforcement learning models resource scheduling problems through Markov decision processes, optimizes strategies through continuous trial and error and feedback, and thus achieves higher resource utilization efficiency. The effectiveness of this method has been verified in many dynamic resource scheduling problems.
[0048] See also Figure 3 , Figure 31 is a schematic diagram of a structure of an embodiment of a heterogeneous resource management module provided by the present application. The heterogeneous resource management module 20 includes: a hybrid parallel module 21 and a hybrid computing unit 22.
[0049] The huge feature tensor leads to memory explosion and limited communication bandwidth, and the mixture of sparse and dense matrix operations makes it difficult to effectively utilize computing resources. To address these challenges, the present application proposes two core technologies: a hybrid parallel module 21 and a hybrid computing unit 22.
[0050] The hybrid parallel module 21 is used to perform neighbor aggregation in feature-level parallelism and node update in node-level parallelism.
[0051] Although traditional partitioned parallelism can divide storage into multiple parts and assign them to different computing units for processing, it will lead to repeated storage and transmission of remote neighbor node data, thereby increasing communication volume and memory overhead. In addition, since real-world files are usually irregular, it is difficult to find a partitioning scheme that balances the load. Based on this, the hybrid parallelism proposed in this application uses feature-level parallelism for neighbor aggregation and node-level parallelism for node update.
[0052] Feature-level parallelism is mainly reflected in the fact that the feature vector of the node is divided by dimension and assigned to different computing units for neighbor aggregation calculation. Each computing unit is only responsible for calculating part of the features, avoiding repeated storage and transmission of data of remote neighbor nodes.
[0053] Node-level parallelism is mainly reflected in the fact that each computing unit obtains the characteristics of all nodes and updates the characteristics of the nodes assigned to it.
[0054] The advantages are,constant communication volume and memory usage, load balancing, and more regular,communication patterns.
[0055] The constant communication volume and memory usage mainly reflects that hybrid parallelism avoids the replication of remote neighbor data, so the communication volume and characteristic memory requirements are independent of the number of computational units.
[0056] Load balancing is mainly reflected in the even division of feature tensors and mixed parallelism to ensure load balancing of aggregation and update computing units.
[0057] A more regular communication pattern is mainly reflected in the mixed parallel use of broadcast communications to everyone, but because the communication volume is constant, the communication pattern is more regular and can improve efficiency.
[0058] The hybrid computing unit 22 includes a sparse computing unit and a dense computing unit. The sparse computing unit is used for sampling sparse matrix multiplication, and the dense computing unit is used for node update.
[0059] The sparse computing unit is used to select the required neighbor features using a hybrid sparse indexing module and to pipeline data processing using local double buffering.
[0060] Real-time computing involves both sparse matrix operations (neighborhood aggregation) and dense matrix operations (node updates). If a single type of computing unit is used to handle all calculations, hardware utilization will be low.
[0061] The hybrid computing unit 22 sets a corresponding sparse computing unit for a sparse operation (neighborhood aggregation) and sets a corresponding dense computing unit for a dense operation (node update).
[0062] The sparse computing unit is mainly reflected in the computing unit (FPGA) dedicated to sampling sparse matrix multiplication. The fusion of two consecutive sparse operations can reduce the amount of calculation and data movement. This computing unit uses a hybrid sparse index module to select the required neighbor features and uses local double buffering to pipeline data processing.
[0063] The dense computing unit is mainly reflected in the use of common dense computing units (such as GPU) for node updates.
[0064] The advantages lie in targeted acceleration, sparse matrix multiplication acceleration, and fine-grained pipelines.
[0065] Targeted acceleration is mainly reflected in selecting appropriate computing units according to the characteristics of the operation, which can improve hardware utilization and computing efficiency.
[0066] The sparse matrix multiplication acceleration is mainly reflected in the fact that the specialized sparse computing unit can efficiently process sampled sparse matrix multiplication operations, further improving the training speed.
[0067] The fine-grained pipeline is mainly reflected in the fact that the hybrid computing unit 22 adopts a fine-grained pipeline with node reordering, which can further improve the utilization and throughput of the computing unit. The nodes are reordered by the reverse Cuthill-McKee algorithm, which reduces the dependency between nodes and thus reduces the idle time of the pipeline.
[0068] The overall workflow is as follows:
[0069] 1. Input: adjacency matrix and node features of input file.
[0070] 2. Feature-level parallelism: Split node features by dimension and assign them to different sparse computing units for neighbor aggregation calculation.
[0071] 3. Fully connected communication: pass the aggregated features to the dense computing unit.
[0072] 4. Node-level parallelism: The dense computation unit updates the features of the nodes assigned to it based on the received features and model weights.
[0073] 5. Fully connected communication: pass the updated features to the sparse computing unit for the next layer of calculation.
[0074] 6. Repeat steps 2-5 until all layers are calculated.
[0075] 7. Back propagation and weight update.
[0076] In some embodiments,
[0077] In some embodiments, the resource virtualization of the memory can be refined, such as extending the resource virtualization to a more fine-grained level, such as dividing the memory into different virtual areas and allocating them according to task type and priority, so as to better utilize the high density characteristics of the embedded memory.
[0078] In some embodiments, fine-grained virtualization can be combined with non-intrusive partitioning of a multi-level partitioning algorithm. For example, non-intrusive partitioning of files can be achieved through a multi-level partitioning algorithm. The multi-level partitioning algorithm is regarded as a fine-grained resource virtualization method, and the file data partitioning is regarded as a dynamic allocation of computing resources. Traditional memory allocation is coarse-grained, which is prone to resource waste and contention. Fine-grained virtualization can better utilize the high-density characteristics of embedded memory and allocate storage space according to task requirements. Improve the efficiency of collaborative reasoning and reduce communication overhead caused by data sharing. The multi-level partitioning algorithm can be regarded as a virtualization method that dynamically allocates file data to different devices, avoids repeated data transmission between devices, and improves reasoning efficiency. Non-intrusive partitioning reduces communication delay and improves the overall performance of the system by reducing the amount of boundary data transmission required in convolution operations. Fine-grained virtualization can make this data allocation more accurate and efficient.
[0079] In some embodiments, the heterogeneous resource management module 20 is also used to reduce the frequency of the computing unit or select computing resources with power consumption lower than a preset power consumption for tasks with a delay requirement lower than a threshold.
[0080] In some embodiments, when scheduling resources, not only performance but also power consumption is considered. For example, for tasks with low latency requirements, the frequency of the computing unit can be reduced or computing resources with lower power consumption can be selected. For example, power consumption-aware resource management and dynamic computing unit allocation can be used. Specifically, power consumption awareness is introduced into dynamic computing unit allocation so that power consumption can be reduced as much as possible while ensuring performance. Embedded systems are usually very sensitive to power consumption and need to strike a balance between performance and power consumption. Power consumption-aware methods can improve system energy efficiency and extend the service life of equipment. Dynamically adjust the use of computing units according to the latency requirements of the task. For tasks with low latency requirements, computing units with lower power consumption can be used to reduce overall power consumption; otherwise, computing units with higher performance are selected. By dynamically adjusting the frequency and type of use of computing units, the power consumption of the system can be effectively reduced while meeting performance requirements. This dynamic adjustment relies on real-time monitoring and dynamic optimization of task loads.
[0081] In some embodiments, resource management of collaborative reasoning and parallel strategies of multi-level partitioning algorithms can be adopted. Specifically, the multi-level partitioning algorithm divides the file and assigns it to different devices for collaborative reasoning, which can further optimize the data transmission and calculation between devices in the multi-level partitioning algorithm. Collaborative reasoning requires efficient use of the computing and communication resources of multiple devices. Heterogeneous resource management extensions can help NPTP better manage these resources and improve the reasoning speed. Further reduce the communication delay of the multi-level partitioning algorithm. Through reinforcement learning or other optimization algorithms, the image partitioning strategy can be further optimized, the amount of data sharing between devices can be reduced, and the allocation of computing tasks can be optimized. Data parallelism, model parallelism or hybrid parallelism are common methods of collaborative reasoning. Through more intelligent resource allocation and scheduling, the delay caused by data transmission can be significantly reduced and the reasoning speed can be improved.
[0082] In some embodiments, the embedded hybrid computing semiconductor memory 100 further includes: a development framework support module for providing a programming abstraction level higher than a preset level, providing an automated code optimization tool, and a formal verification tool.
[0083] High-level programming abstraction mainly provides a higher level of programming abstraction, allowing developers to focus on algorithm logic without paying attention to the underlying hardware details. For example, a framework similar to TensorFlow or PyTorch can be developed to facilitate the development of AI applications.
[0084] In some embodiments, the automated code optimization tool is used for automated code optimization, equivalence verification using an instruction-level equivalence detection tool, and verification of optimization features for embedded systems.
[0085] Automated code optimization tools are mainly developed to optimize the code at compile time or runtime, making full use of the advantages of embedded memory. For example, data layout optimization, parallel computing optimization, etc. can be performed.
[0086] Step 1: Automated code optimization According to the characteristics and requirements of the embedded system, the following optimizations are performed:
[0087] Data layout optimization: Place frequently accessed data in cache or local memory to reduce data access latency.
[0088] Parallel computing optimization: Decompose the code into multiple tasks that are executed in parallel to improve computing throughput.
[0089] Optimization for specific compute units: If the embedded system contains specific compute units (such as neural network accelerators), optimize for those units.
[0090] On-chip data compression and decompression: Add hardware-accelerated data compression and decompression modules during data transmission. For data that needs to be read and written frequently, it can be compressed first to reduce storage usage and transmission bandwidth requirements, and reduce latency. These optimized codes may be very different from the original code on the surface, such as the storage location of variables is changed, the execution order is changed, the loop is unrolled, etc.
[0091] Step 2: Use instruction-level equivalence detection tools to perform equivalence verification. After the optimization is completed, it is necessary to verify whether the optimized code is still semantically equivalent to the original code. At this time, an instruction-level equivalence detection tool is needed. The specific process is as follows:
[0092] SSA transformation: The instruction-level equivalence checker converts both the original code and the optimized code into static single assignment (SSA) form. The SSA form ensures that each variable is assigned a value only once, which simplifies data flow analysis.
[0093] Construct dependency graph: The instruction-level equivalence detection tool constructs data flow dependency graph and control flow dependency graph based on the SSA form of code. Data flow dependency describes the data transfer relationship between instructions, and control flow dependency describes the order in which instructions are executed.
[0094] Generate lemmas: The instruction-level equivalence checker generates a series of lemmas based on data flow and control flow dependencies. A lemma is a proposition in formal logic that expresses the conditions for code equivalence. For example, if two instructions have the same operators and their dependencies are equivalent, then the two instructions are equivalent.
[0095] Inductive proof: Instruction-level equivalence detection tools use inductive proofs to determine whether two instructions are equivalent. The basic case is function parameters and constants. The inductive step uses the lemma. If all dependencies of the instruction are equivalent and the operators are the same, then the instructions themselves are also equivalent. For loop structures, instruction-level equivalence detection tools use special loop guidance methods to handle loop dependencies.
[0096] Equivalence alignment: The output of the instruction-level equivalence detection tool is an equivalence alignment, which clearly indicates which instructions in the original code and the optimized code are equivalent. If the instruction-level equivalence detection tool determines that the optimized code is equivalent to the original code, it means that the code optimization has not introduced semantic errors and can be safely deployed in the embedded system. If the instruction-level equivalence detection tool determines that the optimized code is not equivalent to the original code, then it is necessary to analyze the cause of the error, and it may be necessary to adjust the optimization strategy or fix the code error.
[0097] Step 3: Verify the optimization features of the embedded system.
[0098] Instruction-level equivalence detection tools can detect the impact of various optimization techniques on code semantics.
[0099] Data layout optimization: Even if the data layout in the code changes, as long as the logic is equivalent, the instruction-level equivalence detection tool can identify it.
[0100] Parallel computing optimization: Even if the execution order of the program changes, instruction-level equivalence checking tools can verify semantic equivalence based on data and control flow dependencies.
[0101] Loop optimization: The instruction-level equivalence detection tool handles loop dependencies through a loop guidance method to ensure that loop optimization does not introduce semantic errors.
[0102] Specific compute unit optimization: Instruction-level equivalence detection tools can verify whether the code after using a specific compute unit is equivalent to the original code.
[0103] There may be different computing units in an embedded system. Automated code optimization tools can customize optimizations for these units, and instruction-level equivalence detection tools can verify whether these optimizations maintain the original semantics of the code. This allows the optimized code to achieve better performance on specific hardware.
[0104] Improve optimization efficiency: Automated code optimization tools can automatically complete various optimization tasks, reducing the cost and time of manual optimization. These tools can automatically adjust the optimization strategy according to the characteristics of the embedded system to improve optimization efficiency. Instruction-level equivalence detection tools can automatically verify the optimized code, avoiding the tedious process of manual testing.
[0105] Ensure optimization correctness: Traditional code optimization methods may introduce errors that are difficult to detect manually. The equivalence verification of the instruction-level equivalence checker provides a formal and reliable method to ensure that the optimized code still maintains semantic correctness. During the optimization process, if an error occurs, the instruction-level equivalence checker can accurately point out the location and cause of the error, helping developers to quickly fix the error.
[0106] Improve code readability: Although the main purpose of the instruction-level equivalence detection tool is to verify equivalence, it can provide variable mapping information, which helps to understand the variable transformation during the optimization process. Understanding the correspondence of variables can help developers better understand the optimized code, thereby improving the readability of the code.
[0107] Optimization based on reinforcement learning: The introduction of reinforcement learning algorithms can make automated code optimization tools more intelligent and dynamically adjust optimization strategies according to different application scenarios. Reinforcement learning can learn the best optimization solution through trial and error and feedback. Instruction-level equivalence detection tools can be used as a feedback mechanism for reinforcement learning to evaluate the optimization effect and ensure that the optimized code is still semantically correct.
[0108] On-chip data compression / decompression: Combining on-chip data compression and decompression technology can reduce data transmission overhead and improve memory access efficiency. Instruction-level equivalence detection tools can verify whether the compression / decompression process introduces errors and ensure the correctness of the data.
[0109] Formal verification tools: Formal verification tools are provided to help developers verify the correctness and reliability of programs. This is very important for ensuring the stability of the system, especially in embedded systems.
[0110] Assistance to formal verification: Using the equivalence verification results of the instruction-level equivalence detection tool as the input of the formal verification tool can improve the coverage and accuracy of the verification.
[0111] In some embodiments, the embedded hybrid computing semiconductor memory 100 of the present application can be applied to data-intensive computing, reducing computing latency by 65%, increasing data throughput by 80%, reducing energy consumption by 45%, and increasing resource utilization by 50%.
[0112] The 65% reduction in computing latency is mainly due to significant reduction in computing latency through near-data computing and optimized data transmission paths.
[0113] The 80% increase in data throughput is mainly due to the improvement of data processing capabilities through heterogeneous computing integration and efficient resource management.
[0114] The 45% reduction in energy consumption is mainly due to intelligent scheduling and dynamic load balancing that optimizes system energy consumption.
[0115] The 50% increase in resource utilization is mainly reflected in the unified resource pooling and intelligent scheduling strategies that improve resource utilization efficiency.
[0116] In some embodiments, the embedded hybrid computing semiconductor memory 100 of the present application can be applied to AI acceleration applications, with model reasoning acceleration of 3-5 times, batch processing performance improved by 4 times, power consumption efficiency improved by 70%, and resource utilization optimized by 55%.
[0117] The 3-5 times acceleration of model reasoning is mainly reflected in the significant improvement of model reasoning speed through the integration of heterogeneous computing units such as GPU and FPGA.
[0118] The four-fold improvement in batch processing performance is mainly due to the optimization of task scheduling and data flow paths, which improves batch processing performance.
[0119] The 70% improvement in power consumption efficiency is mainly reflected in the use of near-data computing and intelligent resource management to optimize system energy consumption.
[0120] Resource utilization optimization of 55% is mainly reflected in unified resource management and intelligent scheduling strategies to ensure efficient use of resources.
[0121] That is, by combining the characteristics of embedded semiconductor memory, the new hybrid computing storage architecture can significantly improve the performance and efficiency of the system, reduce data transmission latency, optimize resource utilization, and meet growing computing needs.
[0122] In summary, the embedded hybrid computing semiconductor memory 100 provided by the present application can reduce data movement by allowing dynamic changes in data transmission paths to improve computing efficiency, optimize energy consumption ratio, and reduce system latency. It also uses heterogeneous resource management modules to perform flexible computing unit allocation, intelligent task scheduling, resource collaborative optimization, and a unified management framework.
[0123] In the several embodiments provided in this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation described above is only illustrative, for example, the division of the modules or units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0124] If the integrated units in the above other embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, RandomAccess Memory), disk or optical disk and other media that can store program codes.
[0125] The above description is only an implementation method of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly used in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. An embedded hybrid computing semiconductor memory, characterized in that: include: A computing storage fusion engine, comprising several types of computing units; wherein a reconfigurable interconnection network is formed between the computing units and the memory, allowing the data transmission path to be dynamically changed; A heterogeneous resource management module for dynamically adjusting the use of computing units according to real-time load based on reinforcement learning; The several types of computing units include deep learning neural network computing units, encryption algorithm computing units, FPGAs and GPUs, and pulse neural network computing units; The neural network computing unit of the deep learning includes a shallow embedding sub-model and a deep decision sub-model, the shallow embedding sub-model is used to process low-dimensional features, the deep decision sub-model is used to process high-dimensional features, and the shallow embedding sub-model and the deep decision sub-model are allocated different computing resources, the shallow embedding sub-model is allocated for global collaborative learning, and the deep decision sub-model is allocated for intra-cluster collaborative learning; The interconnection network comprises: An optical path switching regional interconnection architecture is used to realize the interconnection between computing units and memories by using optical path switching; Dynamically adjust the interconnection network module, which uses traffic monitors to track regional network requirements and uses a greedy algorithm to generate OCS topology and dynamically adjust optical path connections; The computing storage fusion engine also includes an on-chip data compression and decompression module, which is used to compress and / or decompress data during data transmission, and adopts hardware acceleration when compressing and / or decompressing data; The heterogeneous resource management module includes: A hybrid parallel module is used to perform neighbor aggregation using feature-level parallelism and node-level parallelism for node updating; the heterogeneous resource management module is also used to reduce the frequency of the computing unit or select computing resources with power consumption lower than a preset power consumption for tasks with a delay requirement lower than a threshold; A hybrid computing unit includes a sparse computing unit and a dense computing unit, wherein the sparse computing unit is used for sampling sparse matrix multiplication, and the dense computing unit is used for node update.
2. The embedded hybrid computing semiconductor memory according to claim 1, characterized in that: The sparse computing unit is used to select required neighbor features using a hybrid sparse indexing module and use local double buffering to pipeline data processing.
3. The embedded hybrid computing semiconductor memory according to claim 1, characterized in that: The embedded hybrid computing semiconductor memory also includes: a development framework support module, which is used to provide a programming abstraction level higher than a preset level, and to provide an automated code optimization tool and a formal verification tool.
4. The embedded hybrid computing semiconductor memory according to claim 3, characterized in that: The automated code optimization tool is used for automated code optimization, equivalence verification using an instruction-level equivalence detection tool, and verification of optimization characteristics for embedded systems.
Citation Information
Patent Citations
Method and system for improving computing power efficiency
CN118550711A