Real-time data processing method and chip for low-power AI processor
Through the combination of neuromimicry classification processing, multi-dimensional mapping, integrated storage and computing data path design, dynamic voltage frequency mapping and pulse neural network model, the problems of high power consumption and performance bottlenecks in real-time data processing are solved, and efficient and low-power computing capabilities are achieved.
Patent Information
- Application Number
- CN202510294735.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-03-13
AI Technical Summary
Existing AI processors have problems such as excessive power consumption, obvious performance bottlenecks, and low computing efficiency in real-time data processing, which is difficult to meet the needs of scenarios such as edge computing and mobile devices.
Through neuromimicry classification processing and multi-dimensional mapping operations, the computing task characteristics are extracted, core resource allocation is allocated based on heterogeneous computing space vectors, integrated data paths for storage and computing are built, dynamic voltage frequency mapping is performed, and pulsed neural network model is introduced for computing power load balancing calculation.
It significantly reduces system power consumption, improves computing efficiency, realizes intelligent resource scheduling of CPU and NPU core groups, fully utilizes the advantages of the dual-mode processing architecture, and improves data processing capabilities and overall system utilization.
Smart Images

Figure CN119806849B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of low-power AI processors, and in particular to a real-time data processing method and chip for a low-power AI processor. Background Art
[0002] With the rapid development of artificial intelligence technology, the demand for computing power is growing, especially in scenarios such as edge computing and mobile devices, which puts higher requirements on the performance and power consumption of AI processors. Traditional processor architectures have problems such as excessive power consumption, obvious performance bottlenecks, and low computing efficiency when processing AI tasks, making it difficult to meet actual application needs.
[0003] When performing real-time data processing, current AI processors mainly rely on a general random task scheduling model, which often cannot accurately adapt to the processing requirements in specific computing scenarios or application environments, leading to problems such as uneven resource allocation, low computing efficiency, and energy waste. At the same time, the existing processor architecture generally adopts a storage-computing separation design mode, which consumes a lot of energy during data transfer, seriously affecting the overall performance of the processor. In addition, the existing technology still has major deficiencies in processor power consumption control, lacks an effective dynamic power consumption management mechanism, and cannot flexibly adjust the working state according to the actual computing load, resulting in low energy utilization efficiency. At the same time, there is also a lack of intelligent load balancing mechanism in task scheduling, which cannot fully utilize heterogeneous computing resources, affecting the overall performance of the processor. Summary of the invention
[0004] The present invention provides a real-time data processing method and chip for a low-power AI processor, which are used to improve the real-time data processing capability of the low-power AI processor.
[0005] In a first aspect, the present invention provides a real-time data processing method for a low-power AI processor, the real-time data processing method for a low-power AI processor comprising:
[0006] Perform neuromorphic classification processing on the computing tasks input by the low-power AI processor to obtain the computing task feature matrix;
[0007] Performing a multi-dimensional mapping operation on the computing task feature matrix to obtain a heterogeneous computing space vector;
[0008] Performing core resource allocation based on the heterogeneous computing space vector to obtain a dual-mode processing unit configuration matrix;
[0009] Based on the dual-mode processing unit configuration matrix, a storage-computation integrated data path is constructed to obtain a parallel computing data flow graph;
[0010] Performing dynamic voltage-frequency mapping on the parallel computing data flow graph to obtain a power consumption optimization control vector;
[0011] The power consumption optimization control vector is input into the trained pulse neural network model to perform computing load balancing calculation and generate computing load balancing parameters.
[0012] In a second aspect, the present invention provides a real-time data processing chip of a low-power AI processor, wherein the real-time data processing chip of the low-power AI processor comprises:
[0013] The classification module is used to perform neuromorphic classification processing on the computing tasks input by the low-power AI processor to obtain the computing task feature matrix;
[0014] A computing module, used for performing a multi-dimensional mapping operation on the computing task feature matrix to obtain a heterogeneous computing space vector;
[0015] An allocation module, configured to allocate core resources based on the heterogeneous computing space vector to obtain a dual-mode processing unit configuration matrix;
[0016] A construction module, used to construct a storage-computation integrated data path based on the dual-mode processing unit configuration matrix to obtain a parallel computing data flow graph;
[0017] A mapping module, used for performing dynamic voltage-frequency mapping on the parallel computing data flow graph to obtain a power consumption optimization control vector;
[0018] A generation module is used to input the power consumption optimization control vector into the trained pulse neural network model to perform computing load balancing calculation and generate computing load balancing parameters.
[0019] The third aspect of the present invention provides a real-time data processing device for a low-power AI processor, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory so that the real-time data processing device for the low-power AI processor executes the above-mentioned real-time data processing method for the low-power AI processor.
[0020] A fourth aspect of the present invention provides a computer-readable storage medium, which stores instructions that, when executed on a computer, enable the computer to execute the real-time data processing method of the low-power AI processor described above.
[0021] In the technical solution provided by the present invention, through neuromorphic classification processing and multi-dimensional mapping operations, the feature extraction of computing tasks is more accurate, the task classification is more reasonable, and the task processing efficiency of the system is improved; based on the core resource allocation mechanism of heterogeneous computing space vectors, the intelligent resource scheduling of the CPU core group and the NPU core group is realized, and the advantages of the dual-mode processing architecture are fully utilized; the storage and computing integrated data path design is adopted to reduce the data handling overhead, optimize the data flow transmission efficiency, and significantly reduce the system power consumption; through the dynamic voltage frequency mapping technology, the precise regulation of the processor working state is realized; the pulse neural network model is introduced for computing power load balancing calculation, and event-driven asynchronous computing scheduling is realized, which improves the computing efficiency of the system; the multi-level cache and high-bandwidth design are adopted to achieve 16PB / s memory bandwidth, 3.5PB / s core-to-core communication bandwidth and 5TB / s chip-to-chip communication bandwidth, which greatly improves the data processing capability; the neural network structure constructed by neurons and synapses provides powerful parallel computing capabilities and realizes efficient task processing; dynamic migration and load balancing of tasks are realized, and excessive concentration or idleness of computing resources is avoided, which improves the overall utilization of the system.
[0022] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention are realized and obtained by the structures particularly pointed out in the description, claims and drawings.
[0023] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 Schematic diagram of an embodiment of a real-time data processing method of a low-power AI processor in an embodiment of the present invention;
[0025] Figure 2 This is a schematic diagram of an embodiment of a real-time data processing chip of a low-power AI processor in an embodiment of the present invention;
[0026] Figure 3 Schematic diagram of an embodiment of a real-time data processing device of a low-power AI processor in an embodiment of the present invention. DETAILED DESCRIPTION
[0027] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0028] The terms "including" and "having" and any variations thereof mentioned in the embodiments of the present invention are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device end including a series of steps or units is not limited to the listed steps or units, but may optionally include other steps or units that are not listed, or may optionally include other steps or units that are inherent to these processes, methods, products or device ends.
[0029] To facilitate understanding of this embodiment, a real-time data processing method of a low-power AI processor disclosed in an embodiment of the present invention is first described in detail. Figure 1 As shown, the method comprises the following steps:
[0030] 101. Perform neuromorphic classification processing on the computing tasks input by the low-power AI processor to obtain a computing task feature matrix;
[0031] It is understandable that the execution subject of the present invention can be a real-time data processing chip of a low-power AI processor, or a terminal or a server, which is not limited here. The embodiment of the present invention is described by taking the server as the execution subject as an example.
[0032] Specifically, the computing task input to the low-power AI processor is analyzed and calculated on a scale to obtain the complexity vector of the task. The task complexity vector includes a data scale parameter and a computational complexity parameter. The data scale parameter is used to indicate the size of the task data, and the computational complexity parameter is used to evaluate the computational difficulty of the task and the resources required for the operation. The task complexity vector is input into the neuromorphic classifier. The neuromorphic classifier uses the neuromorphic core to perform parallel classification operations on the computing tasks to obtain a neuron state distribution matrix. The use of the neuromorphic core makes the processing process more efficient and has the simulation characteristics of the biological nervous system. By processing the complex relationships in the task in a parallel manner, the response speed of the processor can be significantly improved. The neuron state distribution matrix describes the activation of each neuron when receiving the task complexity vector. Based on the neuron state distribution matrix, a pulse coding operation is performed. The pulse coding operation encodes information through neuron nodes to generate a neuron activation vector. The neuron activation vector is an expression of the activation characteristics of the task in the neural network. Based on the neuron activation vector, synaptic connection calculation is performed. The role of synapses in this process is like a bridge connecting neurons in a neural network. The synaptic calculation weight is updated to reflect how the characteristics of the task are propagated and processed in the network, and the synaptic weight vector is obtained. The synaptic weight vector is subjected to data flow dependency analysis to establish the relationship between the previous task and the subsequent task, clarify the dependency structure between tasks, and obtain the task dependency graph. The task dependency graph is input into the topological sorting processor for data caching to obtain the optimized task dependency vector. The topological sorting processor can effectively reduce the blocking phenomenon between tasks, so that the processor can process multiple tasks with a higher degree of parallelism to form an optimized task dependency vector. A three-dimensional feature space is constructed through the neuron activation vector, the synaptic weight vector and the optimized task dependency vector, and the feature fusion is performed using an asynchronous event-driven mechanism to obtain the initial task feature matrix. The initial task feature matrix is batch normalized, and the final calculation task feature matrix is obtained through dynamic scaling and translation operations. Batch normalization can eliminate the scale differences between different task features and avoid excessive influence of certain features on the processing results. Dynamic scaling and translation operations enhance the stability of task features, so that the feature matrix can better adapt to subsequent calculation steps.
[0033] 102. Perform multi-dimensional mapping operation on the computing task feature matrix to obtain a heterogeneous computing space vector;
[0034] Specifically, the feature decomposition operation is performed on the computing task feature matrix to extract the main features of the task and generate the computing feature principal component matrix. The feature matrix is decomposed into a low-dimensional matrix representing the most important features by the principal component analysis method. Based on the computing feature principal component matrix, a task similarity measurement space is constructed, and the similarity between different tasks is measured by the measurement space to obtain the task similarity matrix. According to the task similarity matrix, the core affinity is calculated to determine the execution priority of the task on different cores of the processor respectively, and the CPU core affinity vector and the NPU core affinity vector are obtained. By combining the task characteristics with the characteristics of the processor core, the adaptability of each task on the CPU and NPU is determined, and two types of core affinity vectors are generated. The affinity vector indicates the degree of fit between the task and the different computing cores, which helps to improve the processor resource utilization and ensure that the task is assigned to the most suitable computing unit. Based on the CPU core affinity vector and the NPU core affinity vector, the computing load distribution analysis is performed to generate the load distribution mapping matrix. The load distribution mapping matrix is used to describe the task load between different computing units, ensuring that each task can get reasonable computing resource support and avoiding the situation where some cores are overloaded while other cores are idle, so as to achieve load balancing. The load distribution mapping matrix is subjected to dimension compression operation to simplify the computational complexity and retain the most important load information, and the computational dimension projection vector is obtained. The dimension compression operation can effectively reduce the computational burden, so that the processor can process tasks more efficiently in subsequent steps. The task relevance calculation is performed according to the computational dimension projection vector, and the relevance between tasks is calculated by the cosine similarity function to generate a task relevance vector. Based on the task relevance vector, a heterogeneous computing resource graph is constructed. The heterogeneous computing resource graph is an abstract representation of the resources of the entire system. By matching tasks with different types of computing resources, the optimal resource allocation scheme is determined and a computing resource allocation vector is generated. The computing resource allocation vector specifically represents the distribution of tasks on heterogeneous computing resources, ensuring that each task can be assigned to the most suitable computing unit for processing, thereby improving the overall computing efficiency and response speed of the system. In order to optimize the processing of computing tasks, the computing resource allocation vector is transformed through spatial coordinates to obtain a heterogeneous computing space vector. The process of spatial coordinate transformation converts the resource allocation vector from the task feature space to the heterogeneous computing space, so that the characteristics of the task can be better matched and adapted to the heterogeneous computing resources.
[0035] 103. Perform core resource allocation based on the heterogeneous computing space vector to obtain a dual-mode processing unit configuration matrix;
[0036] Specifically, the heterogeneous computing space vector is hierarchically divided into resources, and the computing tasks are assigned to different processor units. The tasks are divided according to their dependence characteristics on the CPU and NPU to obtain the CPU task set and the NPU task set. Based on the demand characteristics of the computing tasks, it is reasonably determined which tasks should be processed by the CPU and which tasks should be handed over to the NPU to ensure that the tasks can fully utilize the advantages of different processor units. The CPU task set contains tasks that require high computing accuracy and low parallelism, while the NPU task set includes tasks that require large-scale parallel computing and high throughput. Through division, the resource utilization of each processor unit can be improved. According to the CPU task set, the utilization analysis of the neuromorphic core is performed, the load of the CPU core when processing these tasks is evaluated, and the CPU core group allocation matrix is obtained. The CPU core group allocation matrix reflects the workload distribution of each CPU core when processing a specified task, helps the system to reasonably plan task allocation, avoid overloading some cores while other cores are idle, and improve the computing efficiency of the system. For the NPU task set, a similar computing power analysis is performed. The NPU core is used to process high-parallel computing tasks such as deep learning. The computing power of the NPU task set is analyzed, and the computing resource requirements of the NPU core when processing these tasks are evaluated to obtain the NPU core group allocation matrix. The NPU core group allocation matrix can effectively reflect the distribution of tasks on different NPU cores, help the system to reasonably allocate computing tasks, and ensure that the NPU core can perform efficient computing processing according to the requirements of the tasks. Cache capacity is allocated based on the CPU core group allocation matrix and the NPU core group allocation matrix. According to the distribution of tasks among the cores, cache resources are reasonably allocated to improve data access efficiency and reduce cache conflicts and bandwidth bottlenecks. The result of cache capacity allocation is a cache allocation vector, which represents the allocation of cache resources among different computing cores. On this basis, the cache allocation vector is divided into 192KB units to obtain cache unit groups. Each cache unit group contains the specific allocation of cache, which is used to effectively manage the cache requirements of different cores. Memory bandwidth is allocated according to the cache unit group to ensure that different tasks can be efficiently executed within the required memory bandwidth range. Memory bandwidth allocation is not only the management of cache capacity, but also requires reasonable adjustment of bandwidth allocation ratio according to the task execution mode and data access mode to avoid bottlenecks in certain tasks due to insufficient bandwidth. Through memory bandwidth allocation, a bandwidth allocation matrix is obtained to reflect the bandwidth allocation between each core and memory, ensuring that data flows can flow efficiently between the processor and memory. Based on the bandwidth allocation matrix, processing unit combination operations are performed to obtain the initial processing unit configuration matrix. According to the cache and bandwidth allocation results, the computing core and storage resources are reasonably combined to achieve efficient execution of tasks. The initial processing unit configuration matrix reflects the matching between tasks and processing units under specific resource allocation.The initial processing unit configuration matrix is optimized for resource balancing to obtain the final dual-mode processing unit configuration matrix. The goal of resource balancing optimization is to avoid resource waste and ensure load balancing of computing units by further adjusting tasks and resources. During the optimization process, the allocation ratio of each core is dynamically adjusted according to the priority of the task, the load of the computing core, and the demand for memory bandwidth to generate a configuration matrix that achieves the optimal balance between performance and power consumption.
[0037] 104. Construct a storage-computation integrated data path based on the dual-mode processing unit configuration matrix to obtain a parallel computing data flow graph;
[0038] Specifically, the data flow analysis of the dual-mode processing unit configuration matrix is performed to extract the data interaction relationship between different processing units. The data flow between the CPU core group and the NPU core group is understood, and the computing unit communication matrix is obtained through this analysis. The computing unit communication matrix records the data exchange path and interaction frequency between each core group. By clarifying the communication relationship between different core groups, data transmission is accurately planned to avoid data flow conflicts and bandwidth bottlenecks. Based on the computing unit communication matrix, a memory access topology map is constructed to describe the data access mode between each computing core and the memory, and a memory bandwidth allocation map is generated. The memory bandwidth allocation map ensures that each core obtains the corresponding memory bandwidth according to its computing load by segmenting the total bandwidth of 16PB / s. Segmentation can effectively manage the memory access requirements of different cores, avoid bandwidth contention, and ensure that each task in parallel computing can get sufficient bandwidth support, thereby improving computing efficiency. The storage units in the core are mapped according to the memory bandwidth allocation map to generate the in-core computing channel. The in-core computing channel includes read and write ports of 192KB cache, which are used to quickly read and write data during the computing process and provide efficient cache support. Through this mapping, efficient data exchange is achieved between the computing core and the storage unit, reducing data access latency and improving processing speed. The role of the cache is to provide a fast data access channel, reduce the waiting time of the task, and ensure that the computing core can process data efficiently. Data flow scheduling is performed based on the computing channel within the core to generate inter-core data channels. The inter-core data channel connects different core groups through a bandwidth of 3.5PB / s to ensure that data can be quickly transmitted between different computing units. The purpose of data flow scheduling is to optimize the data transmission path between multiple computing cores, avoid network congestion and bandwidth waste, and thus improve the computing performance of the overall system. The design of the inter-core data channel needs to take into account the computing needs and data interaction frequency of each core group, and ensure that the computing task can proceed smoothly by reasonably scheduling the data flow. The inter-core data channel is analyzed for parallelism to obtain a parallel communication topology diagram, depict the data exchange path between multiple cores, and clarify the communication method and communication frequency between each core. Through parallelism analysis, the data flow between multiple cores is optimized, so that parallel computing can be better supported. The parallel communication topology diagram shows how the cores work together to ensure that tasks can be processed in parallel and maximize the use of the system's computing resources. Based on the parallel communication topology, a chip-level interconnection network is constructed to support efficient data transmission between chips, ensuring that high-speed data exchange can be maintained between computing units in large-scale computing tasks. The chip-to-chip interaction channel supports 5TB / s data transmission, which can meet the needs of large-scale parallel computing and ensure that data can be quickly transmitted at the chip level to avoid bottlenecks and delays. The data flow of the intra-core computing channel, the inter-core data channel, and the inter-chip interaction channel are integrated to obtain the initial data flow graph.Integrate data transmission paths at different levels into a unified flow graph to better understand the flow of data between various computing units. Perform data flow optimization calculations based on the initial data flow graph. By adjusting the data flow path, optimize the order and method of data transmission, reduce redundant transmission and delay, and finally obtain a parallel computing data flow graph.
[0039] 105. Perform dynamic voltage-frequency mapping on the parallel computing data flow graph to obtain a power consumption optimization control vector;
[0040] Specifically, the computing load distribution analysis is performed on the parallel computing data flow graph, and the core load matrix is obtained according to the distribution of tasks on each computing core. The core load matrix contains the real-time computing load data of the CPU core group and the NPU core group, which reflects the computing workload undertaken by different cores in actual operation. By analyzing the core load, accurate feedback on the current computing requirements of each computing core is obtained. The voltage-frequency correlation calculation is performed based on the core load matrix to obtain a voltage-frequency mapping table. The voltage-frequency mapping table contains a voltage range of 0.8V to 1.2V and a frequency range of 1GHz to 3GHz. This mapping table reflects the relationship between voltage and frequency, and provides the voltage and frequency settings of the core under different working conditions according to the load conditions. The purpose of the voltage-frequency correlation calculation is to determine the optimal ratio of voltage and frequency required for each core under different loads, so as to ensure that the system can meet the performance requirements while minimizing power consumption. Based on the voltage-frequency mapping table, the power consumption characteristics are analyzed to obtain the initial power consumption control parameters. According to the change law of voltage and frequency, the power consumption performance of the system under different voltage-frequency combinations is evaluated. The initial power consumption control parameters include the core temperature threshold and the power consumption threshold, which are used to determine whether the core is in an overheated or overloaded state. When the core load is too high, the temperature may rise, affecting the stability and long-term performance of the system. By setting the power consumption threshold and temperature threshold reasonably, the working state of the computing core is monitored in real time to avoid excessive power consumption or damage to the hardware. A dynamic adjustment model is constructed based on the initial power consumption control parameters. The dynamic adjustment model enables the system to maintain the optimal working state under different load conditions by adjusting the voltage and frequency in real time. Through this model, the voltage and frequency settings are dynamically adjusted to adapt to different computing needs. According to the dynamic adjustment model, a voltage and frequency adjustment matrix is obtained, which records the optimal voltage and frequency settings under different load conditions. The voltage and frequency adjustment matrix is dynamically divided into a voltage adjustment subspace and a frequency adjustment subspace. The adjustment range of voltage and frequency is subdivided into multiple subspaces, so that more precise adjustments can be made in each subspace. The voltage adjustment subspace and the frequency adjustment subspace represent the optimal adjustment areas of voltage and frequency under different working conditions, respectively. On this basis, the core voltage adjustment coefficient is calculated according to the voltage adjustment subspace to obtain the voltage control vector. The voltage control vector corresponds to the voltage adjustment strategy of different cores, providing each core with a voltage adjustment scheme for load changes. According to the core load, the voltage control vector dynamically adjusts the voltage setting of each core to ensure that the power consumption of each core is optimized under different loads, while avoiding energy waste caused by overpowering. Similarly, the system calculates the operating frequency adjustment coefficient based on the frequency adjustment subspace to obtain the frequency control vector. The frequency control vector defines the frequency adjustment strategy of each core under different load conditions and provides the best frequency setting for different cores.By dynamically adjusting the frequency, the system can reduce unnecessary high-frequency operation according to the computing requirements of the task, thereby reducing power consumption while ensuring computing performance. The voltage control vector and the frequency control vector are jointly optimized to obtain the final power optimization control vector. Through joint optimization, while considering voltage and frequency, the overall power consumption is optimized without affecting system performance. The optimized power control vector combines the optimal settings of voltage and frequency, and can accurately adjust the working state of each computing core for different workloads to minimize power consumption and ensure that the low-power AI processor maintains low energy consumption while running efficiently.
[0041] 106. Input the power consumption optimization control vector into the trained pulse neural network model to perform computing load balancing calculation and generate computing load balancing parameters.
[0042] Specifically, asynchronous event encoding is performed on the power optimization control vector to obtain a pulse event sequence. The pulse event sequence contains the discharge timing data of neurons, reflecting the response of neurons to input stimuli. Through asynchronous event encoding, the computing loads of different cores are converted into discharge events of neurons to form timing data. A pulse neural network structure is constructed based on the pulse event sequence to obtain a network topology matrix. The network topology matrix connects multiple neuron layers through synapses to form a complex computing graph. The network structure can capture the dependencies between computing tasks and effectively simulate the interactions between cores in large-scale parallel computing. According to the network topology matrix, timing dependency processing is performed to obtain a neuron dynamic response model. The neuron dynamic response model describes the response characteristics of neurons after receiving input stimuli, characterizes how neurons respond to input signals of different intensities and frequencies, and the timing characteristics of neuron discharge. Through this model, the response behavior of each computing core under different load conditions is more accurately predicted to provide support for load balancing. Based on the neuron dynamic response model, a network training operation is performed to obtain a trained pulse neural network model. The training process optimizes the parameters in the neural network so that it can process asynchronous computing tasks based on an event-driven mechanism. In the processing of asynchronous computing tasks, each neuron decides whether to discharge according to the arrival time of the input event. Therefore, the neural network can handle highly dynamic and asynchronous computing tasks and simulate the computing load changes in parallel computing. The trained neural network can flexibly respond to the needs of real-time computing tasks. The trained spiking neural network model is used to perform load analysis to obtain the core load balancing vector. The core load balancing vector records the load distribution information of the CPU core group and the NPU core group. This vector can reflect the workload of each computing core and help the system understand which cores are busy when performing computing tasks and which cores are idle. By analyzing the load balancing vector, the load imbalance is identified. According to the core load balancing vector, a dynamic task migration strategy is constructed to obtain the task migration matrix. The task migration matrix contains the task scheduling scheme across the core group, which aims to migrate computing tasks from the core with heavier load to the core with lighter load according to the current load distribution to achieve load balancing. Through dynamic task migration, the performance degradation of some cores due to excessive load is avoided, while the utilization of other cores is improved to achieve the optimal use of overall computing resources. Based on the task migration matrix, the load dynamic balancing calculation is performed to obtain the initial balancing parameters. The initial balancing parameters include the allocation ratio of computing resources, which determines the amount of tasks that each core should undertake during the load balancing process. Through these initial parameters, the resource allocation is reasonably adjusted to ensure that the load of each computing core is in the best state. On this basis, the initial balancing parameters are optimized to generate the final computing power load balancing parameters.By optimizing calculations, the load balancing effect is improved, ensuring that the computing load of each core is within the most appropriate range, while minimizing energy consumption and improving the computing performance of the overall system.
[0043] In the embodiment of the present invention, through neuromorphic classification processing and multi-dimensional mapping operations, the feature extraction of computing tasks is more accurate, the task classification is more reasonable, and the task processing efficiency of the system is improved; based on the core resource allocation mechanism of heterogeneous computing space vectors, the intelligent resource scheduling of the CPU core group and the NPU core group is realized, and the advantages of the dual-mode processing architecture are fully utilized; the storage and computing integrated data path design is adopted to reduce the data handling overhead, optimize the data flow transmission efficiency, and significantly reduce the system power consumption; through the dynamic voltage frequency mapping technology, the precise regulation of the processor working state is realized; the pulse neural network model is introduced for computing power load balancing calculation, and event-driven asynchronous computing scheduling is realized, which improves the computing efficiency of the system; the multi-level cache and high-bandwidth design are adopted to achieve 16PB / s memory bandwidth, 3.5PB / s core-to-core communication bandwidth and 5TB / s chip-to-chip communication bandwidth, which greatly improves the data processing capability; the neural network structure constructed by neurons and synapses provides powerful parallel computing capabilities and realizes efficient task processing; dynamic migration and load balancing of tasks are realized, avoiding excessive concentration or idleness of computing resources, and improving the overall utilization of the system.
[0044] In a specific embodiment, the process of executing step 101 may specifically include the following steps:
[0045] Perform scale analysis and calculation on the computing tasks input to the low-power AI processor to obtain a task complexity vector, which includes a data scale parameter and a computing complexity parameter.
[0046] The task complexity vector is input into the neuromorphic classifier, and the computational tasks are classified in parallel through the neuromorphic core to obtain the neuron state distribution matrix;
[0047] Perform pulse coding operation according to the neuron state distribution matrix, encode information through neuron nodes to obtain neuron activation vectors, calculate synaptic connections based on the neuron activation vectors, update weights through synapses to obtain synaptic weight vectors;
[0048] Perform data flow dependency analysis on the synaptic weight vector, obtain a task dependency graph by establishing a mapping relationship between the preceding task and the subsequent task, and input the task dependency graph into the topological sorting processor for data caching to obtain an optimized task dependency vector;
[0049] A three-dimensional feature space is constructed based on the neuron activation vector, synaptic weight vector and optimized task dependency vector, and feature fusion is performed through an asynchronous event-driven mechanism to obtain the initial task feature matrix.
[0050] Batch normalization operation is performed according to the initial task feature matrix, and the calculation task feature matrix is obtained through dynamic scaling and translation operations.
[0051] Specifically, the scale of the computing task input to the low-power AI processor is analyzed and calculated to obtain a task complexity vector, which contains a data scale parameter and a computational complexity parameter. The data scale parameter is determined by factors such as the dimension and amount of the input data. For example, in image recognition tasks, the data scale parameter is defined based on the number of pixels and channels of the image. The computational complexity parameter is measured based on the algorithmic complexity of the computing task. For example, in convolutional neural networks in deep learning, the computational complexity is related to factors such as the size of the convolution kernel, the size of the input image, and the number of network layers. The task complexity vector represents the two core attributes of the computing task: data scale and computational complexity. The following task complexity vector is obtained through calculation:
[0052] ;
[0053] in, Indicates the data size, Indicates computational complexity. Data size By the dimensions of the input data and sample size Calculations, e.g. , and the computational complexity It is the number of operations in the task. For example, in the forward propagation of a deep neural network, the computational complexity is estimated by the number of layers in the network, the number of neurons in each layer, and the type of operation. The task complexity vector is input into the neuromorphic classifier, and the neuromorphic core performs parallel classification operations on the computing tasks to obtain the neuron state distribution matrix. The neuromorphic core simulates the working principle of brain neurons. Through the input task complexity vector, the neuromorphic classifier judges the characteristics of the task and decides how to assign the task to different processing units. In this process, the neuron state distribution matrix reflects the response status of different computing units (such as CPU cores, NPU cores, etc.) to the task. Neuron state distribution matrix It is expressed as:
[0054] ;
[0055] in, Indicates The activation state of a neuron, is the total number of neurons. The activation state of each neuron reflects the responsiveness of the processor when facing different computing tasks, which determines how to classify and allocate tasks in parallel. According to the neuron state distribution matrix, pulse coding operation is performed. Pulse coding operation encodes the input task information through neuron nodes to obtain a neuron activation vector. The pulse coding method determines the size and nature of the input signal according to the discharge timing of the neuron. The neuron activation vector represents the response intensity of the neuron to each input signal during the task execution process, which is expressed as:
[0056] ;
[0057] in, Indicates The activation value of each neuron is quantified by the firing sequence of the neuron. These activation values reflect the degree of response of the neuron to the task. Based on the neuron activation vector, the synaptic connection calculation is performed, and the connection weights between neurons are updated through synapses to obtain the synaptic weight vector. The synaptic weight represents the strength or influence of the signal transmitted between neurons. According to the learning rules of the neural network (such as Hebbian learning or back propagation algorithm), the update of the synaptic weight is expressed as:
[0058] ;
[0059] in, Represents neurons To the neuron The synaptic weights of is the learning rate, and Represents neurons and neurons The synaptic weight update process optimizes the information transfer between neurons, allowing the processor to effectively process input signals and optimize the calculation path when executing tasks. After the synaptic weight is updated, data flow dependency analysis is performed, and the task dependency graph is obtained by establishing a mapping relationship between the previous task and the subsequent task. The task dependency graph reflects the order of tasks in the execution process and the dependencies between tasks. In this process, by establishing the data flow relationship between tasks, the task dependency graph is represented as:
[0060] ;
[0061] in, Represents a set of tasks, Represents the dependency relationship (edge) between tasks. Each task node represents a computing task, and the edge represents the order or data dependency between tasks. This graph helps the system to perform efficient scheduling when processing tasks and avoid computing bottlenecks caused by improper dependencies. After the task dependency graph is input into the topological sorting processor, the tasks are optimized through topological sorting to obtain the optimized task dependency vector. The optimized task dependency vector represents the best order for task execution, which is calculated through the graph sorting algorithm to ensure that there are no unnecessary delays or conflicts during the calculation process. The optimized task dependency vector is expressed as:
[0062] ;
[0063] in, Represents the execution order or priority of tasks. The optimization of task dependencies improves computing efficiency and makes the transmission of data streams more efficient, avoiding redundant calculations caused by unclear dependencies during the calculation process. A three-dimensional feature space is constructed based on the neuron activation vector, synaptic weight vector, and optimized task dependency vector. And through an asynchronous event-driven mechanism, these features are dynamically fused according to the occurrence of events to generate an initial task feature matrix. The initial task feature matrix reflects all key information of the computing task, including data size, computational complexity, task dependencies, and other dimensions. This matrix is expressed as:
[0064] ;
[0065] in, Represents the characteristics of the task, including multi-dimensional information, such as data size, computing requirements, dependencies, etc. These characteristics help the processor make more accurate decisions when executing tasks. According to the initial task feature matrix, batch normalization operation is performed, and the computing task feature matrix is obtained through dynamic scaling and translation operations. Batch normalization is a technology in deep learning. Through scaling and offset operations, the task feature matrix is standardized to eliminate the differences between features and avoid the problem of gradient disappearance or explosion during training. The normalized computing task feature matrix is expressed as:
[0066] ;
[0067] in, is the mean of the feature matrix, is the standard deviation, Represents the normalized feature matrix. Through normalization, subsequent computing task scheduling and optimization can be performed more effectively.
[0068] In a specific embodiment, the process of executing step 102 may specifically include the following steps:
[0069] Performing feature decomposition operation on the computing task feature matrix to obtain the computing feature principal component matrix, and constructing the task similarity measurement space based on the computing feature principal component matrix to obtain the task similarity matrix;
[0070] Perform core affinity calculation according to the task similarity matrix to obtain a CPU core affinity vector and an NPU core affinity vector, and perform computational load distribution analysis based on the CPU core affinity vector and the NPU core affinity vector to obtain a load distribution mapping matrix;
[0071] Perform dimension compression operation on the load distribution mapping matrix to obtain the calculation dimension projection vector, and perform task relevance calculation based on the calculation dimension projection vector, and obtain the task relevance vector through the cosine similarity function;
[0072] A heterogeneous computing resource graph is constructed based on the task correlation vector to obtain a computing resource allocation vector, and the computing resource allocation vector is transformed into a spatial coordinate to obtain a heterogeneous computing space vector.
[0073] Specifically, the feature matrix of the computing task is decomposed to extract the main features of the task. The feature decomposition operation is a common operation in linear algebra. By performing principal component analysis on the feature matrix, the main components of the computing task are extracted, and the task similarity measurement space is constructed based on these components. Suppose there is a computing task feature matrix , contains various characteristics of different computing tasks, such as data size, computational complexity, dependencies, etc. In order to perform feature decomposition, singular value decomposition or principal component analysis is performed on the feature matrix. Assuming that the principal component analysis method is used, the feature matrix is calculated The covariance matrix of :
[0074] ;
[0075] in, is the number of tasks, is the transpose of the feature matrix. By performing eigenvalue decomposition on the covariance matrix, eigenvalues and eigenvectors are obtained, among which the eigenvectors constitute the calculation feature principal component matrix , represents the main characteristic direction of the computing task. Characteristic principal component matrix It is expressed as:
[0076] ;
[0077] in, Indicates principal components, is the number of principal components. The variance value of each principal component reflects the proportion of the principal component in the feature space, and the main features of the computing task are measured according to these principal components. A task similarity metric space is constructed based on the principal component matrix of the computing features. By calculating the distance between tasks in the principal component space, a task similarity matrix is obtained :
[0078] ;
[0079] Among them, represents the similarity between task and task , and the cosine similarity is used to measure the similarity between tasks. The task similarity matrix is used for subsequent core affinity calculations to help the system identify which tasks are computationally similar, so that reasonable grouping and scheduling can be performed when allocating computing resources. According to the task similarity matrix , core affinity calculations are performed. The affinity of tasks on CPU cores and NPU cores is determined according to the similarity of tasks. Assuming that CPU cores and NPU cores have different computing capabilities and task adaptabilities, the appropriate tasks are allocated to different cores by calculating the affinity between tasks and cores. Calculate the core affinity vectors and are expressed as:
[0080] ;
[0081] Among them, and represent the adaptabilities of each task on the CPU and NPU respectively. Through the above affinity vectors, the priority allocation of different tasks on CPU and NPU cores is determined. According to the CPU core affinity vector and the NPU core affinity vector , a computational load distribution analysis is performed to obtain a load distribution mapping matrix . The load distribution mapping matrix is used to describe the load distribution of each core among different computing tasks. Assuming that the load distribution mapping matrix represents the allocation relationship of tasks among different cores, where represents the load ratio of task on core . The load distribution mapping matrix is calculated as follows:
[0082] ;
[0083] After obtaining the load distribution mapping matrix, a dimension compression operation is performed to reduce the computational burden and increase the processing speed. Dimension compression is to obtain a more concise calculation dimension projection vector by retaining the key relationship between tasks and reducing the influence of irrelevant features. The goal of dimensionality compression is to map a multidimensional feature space into a low-dimensional space, thereby reducing the need for computing resources. The dimensionality compression process is performed using methods such as PCA, t-SNE, or AutoEncoder, where the calculated dimensionality projection vector is expressed as:
[0084] ;
[0085] In the calculation dimension projection vector After that, the cosine similarity function is used to calculate the task relevance. The cosine similarity function measures the correlation between tasks in the projection space, and the calculation formula is:
[0086] ;
[0087] in, and Represent the feature vectors of the two tasks respectively, and Represent their modulus lengths respectively. By calculating the cosine similarity between tasks, we can get the task association vector :
[0088] ;
[0089] Task relevance vector It is used to describe the similarities between tasks in terms of computing resource requirements, computing methods, and scheduling strategies. , build a heterogeneous computing resource graph , which describes the relationship between tasks and computing resources. The heterogeneous computing resource graph consists of nodes and edges. Nodes represent tasks or computing resources, and edges represent the relationship between tasks and computing resources. By constructing a heterogeneous computing resource graph, computing resources can be reasonably allocated according to the computing requirements of tasks and the availability of resources. Computing resource allocation vector It is expressed as:
[0090] ;
[0091] in, Represents the computing resource allocation vector, which indicates the computing resource requirements corresponding to each task. Provide decision-making basis for subsequent computing task scheduling and optimization. By performing spatial coordinate transformation on the computing resource allocation vector A, the task scheduling is mapped to a specific computing space to obtain the heterogeneous computing space vector E. The heterogeneous computing space vector E describes the allocation of different tasks on different computing resources. The computing space transformation is implemented through matrix multiplication, nonlinear mapping or deep learning algorithm, and finally an optimized computing resource allocation scheme is obtained to ensure that the system can efficiently complete task processing under low power consumption conditions.
[0092] In a specific embodiment, the process of executing step 103 may specifically include the following steps:
[0093] Perform hierarchical resource partitioning on the heterogeneous computing space vector to obtain a CPU task set and an NPU task set;
[0094] Analyze the utilization of neuromorphic cores according to the CPU task set and obtain the CPU core group allocation matrix;
[0095] Analyze the computing power of the neuromorphic cores according to the NPU task set and obtain the NPU core group allocation matrix;
[0096] Allocate cache capacity based on the CPU core group allocation matrix and the NPU core group allocation matrix to obtain a cache allocation vector;
[0097] The cache allocation vector is divided into 192KB units to obtain cache unit groups, and memory bandwidth is allocated according to the cache unit groups to obtain a bandwidth allocation matrix;
[0098] The processing unit combination operation is performed based on the bandwidth allocation matrix to obtain an initial processing unit configuration matrix, and the initial processing unit configuration matrix is optimized for resource balancing to obtain a dual-mode processing unit configuration matrix.
[0099] Specifically, the heterogeneous computing space vector is divided into layers and resources are allocated to different computing units, including CPU and NPU. Describes the distribution of computing tasks on heterogeneous computing platforms, including the mapping information between tasks and computing units. By analyzing the vector, tasks are divided into two main sets: CPU task set and NPU task set according to the nature of computing tasks, computing requirements and characteristics of computing units. This division is achieved by classifying and analyzing the characteristics of computing tasks. Clustering algorithms (such as K-means or hierarchical clustering) are used to analyze the task characteristics and assign tasks to different core groups according to the complexity of the task, data requirements and execution mode. Assume that the computing task feature matrix After feature extraction, we get the task complexity and data requirement parameters. Through algorithm analysis, we get the CPU task set. and NPU task set , which includes tasks executed on the CPU and NPU cores respectively. This process is represented as follows:
[0100] ;
[0101] Suitable for NPU execution ;
[0102] Analyze the CPU task set, especially the utilization analysis of the neuromorphic core. The goal of utilization analysis is to evaluate the load of each CPU core when executing tasks and determine the specific allocation of tasks by calculating the core resource utilization. For each task The execution on the CPU core group is expressed by the following formula:
[0103] ;
[0104] in, Indicates the utilization of the CPU core, Indicates the load of the CPU core, is the computational complexity of the task, Represents the total computing power of the CPU core. By calculating the utilization of all tasks on different cores, the CPU core group allocation matrix is obtained: , which represents the distribution of each task on the CPU core group. It is expressed as:
[0105] ;
[0106] in, Indicates the task In the The distribution ratio on the CPU cores. According to the computing power of the core and the load of the task, the task is reasonably distributed to multiple cores to ensure that the load of each core is as balanced as possible to improve computing efficiency. Similarly, according to the NPU task set Analyze the computing power of the neuromorphic core. Unlike CPU cores, NPU cores have stronger parallel computing capabilities and higher throughput. When assigning NPU tasks, consider the computing power of each NPU core and the adaptability of the task. By calculating each task The computing power requirements on the NPU cores are compared with the computing power of the NPU cores to obtain the NPU core group allocation matrix , which represents the distribution of tasks on the NPU core group. The formula is as follows:
[0107] ;
[0108] in, Indicates the computing power of the NPU core. By calculating the allocation ratio of each task on the NPU core group, the NPU core group allocation matrix is obtained. , similar to the CPU core group allocation matrix .matrix It is expressed as:
[0109] ;
[0110] Based on CPU core group allocation matrix and NPU core group allocation matrix Cache capacity allocation is performed to ensure that each computing unit can obtain sufficient cache resources to reduce access latency and improve computing efficiency. Cache capacity allocation is based on the core load and the computing requirements of the task. In this process, based on the obtained core group allocation matrix, the cache allocation vector is calculated by the following formula :
[0111] ;
[0112] in, Indicates the allocation of cache. and They are the CPU and NPU core group allocation matrices, and is the computational requirement of the task set, and is the computing power of each computing unit. By calculating the cache allocation vector, reasonable cache resources are allocated to different computing units. In order to optimize the utilization efficiency of the cache, the cache allocation vector The 192KB unit is divided into cache unit groups. Each cache unit will be allocated a certain amount of memory bandwidth to meet the needs of the computing task. The cache unit group division is expressed by the following formula:
[0113] ;
[0114] in, Represents the cache unit group. Memory bandwidth is allocated according to the cache unit group to obtain the bandwidth allocation matrix The bandwidth allocation matrix represents the allocation of memory bandwidth shared between different cores. The bandwidth allocation matrix is calculated using the following formula:
[0115] ;
[0116] Bandwidth Allocation Matrix Ensure that each computing unit obtains appropriate bandwidth resources according to load and cache requirements. Based on the bandwidth allocation matrix, perform processing unit combination operations to obtain the initial processing unit configuration matrix The initial processing unit configuration matrix represents the resource configuration required by each computing unit when processing tasks. The calculation of the processing unit combination is expressed by the following formula:
[0117] ;
[0118] By performing resource balancing optimization on the initial processing unit configuration matrix, the dual-mode processing unit configuration matrix is obtained. The goal of balanced optimization is to distribute the load of different cores as evenly as possible, avoid overload or idleness, and ensure that the processing unit can run efficiently with the lowest power consumption. Dual-mode processing unit configuration matrix Represents the final processing unit configuration, ensuring that the CPU and NPU cores can balance loads, allocate resources reasonably, and maximize system performance when performing tasks.
[0119] In a specific embodiment, the process of executing step 104 may specifically include the following steps:
[0120] Perform data flow analysis on the dual-mode processing unit configuration matrix to obtain the computing unit communication matrix, which records the data interaction relationship between the CPU core group and the NPU core group;
[0121] A memory access topology map is constructed based on the computing unit communication matrix to obtain a memory bandwidth allocation map, which divides the total bandwidth of 16PB / s into segments.
[0122] Mapping the in-core storage units according to the memory bandwidth allocation map to obtain the in-core computing channel, which includes the read and write ports of the 192KB cache;
[0123] Data flow scheduling is performed based on the intra-core computing channel to obtain the inter-core data channel, which connects different core groups through a 3.5PB / s bandwidth;
[0124] Perform parallel analysis on the data channels between cores to obtain a parallel communication topology diagram, which describes the data exchange path between multiple cores.
[0125] Based on the parallel communication topology, a chip-level interconnection network is constructed to obtain the chip-to-chip interaction channel, which supports 5TB / s data transmission.
[0126] The data flow of the intra-core computing channel, the inter-core data channel and the inter-chip interaction channel is integrated to obtain an initial data flow graph, and data flow optimization calculation is performed based on the initial data flow graph to obtain a parallel computing data flow graph.
[0127] Specifically, in the dual-mode processing unit configuration matrix Based on this, the data flow is analyzed to obtain the computing unit communication matrix , records the data interaction relationship between the CPU core group and the NPU core group. Computing unit communication matrix By modeling the computing task flow of each core, the mathematical expression is:
[0128] ;
[0129] in, and They represent the data transmission requirements between the CPU core and the NPU core, and between the NPU core and the CPU core. This matrix reflects how the CPU and NPU cores can efficiently exchange data when processing multiple computing tasks. By analyzing the dependencies of data flows, the interaction mode of tasks between the cores and the flow path of data are determined. According to the computing unit communication matrix Constructing a memory access topology graph , which is used to show how different cores communicate through memory. The memory access topology is established by identifying the bandwidth requirements and communication paths of each computing unit. For example, for each pair of cores that need to communicate, their memory access requirements It is expressed by the following formula:
[0130] ;
[0131] in, Indicates the bandwidth limit of memory, It represents the bandwidth required for memory access. In order to ensure efficient data transmission memory bandwidth allocation, the computing requirements and bandwidth limitations of each core are considered. By segmenting the total bandwidth of 16PB / s, a memory bandwidth allocation diagram is constructed. , the total bandwidth is allocated to each computing unit and communication path according to demand. The calculation of the memory bandwidth allocation graph is expressed by the following formula:
[0132] ;
[0133] This formula analyzes the memory access pattern of the computing task to obtain the bandwidth required for each communication path. The memory bandwidth allocation diagram helps understand the bandwidth requirements of each computing unit and make reasonable bandwidth resource allocation. According to the memory bandwidth allocation diagram, the storage units in the core are mapped to obtain the computing channels in the core. , which is used to connect the computing unit inside the core with the memory. Each computing channel contains a 192KB cache to store data and provide a fast access path during read and write operations. The mapping of the computing channel in the core is expressed as:
[0134] ;
[0135] in, is the total capacity of the memory, and 192KB is the fixed size of each cache unit. By mapping the computing channels within the core, it is ensured that each computing unit has sufficient cache space when accessing the memory, so as to reduce memory access latency and improve computing efficiency. Data flow scheduling is performed based on the computing channels within the core to obtain the data channels between cores. The inter-core data channel connects the data transmission paths between different core groups, and the bandwidth of each data channel is 3.5PB / s. Through data flow scheduling, it is determined which tasks need to communicate across core groups, and bandwidth resources are dynamically allocated according to communication requirements. The bandwidth calculation of the inter-core data channel is expressed by the following formula:
[0136] ;
[0137] in, is the bandwidth of the inter-core data channel, The bandwidth allocation is derived from the memory bandwidth allocation diagram. In this way, the data transmission between different core groups is optimized and the maximum utilization of bandwidth resources is ensured. The parallelism analysis of the data channel between cores is performed to obtain the parallel communication topology diagram The parallel communication topology diagram describes the data exchange path between multiple cores, indicating which communications between cores can be executed in parallel and which communications need to be processed serially. The purpose of parallelism analysis is to improve the parallelism of data transmission, thereby reducing the waiting time for data transmission and improving the overall computing efficiency of the system. The parallel communication topology diagram is determined by calculating the dependencies and bandwidth requirements of data transmission. For example, if the communications between two computing units do not conflict, they transmit data simultaneously through parallel data channels to reduce the total execution time. In the parallel communication topology diagram After completion, a chip-level interconnect network is constructed based on the graph. , which is used to connect the communication channels between chips. The chip-to-chip interaction channel supports a data transmission bandwidth of 5TB / s, ensuring high-speed data exchange between different computing units. The construction of the chip-level interconnection network is achieved through a parallel communication topology graph, and the network topology analysis is used to determine the optimal path and bandwidth requirements for inter-chip communication. The bandwidth calculation of the chip-to-chip interaction channel is expressed by the following formula:
[0138] ;
[0139] in, is the bandwidth of the chip-to-chip interaction channel, and 5TB / s is the total bandwidth. This calculation ensures that data exchange between different chips can be completed quickly to improve computing performance. The data flows of the intra-core computing channel, the inter-core data channel, and the inter-chip interaction channel are integrated to obtain the initial data flow terminal. The bandwidth and latency of each computing unit and communication path are comprehensively considered to optimize the overall computing process. Initial data flow diagram Describe the data transmission and task scheduling methods of the entire system. Based on the initial data flow graph, perform data flow optimization calculations to obtain a parallel computing data flow graph Data flow optimization is achieved by adjusting the data flow path, bandwidth and scheduling strategy to improve the system's parallel computing capabilities and reduce power consumption. The goal of data flow optimization is to ensure the efficiency of data transmission and to evenly distribute computing tasks among computing units to avoid resource waste. The resulting parallel computing data flow diagram Represents the optimized data flow path, ensuring that the system has efficient data flow management and resource scheduling capabilities when processing large-scale computing tasks.
[0140] In a specific embodiment, the process of executing step 105 may specifically include the following steps:
[0141] Perform computational load distribution analysis on the parallel computing data flow graph to obtain a core load matrix, which contains real-time computational load data of the CPU core group and the NPU core group.
[0142] A voltage-frequency correlation calculation is performed based on the core load matrix to obtain a voltage-frequency mapping table, where the voltage-frequency mapping table includes a voltage range of 0.8V to 1.2V and a frequency range of 1GHz to 3GHz;
[0143] Perform power consumption characteristic analysis according to the voltage-frequency mapping table to obtain initial power consumption control parameters, which include a core temperature threshold and a power consumption threshold;
[0144] A dynamic regulation model is constructed based on the initial power consumption control parameters to obtain a voltage-frequency regulation matrix, which records the optimal working state under different load conditions.
[0145] Perform dynamic range division on the voltage-frequency regulation matrix to obtain a voltage regulation subspace and a frequency regulation subspace;
[0146] The core voltage regulation coefficient is calculated according to the voltage regulation subspace to obtain a voltage control vector, and the voltage control vector corresponds to the voltage regulation strategy of different cores;
[0147] The operating frequency adjustment coefficient is calculated based on the frequency adjustment subspace to obtain a frequency control vector, which corresponds to the frequency adjustment strategy of different cores. The voltage control vector and the frequency control vector are jointly optimized to obtain a power consumption optimization control vector.
[0148] Specifically, for parallel computing data flow graph Perform computational load distribution analysis to obtain the core load matrix The core load matrix records the real-time load data of the CPU core group and the NPU core group when performing computing tasks, reflecting the computing requirements of each core. The definition of the core load matrix is expressed as:
[0149] ;
[0150] in, and Represents the load data of the CPU core group and the NPU core group respectively. Each element and Describes the CPU core group and NPU core group The core and By analyzing these load data, we can understand the performance requirements of each computing unit when processing specific tasks. , calculate the voltage-frequency correlation and obtain the voltage-frequency mapping table The voltage-frequency mapping table maps the load of the computing core to the voltage and frequency adjustment range. Usually, the voltage range is set to 0.8V to 1.2V and the frequency range is 1GHz to 3GHz. Through voltage-frequency mapping, the core voltage and frequency are adjusted to minimize power consumption while ensuring processing performance. The formula for generating the voltage-frequency mapping table is as follows:
[0151] ;
[0152] in, is a function that calculates voltage and frequency based on the core load matrix, It is the input core load data. The voltage-frequency mapping table can provide the optimal voltage and frequency combination for each core to achieve the purpose of power consumption optimization. According to the voltage-frequency mapping table, the power consumption characteristics are analyzed to obtain the initial power consumption control parameters. The initial power consumption control parameters include the core temperature threshold and power consumption threshold. These two parameters are key to ensure that the processor can run stably under different load conditions. The calculation of power consumption control parameters is expressed by the following formula:
[0153] ;
[0154] in, is the core temperature threshold, indicating that the core will experience thermal runaway when this temperature is exceeded; is the power consumption threshold, which means that the core will overheat or reduce efficiency when the power consumption exceeds this value. , build a dynamic regulation model , and get the voltage-frequency regulation matrix The matrix records the optimal voltage and frequency configuration under different load conditions, aiming to achieve dynamic adjustment of voltage and frequency to ensure processor stability and optimal power consumption. Voltage and frequency adjustment matrix Generated by the following formula:
[0155] ;
[0156] in, Represents a calculation function for voltage and frequency adjustment based on the initial power control parameters. This function uses temperature and power thresholds to adjust the voltage and frequency of each core to adapt to different computing load conditions. Divide the dynamic range to obtain the voltage regulation subspace and frequency-adjusted subspace The purpose of dynamic range partitioning is to subdivide the voltage and frequency adjustment range so that the voltage and frequency of the core can be adjusted more accurately according to different load conditions. The divided voltage adjustment subspace and frequency adjustment subspace are expressed by the following formula:
[0157] ;
[0158] ;
[0159] in, and are the minimum and maximum values of voltage regulation, and are the minimum and maximum values of frequency regulation, respectively. These subspaces provide finer-grained voltage and frequency adjustment strategies to better adapt to changes in computing tasks. Calculate the core voltage adjustment coefficient based on the voltage adjustment subspace , and the voltage control vector . Voltage Regulation Factor is a coefficient that determines the voltage regulation strategy for each core under a given computational load. The voltage control vector It is expressed as:
[0160] ;
[0161] in, Based on core load The calculated voltage regulation coefficient reflects how to reduce power consumption by adjusting the voltage under different load conditions and ensure the stable operation of each core. Similarly, the operating frequency regulation coefficient is calculated based on the frequency regulation subspace. , and obtain the frequency control vector Frequency adjustment factor The purpose of the calculation is to determine how to optimize power consumption by adjusting the frequency under different load conditions. The calculation formula is:
[0162] ;
[0163] in, It is the frequency adjustment coefficient calculated based on the core load data. The adjustment of the frequency adjustment coefficient can significantly affect the performance and power consumption of the system. Through reasonable frequency adjustment, power consumption can be reduced without sacrificing performance. and frequency control vector Perform joint optimization operations to obtain the power consumption optimization control vector The goal of the joint optimization calculation is to comprehensively consider the changes in voltage and frequency to minimize power consumption and ensure stable operation of the system. It is expressed as:
[0164] ;
[0165] in, It is a joint optimization function that achieves the optimal power consumption control strategy by optimizing voltage and frequency at the same time. Provides precise power consumption regulation solutions for low-power AI processors to ensure that the processor can achieve optimal performance with minimal power consumption when performing computing tasks.
[0166] In a specific embodiment, the process of executing step 106 may specifically include the following steps:
[0167] Asynchronous event encoding is performed on the power consumption optimization control vector to obtain a pulse event sequence, where the pulse event sequence includes the discharge timing data of the neuron;
[0168] Based on the pulse event sequence, a spiking neural network structure is constructed to obtain a network topology matrix, which connects multiple neuron layers through synapses;
[0169] According to the network topology matrix, the time-series dependency is processed to obtain a neuron dynamic response model, which describes the response characteristics of neurons to input stimuli;
[0170] Perform network training operations based on the neuron dynamic response model to obtain a trained spiking neural network model, which processes asynchronous computing tasks based on an event-driven mechanism;
[0171] The trained spiking neural network model is used to perform load analysis to obtain a core load balancing vector, which records the load distribution of the CPU core group and the NPU core group.
[0172] A dynamic task migration strategy is constructed based on the core load balancing vector to obtain a task migration matrix, which contains the task scheduling scheme across core groups.
[0173] Based on the task migration matrix, the dynamic load balancing calculation is performed to obtain the initial balancing parameters, which include the computing resource allocation ratio. The initial balancing parameters are optimized to generate the computing power load balancing parameters.
[0174] Specifically, by optimizing the power consumption control vector Asynchronous event encoding is performed to generate a pulse event sequence. The purpose of asynchronous event encoding is to convert the power control vector into a format that can be processed by the pulse neural network, ensuring that the processor's computing tasks can accurately reflect the changes in voltage and frequency. The pulse event sequence contains the discharge timing data of the neuron, describing the activation state of each neuron at different time points. In this way, an accurate event-driven pulse sequence is generated based on the input computing task and power consumption requirements. Pulse event sequence It is expressed as:
[0175] ;
[0176] in, Indicates the time when the pulse occurs, Indicates the activation state of the neuron at that moment (discharging state or dormant state). This pulse coding makes a close connection between the activation of neurons and the computational task. Based on the pulse event sequence, the pulse neural network structure is constructed. , the network consists of multiple neuron layers connected by synapses. Network topology matrix It is the core structure of the spiking neural network, which describes the connection relationship between neurons and transmits information through the weighted connection of synapses. The network topology matrix is defined as follows:
[0177] ;
[0178] in, Representing neurons To the neuron The synaptic weights between is the number of neurons, is the number of neuron layers. This matrix records the connection mode and weights between different neuron layers, and plays a key role in the information transmission process. Through the weighted connection of synapses, the spiking neural network can handle asynchronous computing tasks and respond to different input stimuli. After constructing the spiking neural network structure, the network topology matrix is used to process the timing dependency, capture the time correlation between neurons, and the dynamic response characteristics of neurons to different input stimuli, and obtain the neuron dynamic response model. , which describes the temporal response behavior of neurons to input signals. The neuron dynamic response model is expressed as:
[0179] ;
[0180] in, is a function that represents the response of a neuron to an input signal based on its topological structure and sequence of pulse events. During the network training process, this model can predict how a neuron adjusts its state based on the input pulse. The trained spiking neural network model is obtained by performing network training operations based on the neuron dynamic response model. The training process is to adjust the weights of synapses in the network through back propagation algorithm or other optimization algorithms, so that the pulse neural network can better handle asynchronous computing tasks. The trained pulse neural network model is expressed as:
[0181] ;
[0182] When executing computing tasks, the model controls the scheduling and load distribution of computing tasks according to the event-driven mechanism. Through the event-driven mechanism, the spiking neural network can respond to changes in computing tasks in real time, optimize load distribution, and reduce power consumption. The core load balancing vector is obtained by load analysis through the trained spiking neural network model. , which records the load distribution of the CPU core group and the NPU core group. The calculation formula of the core load balancing vector is:
[0183] ;
[0184] in, and Represents the load distribution of the CPU core group and the NPU core group respectively. Each element and Describes the load level of each core group. Through this vector, the current system load is determined, and the allocation of computing resources is optimized. , build a dynamic task migration strategy , which determines how to schedule computing tasks between different core groups. Task migration matrix It is built based on the core load balancing vector and includes a task scheduling scheme across core groups. The purpose is to schedule more tasks on the core groups with lighter loads, thereby balancing the load of the entire computing system. It is expressed as:
[0185] ;
[0186] in, Indicates that the task From the core Move to the core The task migration matrix provides a specific scheduling solution for load balancing, which enables computing resources to be allocated more efficiently. , perform load dynamic balancing calculation and obtain the initial balancing parameters , which contains the allocation ratio of computing resources. The goal of load dynamic balancing calculation is to ensure a more balanced load distribution between the CPU core group and the NPU core group to optimize the overall computing efficiency. Initial balancing parameters The representation is:
[0187] ;
[0188] in, is the total system load, and The loads of the CPU core group and the NPU core group respectively. This parameter reflects the proportion of each core group in the total load. Perform optimization calculations to obtain computing power load balancing parameters By adjusting the resource allocation ratio, the total power consumption of the system is minimized while ensuring that the computing performance is not affected. The optimal calculation is achieved using mathematical programming algorithms, particle swarm optimization and other methods. Computing power load balancing parameters The calculation formula is:
[0189] ;
[0190] in, It is an optimization function that calculates the final load balancing strategy based on the initial balancing parameters. Through this strategy, the CPU and NPU cores work together more efficiently to achieve a balanced computing load.
[0191] The above describes the real-time data processing method of the low-power AI processor in the embodiment of the present invention. The following describes the real-time data processing chip of the low-power AI processor in the embodiment of the present invention. Figure 2 , an embodiment of a real-time data processing chip of a low-power AI processor in an embodiment of the present invention includes:
[0192] A classification module 201 is used to perform neuromorphic classification processing on the computing tasks input by the low-power AI processor to obtain a computing task feature matrix;
[0193] The operation module 202 is used to perform multi-dimensional mapping operation on the computing task feature matrix to obtain a heterogeneous computing space vector;
[0194] An allocation module 203 is used to allocate core resources based on the heterogeneous computing space vector to obtain a dual-mode processing unit configuration matrix;
[0195] A construction module 204 is used to construct a storage-computation integrated data path based on the dual-mode processing unit configuration matrix to obtain a parallel computing data flow graph;
[0196] A mapping module 205, configured to perform dynamic voltage-frequency mapping on the parallel computing data flow graph to obtain a power consumption optimization control vector;
[0197] The generation module 206 is used to input the power consumption optimization control vector into the trained pulse neural network model to perform computing load balancing calculation and generate computing load balancing parameters.
[0198] Through the collaborative cooperation of the above components, neuromorphic classification processing and multi-dimensional mapping operations, the feature extraction of computing tasks is more accurate, the task classification is more reasonable, and the task processing efficiency of the system is improved; the core resource allocation mechanism based on heterogeneous computing space vectors realizes the intelligent resource scheduling of CPU core group and NPU core group, giving full play to the advantages of dual-mode processing architecture; the storage and computing integrated data path design reduces data handling overhead, optimizes data flow transmission efficiency, and significantly reduces system power consumption; through dynamic voltage-frequency mapping technology, the precise regulation of processor working state is achieved; the pulse neural network model is introduced for computing load balancing calculation, and event-driven asynchronous computing scheduling is realized, which improves the computing efficiency of the system; the multi-level cache and high-bandwidth design are adopted to achieve 16PB / s memory bandwidth, 3.5PB / s core-to-core communication bandwidth and 5TB / s chip-to-chip communication bandwidth, which greatly improves data processing capabilities; the neural network structure constructed by neurons and synapses provides powerful parallel computing capabilities and realizes efficient task processing; dynamic migration and load balancing of tasks are realized, avoiding excessive concentration or idleness of computing resources, and improving the overall utilization of the system.
[0199] above Figure 2 The real-time data processing chip of the low to medium power AI processor in the embodiment of the present invention is described in detail from the perspective of modular functional entities. The real-time data processing device of the low power AI processor in the embodiment of the present invention is described in detail from the perspective of hardware processing.
[0200] Figure 3 It is a structural schematic diagram of a real-time data processing device of a low-power AI processor provided by an embodiment of the present invention. The real-time data processing device 300 of the low-power AI processor may have relatively large differences due to different configurations or performances, and may include one or more processors (central processing units, CPU) 310 (for example, one or more processors) and a memory 320, and one or more storage media 330 (for example, one or more mass storage device terminals) storing application programs 333 or data 332. Among them, the memory 320 and the storage medium 330 can be short-term storage or permanent storage. The program stored in the storage medium 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations in the real-time data processing device 300 of the low-power AI processor. Furthermore, the processor 310 can be configured to communicate with the storage medium 330, and execute a series of instruction operations in the storage medium 330 on the real-time data processing device 300 of the low-power AI processor to implement the steps of the real-time data processing method of the low-power AI processor described above.
[0201] The real-time data processing device 300 of the low-power AI processor may also include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input and output interfaces 360, and / or one or more operating systems 331, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, etc. It can be understood by those skilled in the art that Figure 3 The real-time data processing device structure of the low-power AI processor shown does not constitute a limitation on the real-time data processing device of the low-power AI processor provided by the present invention, and may include more or fewer components than shown in the figure, or a combination of certain components, or a different arrangement of components.
[0202] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions are executed on a computer, the computer executes the steps of the real-time data processing method of the low-power AI processor.
[0203] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0204] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the whole or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program code.
[0205] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A real-time data processing method for a low-power AI processor, characterized in that: The method comprises: Perform neuromorphic classification processing on the computing tasks input by the low-power AI processor to obtain the computing task feature matrix; Performing a multi-dimensional mapping operation on the computing task feature matrix to obtain a heterogeneous computing space vector; Core resource allocation is performed based on the heterogeneous computing space vector to obtain a dual-mode processing unit configuration matrix; specifically comprising: hierarchical resource partitioning of the heterogeneous computing space vector to obtain a CPU task set and an NPU task set; utilization analysis of the neuromorphic core is performed according to the CPU task set to obtain a CPU core group allocation matrix; computing power analysis of the neuromorphic core is performed according to the NPU task set to obtain an NPU core group allocation matrix; cache capacity allocation is performed based on the CPU core group allocation matrix and the NPU core group allocation matrix to obtain a cache allocation vector; 192KB unit partitioning is performed on the cache allocation vector to obtain a cache unit group, and memory bandwidth allocation is performed according to the cache unit group to obtain a bandwidth allocation matrix; processing unit combination operation is performed based on the bandwidth allocation matrix to obtain an initial processing unit configuration matrix, and resource balancing optimization is performed on the initial processing unit configuration matrix to obtain a dual-mode processing unit configuration matrix; Based on the dual-mode processing unit configuration matrix, a storage-computation integrated data path is constructed to obtain a parallel computing data flow graph; Performing dynamic voltage-frequency mapping on the parallel computing data flow graph to obtain a power consumption optimization control vector; The power consumption optimization control vector is input into the trained pulse neural network model to perform computing load balancing calculation and generate computing load balancing parameters.
2. The real-time data processing method of the low-power AI processor according to claim 1, characterized in that: The computing tasks input by the low-power AI processor are subjected to neuromorphic classification processing to obtain a computing task feature matrix, including: Performing scale analysis and calculation on the computing task input by the low-power AI processor to obtain a task complexity vector, wherein the task complexity vector includes a data scale parameter and a computing complexity parameter; Inputting the task complexity vector into the neuromorphic classifier, and performing parallel classification operations on the computing tasks through the neuromorphic core to obtain a neuron state distribution matrix; Performing pulse coding operation according to the neuron state distribution matrix, encoding information through neuron nodes to obtain a neuron activation vector, performing synaptic connection calculation based on the neuron activation vector, and updating weights through synapses to obtain a synaptic weight vector; Performing data flow dependency analysis on the synaptic weight vector, obtaining a task dependency graph by establishing a mapping relationship between a preceding task and a subsequent task, and inputting the task dependency graph into a topological sorting processor for data caching to obtain an optimized task dependency vector; Constructing a three-dimensional feature space based on the neuron activation vector, the synaptic weight vector and the optimized task dependency vector, and performing feature fusion through an asynchronous event-driven mechanism to obtain an initial task feature matrix; A batch normalization operation is performed according to the initial task feature matrix, and a calculation task feature matrix is obtained through dynamic scaling and translation operations.
3. The real-time data processing method of the low-power AI processor according to claim 2, characterized in that: The performing a multi-dimensional mapping operation on the computing task feature matrix to obtain a heterogeneous computing space vector includes: Performing a feature decomposition operation on the computing task feature matrix to obtain a computing feature principal component matrix, and constructing a task similarity measurement space based on the computing feature principal component matrix to obtain a task similarity matrix; Performing core affinity calculation according to the task similarity matrix to obtain a CPU core affinity vector and an NPU core affinity vector, and performing computational load distribution analysis based on the CPU core affinity vector and the NPU core affinity vector to obtain a load distribution mapping matrix; Performing a dimension compression operation on the load distribution mapping matrix to obtain a calculation dimension projection vector, performing task relevance calculation based on the calculation dimension projection vector, and obtaining a task relevance vector through a cosine similarity function; A heterogeneous computing resource graph is constructed based on the task relevance vector to obtain a computing resource allocation vector, and a spatial coordinate transformation is performed on the computing resource allocation vector to obtain a heterogeneous computing space vector.
4. The real-time data processing method of the low-power AI processor according to claim 1, characterized in that: The step of constructing a storage-computation integrated data path based on the dual-mode processing unit configuration matrix to obtain a parallel computing data flow graph includes: Performing data flow analysis on the dual-mode processing unit configuration matrix to obtain a computing unit communication matrix, wherein the computing unit communication matrix records the data interaction relationship between the CPU core group and the NPU core group; Building a memory access topology map based on the computing unit communication matrix to obtain a memory bandwidth allocation map, wherein the memory bandwidth allocation map divides the total bandwidth of 16PB / s into segments; Mapping the in-core storage unit according to the memory bandwidth allocation map to obtain an in-core computing channel, wherein the in-core computing channel includes a read and write port of a 192KB cache; Perform data flow scheduling based on the intra-core computing channel to obtain an inter-core data channel, wherein the inter-core data channel connects different core groups via a 3.5PB / s bandwidth; Performing parallelism analysis on the inter-core data channel to obtain a parallel communication topology diagram, wherein the parallel communication topology diagram describes a data exchange path between multiple cores; Building a chip-level interconnection network based on the parallel communication topology graph to obtain an inter-chip interaction channel, wherein the inter-chip interaction channel supports 5TB / s data transmission; The data flow of the intra-core computing channel, the inter-core data channel and the inter-chip interaction channel is integrated to obtain an initial data flow graph, and data flow optimization calculation is performed based on the initial data flow graph to obtain a parallel computing data flow graph.
5. The real-time data processing method of the low-power AI processor according to claim 4 is characterized in that: The performing dynamic voltage-frequency mapping on the parallel computing data flow graph to obtain a power consumption optimization control vector includes: Performing a computational load distribution analysis on the parallel computing data flow graph to obtain a core load matrix, wherein the core load matrix includes real-time computational load data of a CPU core group and an NPU core group; Perform voltage-frequency correlation calculation based on the core load matrix to obtain a voltage-frequency mapping table, wherein the voltage-frequency mapping table includes a voltage range of 0.8V to 1.2V and a frequency range of 1GHz to 3GHz; Perform power consumption characteristic analysis according to the voltage-frequency mapping table to obtain initial power consumption control parameters, wherein the initial power consumption control parameters include a core temperature threshold and a power consumption threshold; A dynamic adjustment model is constructed based on the initial power consumption control parameters to obtain a voltage-frequency adjustment matrix, wherein the voltage-frequency adjustment matrix records the optimal working state under different load conditions; Performing dynamic range division on the voltage-frequency regulation matrix to obtain a voltage regulation subspace and a frequency regulation subspace; Calculating a core voltage regulation coefficient according to the voltage regulation subspace to obtain a voltage control vector, wherein the voltage control vector corresponds to a voltage regulation strategy for different cores; The operating frequency adjustment coefficient is calculated based on the frequency adjustment subspace to obtain a frequency control vector, which corresponds to the frequency adjustment strategy of different cores, and the voltage control vector and the frequency control vector are jointly optimized to obtain a power consumption optimization control vector.
6. The real-time data processing method of the low-power AI processor according to claim 5, characterized in that: The step of inputting the power consumption optimization control vector into the trained pulse neural network model to perform computing load balancing calculation and generate computing load balancing parameters includes: Performing asynchronous event coding on the power consumption optimization control vector to obtain a pulse event sequence, wherein the pulse event sequence includes discharge timing data of neurons; Building a spiking neural network structure based on the spiking event sequence to obtain a network topology matrix, wherein the network topology matrix connects multiple neuron layers through synapses; Performing time-dependent processing according to the network topology matrix to obtain a neuron dynamic response model, wherein the neuron dynamic response model describes the response characteristics of the neuron to the input stimulus; Performing network training operations based on the neuron dynamic response model to obtain a trained spiking neural network model, wherein the trained spiking neural network model processes asynchronous computing tasks based on an event-driven mechanism; Performing load analysis through the trained pulse neural network model to obtain a core load balancing vector, wherein the core load balancing vector records the load distribution of the CPU core group and the NPU core group; Constructing a dynamic task migration strategy according to the core load balancing vector to obtain a task migration matrix, wherein the task migration matrix includes a task scheduling scheme across core groups; Based on the task migration matrix, a load dynamic balancing calculation is performed to obtain initial balancing parameters, wherein the initial balancing parameters include a computing resource allocation ratio, and the initial balancing parameters are optimized to generate a computing power load balancing parameter.
7. A real-time data processing chip for a low-power AI processor, characterized in that: A real-time data processing method for implementing a low-power AI processor according to any one of claims 1 to 6, wherein the real-time data processing chip of the low-power AI processor comprises: The classification module is used to perform neuromorphic classification processing on the computing tasks input by the low-power AI processor to obtain the computing task feature matrix; A computing module, used for performing a multi-dimensional mapping operation on the computing task feature matrix to obtain a heterogeneous computing space vector; An allocation module, configured to allocate core resources based on the heterogeneous computing space vector to obtain a dual-mode processing unit configuration matrix; A construction module, used to construct a storage-computation integrated data path based on the dual-mode processing unit configuration matrix to obtain a parallel computing data flow graph; A mapping module, used for performing dynamic voltage-frequency mapping on the parallel computing data flow graph to obtain a power consumption optimization control vector; A generation module is used to input the power consumption optimization control vector into the trained pulse neural network model to perform computing load balancing calculation and generate computing load balancing parameters.
8. A real-time data processing device for a low-power AI processor, characterized in that: The real-time data processing device of the low-power AI processor includes: a memory and at least one processor, wherein the memory stores instructions; The at least one processor calls the instruction in the memory so that the real-time data processing device of the low-power AI processor executes the real-time data processing method of the low-power AI processor as described in any one of claims 1-6.
9. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instruction is executed by the processor, the real-time data processing method of the low-power AI processor as described in any one of claims 1-6 is implemented.
Citation Information
Patent Citations
On-chip core compiling and mapping method and device of neural network based on reinforcement learning
CN114492782A
Load balancing method of low-power AI processor, chip and storage medium
CN119271418A