Parallel computing method, chip and storage medium based on low-power AI processor
By building a neuron mapping matrix and storage allocation matrix on a low-power AI processor, combining pulse neural network and integrated storage and computing architecture, the problems of high energy consumption and low efficiency of traditional parallel computing methods are solved, and efficient and robust parallel computing effects are achieved.
Patent Information
- Application Number
- CN202510249062.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-03-04
AI Technical Summary
Traditional parallel computing methods face the problems of high energy consumption and low efficiency when dealing with large-scale parallel computing. Especially in edge computing scenarios, computing resources are limited and the requirements for real-time performance and energy efficiency ratio are strict.
By adopting a parallel computing method based on low-power AI processors, by building a neuron mapping matrix and storage allocation matrix, the refined decomposition and efficient scheduling of computing tasks are achieved, combining pulse neural networks and integrated storage architecture, data processing and storage are optimized, exception monitoring matrix and distributed fault tolerance mechanism are introduced, and the system reliability and adaptability are improved.
It effectively reduces system energy consumption, improves computing efficiency and robustness, realizes efficient utilization of computing resources, and improves the overall performance and reliability of the system.
Smart Images

Figure CN119739488B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of low-power AI processors, and in particular to a parallel computing method, chip and storage medium based on a low-power AI processor. Background Art
[0002] With the rapid development of artificial intelligence technology, the scale and complexity of computing tasks are constantly increasing. Traditional computing architectures face the challenges of high energy consumption and low efficiency when dealing with large-scale parallel computing. Especially in edge computing scenarios, computing resources are limited, while the requirements for real-time performance and energy efficiency are very strict, which makes how to achieve efficient computing under limited resource conditions an urgent problem to be solved.
[0003] Most existing parallel computing methods use a centralized architecture, and often ignore the data dependencies between tasks during task scheduling and resource allocation, resulting in frequent data migration and communication overhead, and difficulty in dealing with failures and abnormal situations during the computing process. Traditional methods lack a flexible dynamic scheduling mechanism when dealing with heterogeneous computing tasks, and cannot fully utilize hardware resources, resulting in low computing efficiency and energy waste. Summary of the invention
[0004] The main purpose of the present invention is to provide a parallel computing method, chip and storage medium based on a low-power AI processor, so as to achieve efficient utilization of computing resources and improve the computing efficiency and robustness of the low-power AI processor as a whole.
[0005] To achieve the above object, the present invention provides a parallel computing method based on a low-power AI processor, comprising the following steps:
[0006] A neuron mapping matrix is constructed according to the computing tasks input by the low-power AI processor, and the neuron mapping matrix is divided into a plurality of subtask groups according to the data communication bandwidth threshold of the processing core to obtain a parallel processing unit allocation scheme;
[0007] Constructing a storage allocation matrix according to the parallel processing unit allocation scheme, performing cache resource allocation and execution sequence arrangement on the subtask groups through pulse neural network operation, and obtaining a computing core mapping scheme;
[0008] An anomaly monitoring matrix is constructed based on the computing core mapping scheme, the anomaly monitoring matrix is deployed to the local storage unit of the neuromorphic core, and a cross-core data recovery channel is established to obtain a distributed fault-tolerant solution;
[0009] Perform parallel matrix operations on the subtask group according to the distributed fault-tolerant scheme, distribute data blocks to corresponding neuron processing units according to the storage-computation integrated architecture, and establish an asynchronous data transmission channel to obtain a parallel computing result set;
[0010] Performing time domain feature analysis on the parallel computing result set, extracting data features by establishing a cross-core pulse signal communication network, and obtaining a global feature vector;
[0011] A dynamic load matrix of S×T is generated according to the global eigenvector, and computing tasks are redistributed to processing cores based on neuron state distribution to obtain a parallel task execution plan.
[0012] The present invention also provides a parallel computing chip based on a low-power AI processor, comprising:
[0013] A partitioning module is used to construct a neuron mapping matrix according to the computing tasks input by the low-power AI processor, and divide the neuron mapping matrix into multiple subtask groups according to the data communication bandwidth threshold of the processing core to obtain a parallel processing unit allocation plan;
[0014] An orchestration module, configured to construct a storage allocation matrix according to the parallel processing unit allocation scheme, and to perform cache resource allocation and execution sequence orchestration on the subtask groups through pulse neural network operations to obtain a computing core mapping scheme;
[0015] Establishing a module for constructing an anomaly monitoring matrix based on the computing core mapping scheme, deploying the anomaly monitoring matrix to a local storage unit of the neuromorphic core, and establishing a cross-core data recovery channel to obtain a distributed fault-tolerant solution;
[0016] A computing module, configured to perform parallel matrix operations on the subtask groups according to the distributed fault-tolerant scheme, distribute data blocks to corresponding neuron processing units according to the storage-computation integrated architecture, and establish an asynchronous data transmission channel to obtain a parallel computing result set;
[0017] An analysis module is used to perform time domain feature analysis on the parallel computing result set, extract data features by establishing a cross-core pulse signal communication network, and obtain a global feature vector;
[0018] The allocation module is used to generate an S×T dynamic load matrix according to the global feature vector, and to reallocate computing tasks to processing cores based on neuron state distribution to obtain a parallel task execution plan.
[0019] The present invention also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned methods are implemented.
[0020] In summary, the technical solution provided by the present invention realizes the refined decomposition and efficient scheduling of computing tasks by constructing a neuron mapping matrix and a storage allocation matrix, thereby effectively reducing the energy consumption of the system; adopts the design of a pulse neural network combined with a storage and computing integrated architecture to make data processing and storage more closely integrated, reducing data handling overhead; introduces an abnormal monitoring matrix and a distributed fault-tolerant mechanism to improve the reliability of the system; uses a dynamic load matrix for real-time task redistribution to enable the system to have adaptive adjustment capabilities; through the design of a multi-level parallel computing pipeline and an asynchronous data transmission channel, the parallel computing potential of the processor is fully utilized, and at the same time, based on time domain feature analysis and cross-core collaborative optimization, efficient utilization of computing resources is achieved, thereby improving the computing efficiency and robustness of the system as a whole. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is a schematic diagram of the steps of a parallel computing method based on a low-power AI processor in one embodiment of the present invention;
[0022] Figure 2 This is a block diagram of the structure of a parallel computing chip based on a low-power AI processor in one embodiment of the present invention.
[0023] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0024] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0025] Reference Figure 1 , this embodiment provides a parallel computing method based on a low-power AI processor, comprising the following steps:
[0026] S1, constructing a neuron mapping matrix according to the computing tasks input by the low-power AI processor, dividing the neuron mapping matrix into multiple subtask groups according to the data communication bandwidth threshold of the processing core, and obtaining a parallel processing unit allocation scheme;
[0027] Among them, the computing tasks input by the low-power AI processor are scanned for data dependencies and the task node directed graph is constructed to obtain the initial task dependency relationship. The preliminary task directed graph reflects the execution order and data transmission requirements between each computing task. According to the initial task dependency relationship, the data communication volume between task nodes is calculated, and the task data flow graph is generated, so that the data transmission requirements between tasks are quantified, and the data source and required bandwidth on which each task depends can be identified. The task nodes in the task data flow graph are divided into neuron units. Considering the hardware characteristics of the low-power AI processor, the cache capacity of each neuromorphic core is 192KB. Data is divided according to this capacity to ensure that each neuron core can effectively store and process task data within its cache capacity. After the neuron units are divided, a neuron mapping matrix is formed. The communication cost is calculated based on the neuron mapping matrix. By comparing the data communication volume between each task node with the data communication bandwidth threshold of the processing core, the data transmission cost table between the cores is obtained. Task clustering operations are performed according to the inter-core data transmission cost table, and task nodes with similar communication costs are assigned to the same neuromorphic core. Through this process, the data transmission burden between cores can be effectively reduced, the communication cost can be reduced, and the dependencies between tasks can be satisfied to obtain the initial subtask group. Since the computing requirements and storage requirements of different task groups are different, the initial subtask group is subjected to load balancing analysis. In the load balancing analysis stage, the load distribution diagram of the task is obtained by calculating the computing density and storage requirements of each subtask group, which shows the computing load and storage requirements of each task group, and can help identify which subtask groups have large computing volume or high storage requirements. Based on the load distribution diagram, the boundary tasks are adjusted. When it is found that the computing load of some subtask groups exceeds the preset threshold, the overloaded task group is split and reorganized, and the excessive computing load is distributed to multiple subtask groups to ensure that the computing volume and storage requirements of each subtask group are within a reasonable range. The splitting and reorganization process can not only balance the load between tasks, but also further optimize the allocation of computing resources. By allocating resources to the multiple subtask groups after splitting, the 128 processing cores of the low-power AI processor are constructed into multiple processing unit groups according to the degree of data dependency between tasks. Each processing unit group is allocated according to the specific needs of the task, ensuring that the resources of each neuromorphic core are optimally utilized, and ultimately forming a parallel processing unit allocation plan.
[0028] S2, builds a storage allocation matrix according to the parallel processing unit allocation plan, allocates cache resources and arranges execution sequences for subtask groups through pulse neural network operations, and obtains a computing core mapping plan;
[0029] Specifically, the parallel processing unit allocation scheme is scanned for storage resources, the ratio between the amount of data of each processing unit group and the 192KB cache capacity of each neuromorphic core is evaluated, and a storage requirement matrix is generated to reflect the storage space required for each subtask group. The tasks are divided into data blocks according to the storage requirement matrix, and the size of each data block is calculated according to the actual situation to obtain a storage allocation matrix. The accuracy of data block segmentation and size calculation directly affects the effective use of storage resources. The storage allocation matrix is input into the pulse neural network for data flow analysis. By setting appropriate synaptic weights, the data transmission path is quantitatively calculated. Synaptic weights can reflect the cost and benefit of data transmission, thereby helping the system to efficiently transfer data between different processing cores. After quantitative calculation, a data path diagram is obtained, which shows the communication path and data flow between tasks, and optimizes the efficiency of data transmission. According to the data path diagram, the subtask groups are analyzed for timing dependency, the execution order between tasks is confirmed, and these orders are sorted to form an initial execution timing table to avoid calculation errors caused by data dependency failure. The initial execution timing table is analyzed for parallelism to evaluate the degree of parallel execution of tasks. The tasks to be executed simultaneously are determined by the overlapping parts between the computing tasks, and the computing resources are reasonably divided according to the synaptic computing power. The division of synaptic computing power takes into account the computing power of each core and the computing requirements of the tasks, ensuring that the computing tasks are properly distributed among multiple processing units to avoid the phenomenon of wasting or overloading computing resources. After parallelism analysis, the pulse neuron allocation scheme is obtained, which describes how to allocate computing tasks to different neuromorphic cores. When mapping this scheme to the neuromorphic core, the data communication link between the cores is constructed based on the storage and computing integrated architecture to ensure unimpeded data transmission between the cores. Based on the constructed core connection topology diagram, the tasks are pipelined. Neuron operations are allocated to the processing cores in the order of data dependency to form an orderly task execution queue. Each task in the task execution queue has a clear execution order, and the dependency relationship between tasks is guaranteed not to be destroyed. After completing the pipeline orchestration, the task execution queue is resource bound. A mapping relationship is established between the task and the corresponding neuromorphic core to ensure that each task can be allocated to the appropriate core and can use sufficient computing resources. Finally, the computing core mapping scheme is obtained.
[0030] S3, build an anomaly monitoring matrix based on the computing core mapping scheme, deploy the anomaly monitoring matrix to the local storage unit of the neuromorphic core, and establish a cross-core data recovery channel to obtain a distributed fault-tolerant solution;
[0031] It should be noted that the abnormal features of the computing core mapping scheme are extracted to obtain a preliminary abnormal pattern set, which reflects various abnormal situations that occur during the computing process, including hardware failures, communication delays, or task execution failures. By extracting these patterns, potential problems that affect the stability of the computing process are identified. Based on the preliminary abnormal pattern set, monitoring rules are constructed for the operating status of each neuromorphic core. These monitoring rules specifically describe how to detect and respond to various abnormal situations in real time during the computing process, so as to ensure that the system can respond quickly and take appropriate measures when facing sudden failures, forming an abnormal monitoring matrix. The abnormal monitoring matrix is mapped in blocks, and the monitoring rules are divided according to the storage capacity threshold (192KB) of each neuromorphic core to ensure that each core can effectively deploy monitoring rules in its local storage unit. The divided monitoring rule distribution table can guide the allocation of monitoring tasks for each neuromorphic core, so that each core only needs to focus on the monitoring tasks directly related to it. By means of block mapping, the utilization of storage resources is optimized to avoid storage overload or redundant monitoring data. According to the monitoring rule distribution table, the state of the computing unit in each neuromorphic core is monitored to build a core-level anomaly detection network. The network can obtain the operating status of each core in real time, detect anomalies in time and issue warnings to ensure the stability of the entire system. Based on the core-level anomaly detection network, the construction of cross-core communication links is performed. The role of the cross-core communication link is to ensure that the cores can share information and recover data in time when an anomaly occurs. In order to achieve efficient data backup and recovery, a data backup channel is established through the 3.5PB / s inter-core communication bandwidth. The high-speed communication bandwidth greatly improves the data transmission speed, ensuring that the lost information can be quickly restored from the backup data when the core fails. Through the data backup channel, a data transmission link group is obtained to connect the data backup and recovery functions between different cores. On this link group, the deployed fault-tolerant strategy can ensure that even if some cores fail, other cores can still maintain the normal operation of the computing task. The data redundant storage structure constructed by neurons enhances the fault tolerance of the system. Through the redundant storage mechanism, the failure of any neuromorphic core will not cause a complete interruption of the computing task, thereby obtaining a distributed backup solution. Based on the distributed backup solution, the construction of the abnormal recovery channel is performed. The role of the recovery channel is to map the backup data to the available computing cores through the storage-computing integrated architecture to ensure that the data can be quickly restored and the computing tasks can continue to be performed after a failure. Through this process, the fault switching mechanism is implemented to ensure that the system can automatically select the backup core for task redistribution after a core failure occurs, minimizing the downtime during the computing process. The coordinated control strategy configuration of the fault switching mechanism enables each core to achieve seamless switching when a failure occurs and maintain the continuity of computing.In order to improve the system's response speed and fault tolerance, an asynchronous event-driven mechanism is used to establish a cross-core data recovery process. The asynchronous event-driven mechanism ensures that the data recovery process will not block the operation of the entire system due to the failure of a single core, optimizes the system's fault tolerance efficiency, and ultimately obtains a complete distributed fault tolerance solution.
[0032] S4, perform parallel matrix operations on the subtask groups according to the distributed fault-tolerant scheme, distribute the data blocks to the corresponding neuron processing units according to the storage-computation integrated architecture, and establish an asynchronous data transmission channel to obtain a parallel computing result set;
[0033] Specifically, the block-level division of the data redundant storage scheme is performed according to the distributed fault-tolerant scheme. The computational data is effectively grouped according to the cache capacity of the neuromorphic core to ensure that the cache space of each neuron processing unit will not be overloaded by too much data. The cache capacity of each neuromorphic core is 192KB, and the computational data is divided into several data blocks to ensure that the size of each data block does not exceed the cache capacity. These data blocks divide the matrix operation tasks on demand to ensure that the computing tasks can adapt to the storage and computing power of the low-power AI processor, and obtain the data block allocation table. Based on the data block allocation table, a computational dependency graph is constructed to complete the mapping process of the matrix operation tasks. The computational dependency graph specifies how each computing node cooperates with other nodes to complete the matrix operation by clarifying the dependency relationship between tasks. The computational task mapping scheme is obtained by mapping the computational tasks to the processing network formed by neurons. This scheme reasonably allocates the matrix operation tasks to each neuromorphic core, thereby optimizing the efficiency of parallel computing. In this process, the mapping of matrix operation tasks not only depends on the storage capacity of each computing node, but also considers the mutual dependence between tasks, the size of the computational amount, and the load balance of the processing unit. The parallelism analysis of the computing task mapping scheme is carried out to maximize the parallel execution capability in the computing process, so that each task can be executed as simultaneously as possible without sacrificing the computing accuracy and task dependency. Through the parallelism analysis, it is possible to identify which parts of the computing task can be executed in parallel and which parts need to be processed serially. Through the storage and computing integrated architecture, data and computing instructions are combined into atomic operation units, and data loading, computing and storage operations are combined into a complete computing process, so that data does not need to be frequently transmitted between the computing and storage processes, improving computing efficiency. The generation and execution of atomic operation units can reduce communication delays and storage access bottlenecks, and form a more efficient operation instruction sequence. According to the generated operation instruction sequence, an asynchronous communication link is established to build a data transmission channel network, so as to ensure that the computing task can efficiently and unimpededly transmit data between each neuromorphic core. The establishment of asynchronous communication links effectively avoids the delay caused by each computing unit in the process of waiting for data transmission, so that each computing unit can start computing immediately after receiving the data without waiting for the calculation results of other cores. Through the design based on the data flow channel graph, the execution order of tasks is refined to form a data pipeline. Divide continuous computing tasks into multiple stages and process them in parallel at each stage to reduce the waiting time between tasks and improve the overall computing throughput. Through data pipeline design, the computing process is decomposed into multi-level parallel execution units, so that each execution unit can process multiple computing tasks at the same time. While optimizing computing performance, the parallel execution method can significantly reduce the execution time of tasks and improve the efficiency of the entire computing process. In order to ensure that the load of each processing core is reasonable and avoid overload or idleness, dynamic adjustment is performed according to the communication bandwidth of each core to optimize the data transmission rate.By dynamically adjusting the communication bandwidth between cores, the balanced distribution of computing tasks between different cores is achieved, so that the overall computing system can operate efficiently. Based on these adjustments, the resulting computing beat sequence is used to synchronize the execution of computing tasks. Each beat corresponds to a computing cycle, ensuring that all tasks are executed in a predetermined time sequence. In the matrix block operation stage, the final calculation results are aggregated through the 5TB / s inter-chip communication bandwidth to form distributed computing results. Matrix block operation divides the calculation process into multiple block-level tasks, each task is executed independently on different cores, and finally all the calculation results are aggregated to obtain the final calculation result set. In order to ensure the consistency of the calculation results, the consistency of the distributed calculation results is verified. The verification process is completed through a cross-core data verification mechanism to ensure that the output data of each processing unit is logically consistent and avoid inconsistency problems caused by communication delays or calculation errors. The output data of each processing unit is merged through a data merging mechanism to obtain the final result set of parallel computing.
[0034] S5, performing time domain feature analysis on the parallel computing result set, extracting data features by establishing a cross-core pulse signal communication network, and obtaining a global feature vector;
[0035] Among them, the parallel computing result set is time-series parsed. In order to fully capture the dynamic changes of data in the time domain, time domain sampling points are established according to the 4096 state configurations of each neuron to form a time feature sequence. The time feature sequence is composed of the state value of each neuron at a time point, reflecting the changes in the activation state of the neuron at different time scales. The data is converted into pulse signals according to the extracted time feature sequence. The pulse neural network maps the time series features into pulse sequences through a synaptic connection network. The pulse sequence can more effectively simulate the working mode of biological neural networks, and its discrete and event-driven characteristics are suitable for running in a low-power hardware environment. Through this pulse conversion process, the time feature sequence is converted into a characteristic pulse stream. When the pulse stream is transmitted in the network, it can represent and process data in a more efficient and low-energy way, which is conducive to the execution of large-scale parallel computing. Based on the characteristic pulse stream, a cross-core channel configuration is performed to establish a cross-core pulse signal communication network, and 128 neuromorphic cores are connected through synapses to build a signal transmission network. In this process, according to the transmission requirements of the pulse stream and the communication bandwidth between cores, a suitable communication topology is designed to ensure that data can be efficiently transmitted between each neuromorphic core, and a pulse transmission topology diagram is obtained. Bandwidth is allocated to the pulse transmission topology diagram to ensure that there will be no bottleneck or overload problems during signal transmission. The bandwidth allocation process dynamically adjusts the communication rate between each core and optimizes the data transmission path, so that the pulse stream can be transmitted with the lowest delay and the highest efficiency, and a feature transmission scheme is obtained. Based on the feature transmission scheme, data feature extraction is performed. Neurons are used to perform feature recognition operations on pulse signals. Neurons can extract key features in the signal, such as mode, frequency, amplitude, etc., by performing activation function operations on the input pulse signal. These features can effectively describe the change law of data in the time domain. Distributed feature sets are fused to obtain a fused feature matrix. Data features from different cores are integrated to form a more comprehensive and comprehensive feature description. Through fusion, data redundancy is reduced and the most representative feature information is retained. The fused feature matrix is mapped to a reduced dimension to obtain a feature encoding vector. High-dimensional data is converted into a low-dimensional space, making the data more compact and removing unnecessary noise. The features after dimensionality reduction retain the most representative information while reducing the burden of calculation and storage. Based on the feature encoding vector, global feature synthesis is performed. All core feature information is integrated into a unified, global feature vector.
[0036] S6, generates an S×T dynamic load matrix based on the global eigenvector, redistributes computing tasks to the processing cores based on the neuron state distribution, and obtains a parallel task execution plan.
[0037] Specifically, the global feature vector is mapped in the first dimension and the second dimension. By mapping the feature data to the S-dimensional task space and the T-dimensional load space respectively, a two-dimensional feature projection matrix is obtained, which reflects the relationship between tasks and loads and helps understand the distribution characteristics of computing tasks in different dimensions. Based on the two-dimensional feature projection matrix, a two-layer neuron encoding process is performed to construct a task distribution network and a load mapping network. The first layer of neurons constructs a task distribution network through 1.15 billion neurons to distribute tasks to different computing cores. These tasks are distributed to appropriate cores according to their computational complexity, memory requirements, and data dependencies. With the help of the second layer of neurons, a load mapping network is constructed to fine-tune the tasks according to the load conditions of each computing core to achieve balanced load distribution. Through the two-layer neuron encoding process, an S×T dynamic load matrix is obtained. The dynamic load matrix is subjected to a three-stage computational density analysis process to ensure that tasks can be evenly distributed among cores and fully utilize computing resources. Coarse-grained task segmentation is performed to divide large-scale tasks into multiple smaller subtasks for better subsequent load evaluation. In the medium-grained load assessment phase, the load density of the task is evaluated by analyzing the computing requirements, memory requirements, and data transmission requirements of each subtask to ensure that each subtask can get reasonable resource support. In the fine-grained resource matching phase, the resources between tasks and cores are matched one-to-one, the resource allocation scheme is optimized, and a multi-level resource allocation strategy is obtained. Based on the multi-level resource allocation strategy, two-way task grouping processing is performed. Through forward task dependency analysis, which tasks need to wait for the results of other tasks are identified to determine the dependencies between tasks. At the same time, through reverse resource demand analysis, the resource requirements of each task are calculated to ensure that tasks can allocate resources on demand. Through these two analysis methods, a task association graph is established, and tasks are hierarchically organized based on this graph to form a hierarchical computing task set. This process can ensure that computing tasks are reasonably scheduled according to dependencies and avoid resource conflicts and computing bottlenecks. Based on the hierarchical computing task set, adaptive task migration processing is performed. According to the status and load of each computing core, the distribution of tasks between cores is flexibly adjusted. Through the memory bandwidth of 16PB / s, the master-slave data channel and the backup data channel are constructed to achieve redundant data transmission during task execution and obtain a redundant transmission network. Even if a computing core fails, the task can be quickly switched to the backup core to ensure the continuity and reliability of computing. Multi-dimensional load balancing is performed on redundant transmission networks. In the time dimension, the execution sequence of each task is reasonably arranged through the scheduling algorithm to ensure that there is no conflict in time; in the spatial dimension, the load of each core is more balanced by adjusting the distribution of tasks between various processing cores. Through the 3.5PB / s inter-core communication bandwidth, the task distribution is dynamically adjusted to make full use of computing resources and obtain an optimized load plan.Based on the optimized load plan, mixed-mode task scheduling is performed. By dividing the 380 trillion synaptic operations per second into real-time processing streams and batch processing streams, computing tasks are divided into two categories: real-time streams require immediate response and processing, while batch processing streams are executed during idle periods. At the same time, during the scheduling process, a master-slave switching mechanism and a data synchronization mechanism are established through 5TB / s of inter-chip communication bandwidth to ensure the coordination and synchronization of tasks between different processing units. Through the asynchronous collaborative control mechanism, it is ensured that each computing task can be seamlessly connected between different cores, and finally a parallel task execution plan is obtained to achieve efficient parallel computing on low-power AI processors.
[0038] In one example, a neuron mapping matrix is constructed according to a computing task input by a low-power AI processor, and the neuron mapping matrix is divided into multiple subtask groups according to a data communication bandwidth threshold of a processing core to obtain a parallel processing unit allocation scheme, including:
[0039] Perform data dependency scanning and task node directed graph construction on the computing tasks input by the low-power AI processor to obtain the initial task dependency relationship, calculate the data communication volume between task nodes based on the initial task dependency relationship, and generate a task data flow graph;
[0040] Divide the task data flow graph into neuron units, divide the data into blocks according to the 192KB cache capacity of each neuromorphic core, obtain the neuron mapping matrix, calculate the communication cost based on the neuron mapping matrix, compare the data communication volume with the data communication bandwidth threshold, and obtain the inter-core data transmission cost table;
[0041] Perform task clustering operations based on the inter-core data transmission cost table, assign task nodes with similar communication costs to the same neuromorphic core, obtain the initial subtask group, and perform load balancing analysis on the initial subtask group, calculate the computing density and storage requirements of each subtask group, and obtain the task load distribution diagram;
[0042] Boundary tasks are adjusted based on the task load distribution diagram, and subtask groups whose computing load exceeds the preset threshold are split and reorganized to obtain multiple subtask groups. Neuromorphic core resources are allocated according to the multiple subtask groups, and 128 processing cores are constructed into processing unit groups according to the degree of data dependency between tasks to obtain a parallel processing unit allocation plan.
[0043] In this example, data dependency scanning is performed on the computing tasks input to the low-power AI processor, and a directed graph of task nodes is constructed based on this information. Data dependency scanning can identify the dependencies between tasks, that is, which tasks’ outputs are the inputs of other tasks. Each task is regarded as a node in the graph, and the dependencies between tasks are represented by directed edges. After the directed graph is initially constructed, the initial task dependencies are obtained, and the task data flow graph is generated by calculating the data communication volume between task nodes. The task data flow graph reflects the input-output relationship and communication volume between computing tasks, and is the key basis for the allocation of parallel computing tasks. Each edge in the task data flow graph represents the communication requirements between task nodes, and the node represents the computing task itself. The weight on each edge is the amount of data transmission, that is, the size of data that needs to be transferred between tasks. According to the task data flow graph, the communication density between tasks is analyzed. The task data flow graph is divided into neuron units. Data is divided into blocks according to the cache capacity of each neuromorphic core. Assume that the cache capacity of each neuromorphic core is 192KB. The data of the computing task is divided into multiple data blocks, and the size of each data block should adapt to the cache capacity of the neuromorphic core. To achieve this goal, the amount of data for each task is calculated according to the computing and storage requirements of the task, forming a data block allocation plan to obtain a neuron mapping matrix, where each matrix element represents the cache allocation of the task to a certain neuromorphic core. The communication cost is calculated based on the neuron mapping matrix. The calculation of the communication cost is based on the communication volume between tasks and the data transmission bandwidth between each neuromorphic core. Assuming that the bandwidth limit between the cores is known (for example, the bandwidth is B), and the communication volume of each task is C, the communication cost D is expressed as:
[0044] ;
[0045] in, is the amount of data communication between tasks, is the communication bandwidth between the two cores. The calculated communication cost can help evaluate the effectiveness of task allocation. If the cost is too high, it is necessary to adjust the task allocation plan or optimize the communication path. According to the communication cost, a cost table for data transmission between cores is obtained. On this basis, task clustering operations are performed. Task nodes with similar communication costs are assigned to the same neuromorphic core to minimize cross-core data communication. During the clustering process, the communication cost of each task is analyzed, and tasks with similar costs are grouped together and assigned to the same core. In this process, the communication volume and computing density between tasks are considered. Computational density is usually expressed as the number of computing resources required for each task. Assume that the computing density of task i is , the computational requirements of each task are quantified by the computational density and the execution time of the task. The load balancing analysis will comprehensively consider the computational density and storage requirements of the task to ensure that the computing resources of each core are efficiently utilized. For the tasks after preliminary clustering, load balancing analysis is performed. The purpose of load balancing is to ensure that the load of each neuromorphic core is within an acceptable range to avoid overloading some cores while other cores are idle. The load balancing process is completed by calculating the computational density and storage requirements of each subtask group. For example, for a subtask group, its total computational density It is expressed as the sum of the computational densities of all tasks in the group:
[0046] ;
[0047] in, is the number of tasks in the subtask group, is the computational density of the ith task. By analyzing the computational density and storage requirements of each subtask group, a task load distribution diagram is drawn to show the computational requirements and resource allocation of each subtask group. According to the task load distribution diagram, boundary tasks are adjusted. For overloaded subtask groups, they are split and recombined to achieve balanced load distribution. Suppose the computational load of a subtask group exceeds the preset threshold , the tasks are split in the following way:
[0048] New Task Set=Split ;
[0049] The multiple subtask groups after splitting are redistributed to different neuromorphic cores. By reorganizing and allocating multiple subtask groups, the load of each core is effectively reduced, thereby improving the overall computing efficiency. Parallel processing units are allocated based on the subtask groups after splitting. The execution order of tasks is often constrained by data dependencies. By analyzing the degree of task data dependency, the 128 processing cores are divided into several processing unit groups according to the data dependency between tasks. These processing unit groups are executed in parallel, and the task dependencies between each unit are effectively managed to ensure that data does not conflict during execution.
[0050] In one example, a storage allocation matrix is constructed according to a parallel processing unit allocation scheme, and cache resource allocation and execution sequence scheduling are performed on subtask groups through pulse neural network operations to obtain a computing core mapping scheme, including:
[0051] Scan the storage resources of the parallel processing unit allocation scheme, calculate the ratio of the data volume of each processing unit group to the 192KB cache capacity, obtain the storage requirement matrix, and divide the subtask groups into data blocks based on the storage requirement matrix to obtain the storage allocation matrix;
[0052] The storage allocation matrix is input into the pulse neural network for data flow analysis. The data transmission path is quantified by setting the synaptic weight to obtain the data path diagram. The subtask groups are analyzed for timing dependency and the data transmission sequence is sorted according to the data path diagram to obtain the initial execution timing table.
[0053] Perform parallelism analysis on the initial execution timing table, divide computing resources according to synaptic computing capabilities, obtain a pulse neuron allocation scheme, and map the pulse neuron allocation scheme to the neuromorphic core. Build data communication links between cores based on the storage-computing integrated architecture to obtain a core connection topology diagram.
[0054] Based on the core connection topology diagram, task pipeline orchestration is performed, and neuron operations are assigned to processing cores in data dependency order to obtain a task execution queue. Resource binding is performed on the task execution queue, and a mapping relationship is established between the task and the corresponding neuromorphic core to obtain a computing core mapping solution.
[0055] In this example, based on the parallel processing unit allocation scheme, the storage resources of each processing unit group are scanned. The ratio of the data volume of each processing unit group to its corresponding cache capacity (assuming 192KB) is calculated. For each processing unit group, the computing data volume and storage requirements of each subtask group are evaluated, and the ratio of its required storage space to cache capacity is calculated. Assume that the data volume of task i is , the required storage space is the cache capacity A certain ratio of , then the ratio is calculated by the formula:
[0056] ;
[0057] in, is the amount of storage data required for task i, is the cache capacity of the processing unit (192KB). Through this step, the storage requirement matrix is obtained, which reflects the specific requirements of each processing unit group for storage resources. The storage requirement matrix can provide a basis for data segmentation to ensure that each subtask group can complete data access within the cache capacity. Divide the subtask groups into data blocks based on the storage requirement matrix. Divide the data according to the storage requirements of each subtask group to ensure that the size of the data block can match the cache capacity of each neuromorphic core. Assume that the storage requirement of subtask group i is , then the data block segmentation process is carried out according to the following formula:
[0058] ;
[0059] in, is the number of computing nodes in subtask group i. Through this method, it is ensured that the computing data blocks of each subtask can be effectively allocated to the storage units of the neuromorphic core, and a storage allocation matrix is obtained, in which each element represents the allocation of a specific data block. The storage allocation matrix is input into the pulse neural network for data flow analysis. The pulse neural network can simulate the signal transmission process between neurons and quantitatively calculate the data flow based on the synaptic connection weight. The weight of the synaptic connection determines the transmission strength of data from one neuron to another. Assume that the connection weight is , represents the connection strength between neuron i and neuron j, then the quantitative calculation of the data transmission path is expressed by the following formula:
[0060] Transmission Data ;
[0061] Among them, Transmission is the data transmission cost from neuron i to neuron j, Data is the amount of data transmitted. Data flow analysis derives a data path diagram and optimizes the data transmission path to reduce unnecessary communication overhead. Based on the data path diagram, a timing dependency analysis is performed on the subtask group to identify the execution order of each task and the order of their data transmission. Timing dependency analysis can help the system reasonably arrange the execution order of tasks, avoid conflicts and resource competition during data transmission, and obtain a preliminary execution timing table. Perform parallelism analysis on the initial execution timing table. Reasonably divide computing resources according to the computing characteristics and storage requirements of the tasks. For example, assuming that the computational complexity of task i is , and its corresponding pulse neuron computing power is , then the computing resources required for task i are expressed by the following formula:
[0062] ;
[0063] in, is the computing power of neuron i, is the computational complexity of task i. Based on parallelism analysis, the resource allocation of each task in the pulse neuron is determined, and the tasks are allocated to different processing units according to the demand for computing resources to obtain a pulse neuron allocation scheme, in which each task is assigned to a suitable processing core to ensure that resources are fully utilized. After mapping the pulse neuron allocation scheme to the neuromorphic core, a data communication link between cores is constructed based on the storage-computing integrated architecture. In the storage-computing integrated architecture, storage and computing are closely integrated, and the construction of the data transmission path needs to fully consider the bandwidth limitations and delays between computing nodes. Assume that the bandwidth between cores is , the delay of data transmission Calculated by the following formula:
[0064] ;
[0065] in, represents the amount of data transferred from core i to core j, is the bandwidth between the two cores. By building data communication links between cores, the data transmission path is effectively planned and the core connection topology is obtained. Task pipeline scheduling is performed based on the core connection topology to ensure that tasks are executed in sequence according to data dependency order and that there is no conflict in communication between tasks. Task pipeline scheduling will assign neuron operations to different processing cores and queue tasks in data dependency order. For example, suppose the execution order of task i is The task execution queue is represented as:
[0066] ;
[0067] In the task execution queue, each task is assigned to the appropriate core according to its data dependency. Through resource binding, each task is mapped to the corresponding neuromorphic core, and finally a computing core mapping scheme is formed. The mapping scheme ensures that tasks can be executed in the appropriate time and space, ensuring the effective use of computing resources and the efficiency of data transmission.
[0068] In one example, an anomaly monitoring matrix is constructed based on the computing core mapping scheme, the anomaly monitoring matrix is deployed to the local storage unit of the neuromorphic core, and a cross-core data recovery channel is established to obtain a distributed fault-tolerant solution, including:
[0069] Extract abnormal features from the computing core mapping scheme to obtain an initial abnormal pattern set, and construct monitoring rules for the operating status of each neuromorphic core based on the initial abnormal pattern set to obtain an abnormal monitoring matrix;
[0070] The anomaly monitoring matrix is mapped into blocks, and the monitoring rules are divided into the local storage units of each neuromorphic core according to the 192KB storage capacity threshold to obtain a monitoring rule distribution table. The status of the computing unit in each neuromorphic core is monitored according to the monitoring rule distribution table to obtain a core-level anomaly detection network.
[0071] Based on the core-level anomaly detection network, a cross-core communication link is constructed, and a data backup channel is established through the 3.5PB / s inter-core communication bandwidth to obtain a data transmission link group. A fault-tolerant strategy is deployed for the data transmission link group, and a data redundant storage structure is constructed using neurons to obtain a distributed backup solution.
[0072] According to the distributed backup solution, the abnormal recovery channel is constructed, the backup data is mapped to the available computing cores through the storage and computing integrated architecture, and the fault switching mechanism is obtained. The collaborative control strategy of the fault switching mechanism is configured, and a cross-core data recovery process is established through the asynchronous event-driven mechanism to obtain a distributed fault-tolerant solution.
[0073] In this example, abnormal features are extracted from the computing core mapping scheme to identify abnormal behaviors that occur during the computing process, such as timeouts, data loss, and calculation errors. It is assumed that by analyzing the execution results of each task node, an initial set of abnormal patterns is obtained, which contains different types of abnormal behavior patterns. For example, the pattern set includes overload of computing nodes, memory leaks, or cache overflows. Based on the preliminary abnormal patterns, monitoring rules are established for the operating status of each neuromorphic core. Each monitoring rule contains several conditions, such as time limit exceedance, computing load of processing units, etc. Based on the monitoring rules, an abnormal monitoring matrix is constructed. Assume that the monitoring rules Detect the neuromorphic core i, where the rules include monitoring the core's computing resource usage , memory usage , and task completion time , then the construction process of the abnormal monitoring matrix is expressed by the following formula:
[0074] ;
[0075] in, is the constructed monitoring matrix, is the monitoring rule for core i. Based on the constructed abnormal monitoring matrix, block mapping is performed. Each monitoring rule is allocated according to the local cache capacity of the neuromorphic core (for example, each core has a cache capacity of 192KB). Since the amount of data involved in each monitoring rule may exceed the cache capacity of the core, the monitoring rules are divided into blocks. The cache capacity of each core is set to , assuming the task The amount of monitoring data is , the storage requirement ratio of each monitoring rule is calculated by the following formula:
[0076] ;
[0077] According to the storage requirement ratio of each monitoring rule, the distribution of each rule in the neuromorphic core is determined, and a monitoring rule distribution table is generated, which records the monitoring rule distribution information of each core. According to the monitoring rule distribution table, the state of the computing unit in each core is monitored to form a core-level anomaly detection network. By monitoring the real-time status of each core, such as memory usage, computing task latency, etc., this network can effectively detect and identify potential anomalies. Based on the core-level anomaly detection network, cross-core communication link construction is performed. In a distributed system, when a core fails or an anomaly occurs, redundant calculations need to be performed by other cores. In order to ensure that the system can continue to run under any circumstances, a data backup channel is established between cores. The communication bandwidth between cores is set to , then the data transfer time from core i to core j is Calculated by the following formula:
[0078] ;
[0079] In this process, the communication bandwidth between cores (for example, 3.5PB / s) is used as a key resource, and a data backup channel is established through the bandwidth to ensure that the core detected by the abnormality is quickly restored. Through this process, the data transmission link group obtained allocates redundant communication paths for each core. Based on the data transmission link group, a fault-tolerant strategy is deployed. The data redundant storage structure is constructed using the neuron model to improve the fault tolerance of the system through redundant backup. For example, assuming that the data block Backup is required between core i and core j. The following formula is used to describe the implementation of redundant storage:
[0080] ;
[0081] in, Indicates the redundant backup of core i, Backup Indicates that the data block Back up to core j. Based on the distributed backup solution, the data backup is mapped to the available computing cores through the storage and computing architecture. This process requires an accurate fault switching mechanism to ensure that when a fault occurs, the system can automatically switch to the backup core and continue computing. The design of the fault switching mechanism requires the system to be able to monitor and sense the health status of the computing core, and when a fault is found, migrate the task to the normal core. Through the asynchronous event-driven mechanism, the fault switching process can be carried out seamlessly. Set the delay of the fault detection mechanism to , when a failure occurs, the system calculates the recovery time using the following formula:
[0082] ;
[0083] Among them, Switching Delay is the time delay of fault switching, Bandwidth is the communication bandwidth of the backup core, and Data Volume is the amount of data that needs to be restored. During the recovery process, the collaborative control strategy configuration is used to ensure that the task migration of each computing core is transparent and that the data synchronization during the recovery process can maintain consistency. Through the asynchronous event-driven mechanism, during the task migration and data recovery process, different cores can execute tasks according to the predetermined order, thereby avoiding task conflicts. Through the above steps, a stable distributed fault-tolerant solution is established, so that the system can effectively detect, recover and switch when facing a computing core failure, ensuring the high reliability of the system.
[0084] In one example, parallel matrix operations are performed on subtask groups according to a distributed fault-tolerant solution, data blocks are allocated to corresponding neuron processing units according to a storage-computation integrated architecture, and an asynchronous data transmission channel is established to obtain a parallel computing result set, including:
[0085] The data redundancy storage scheme in the distributed fault-tolerant scheme is divided into blocks, and the computing data is grouped according to the 192KB cache capacity of the neuromorphic core to obtain a data block allocation table;
[0086] Based on the data block allocation table, a computational dependency graph is constructed to map the matrix operation tasks to the processing network formed by neurons, and a computational task mapping scheme is obtained;
[0087] Perform parallel analysis on the computing task mapping scheme, combine data and computing instructions into atomic operation units through the storage-computing integrated architecture, and obtain the operation instruction sequence;
[0088] According to the operation instruction sequence, an asynchronous communication link is established, a data transmission channel network is constructed, a data flow channel graph is obtained, and a data pipeline is designed based on the data flow channel graph, and the continuous calculation process is decomposed into multi-level parallel execution units to obtain a parallel calculation pipeline;
[0089] Load control is performed on the parallel computing pipeline, and the data transmission rate is dynamically adjusted using the inter-core communication bandwidth to obtain the computing beat sequence. Matrix block operations are performed according to the computing beat sequence, and the calculation results are summarized through the 5TB / s inter-chip communication bandwidth to obtain the distributed computing results.
[0090] The consistency of distributed computing results is verified, and the outputs of each processing unit are merged through the cross-core data verification mechanism to obtain the parallel computing result set.
[0091] In this example, considering that each neuromorphic core has a cache capacity of 192KB, the computational data is grouped so that each data block can fit into the core’s cache capacity, thus avoiding cache overflow or excessive latency caused by excessive data size. Assume that the size of the data block is , the core cache capacity is , the allocation of each data block is described by the following formula:
[0092] ;
[0093] in, is the number of data blocks, It ensures that the size of each data block can adapt to the cache capacity of the core. In the specific implementation, the computing data will be divided into multiple data blocks of the same size to generate a data block allocation table. In the data block allocation table, each data block will be bound to a specific computing task or core to ensure that the computing task can maximize the use of storage resources when executing. Through this allocation strategy, the efficient flow of computing data is ensured, over-reliance on a single core or cache is avoided, and the parallelism and efficiency of computing are improved. Based on the data block allocation table, a computing dependency graph is constructed. The computing dependency graph is a graphical model that describes the data dependency relationship between computing tasks. Each node in the graph represents a computing task, and each edge represents the data transmission dependency between tasks. The purpose of constructing a computing dependency graph is to effectively map computing tasks to the processing network formed by neurons. The nodes in the graph are mapped to each neuromorphic core, and the edges represent the task execution order and the data transmission path. For example, if computing task A depends on the result of task B, there will be a directed edge between A and B in the computing dependency graph. In this way, the execution order of tasks can be clearly specified to avoid data conflicts and deadlock problems in parallel computing. The generation process of the computing dependency graph is expressed by the following formula:
[0094] ;
[0095] in, To calculate the task set, is the set of dependencies between tasks. By computing the dependency graph, the mapping scheme of computing tasks is effectively constructed. The parallelism of the computing task mapping scheme is analyzed to analyze the parallelism of each computing task. Through the storage-computing integrated architecture, data and computing instructions are combined into atomic operation units for processing. In this architecture, data and computing instructions coexist in the same physical unit, which greatly improves computing efficiency. Assume that the computing instruction sequence of each operation unit is , then the total calculation sequence of all operation units is obtained by the following formula:
[0096] ;
[0097] in, represents the total number of operation units, and Represents the instruction sequence of the jth operation unit. Through this combination, the computing tasks are split and optimized according to the instruction stream, the degree of parallelism is improved, and it is ensured that each computing unit can complete the task in the shortest time. An asynchronous communication link is established according to the operation instruction sequence, and a data transmission channel network is constructed. The design of the asynchronous communication link is to ensure that data between tasks can be transmitted without waiting, so that each task unit can run independently and improve the throughput of the system. In the process of constructing the data flow channel diagram, tasks are assigned to different processing units and the data transmission path is defined. Assume that the bandwidth of the transmission path is , then the delay time of data flow is calculated by the following formula:
[0098] ;
[0099] in, is the amount of data transferred between task i and task j, is the bandwidth between the two. According to the calculation, the data flow channel can be effectively designed to ensure that the data transmission between tasks does not become a bottleneck. Based on the data flow channel diagram, the data pipeline design is performed to decompose the calculation process into multiple levels of parallel execution units, and each unit executes its task in chronological order. By designing a multi-stage pipeline, the calculation process is decomposed into different stages so that each stage can be executed independently without interfering with each other. Assume that the number of stages of the pipeline is The execution time of each stage is , the overall pipeline time is:
[0100] ;
[0101] In this way, computing tasks can be completed in parallel by different execution units, improving the processing efficiency of the system. Load control is performed on the parallel computing pipeline. In the load control stage, the data transmission rate is dynamically adjusted through the communication bandwidth between cores to ensure that the computing load is evenly distributed among the cores. Assume that the bandwidth between cores is , then the load of each core is balanced by the following formula:
[0102] ;
[0103] in, is the amount of data processed by core k, is the bandwidth of the core. By dynamically adjusting the data transmission rate and adjusting the task allocation according to the real-time load, the system load is balanced. The matrix block operations are aggregated through the 5TB / s inter-chip communication bandwidth to obtain the distributed computing results. In this process, high-bandwidth communication links are used to ensure that the results of large-scale matrix calculations can be quickly aggregated from multiple processing cores to obtain the final calculation results. The cross-core data verification mechanism verifies the consistency of the output results of each processing unit to ensure that the final parallel computing result set is correct.
[0104] In one example, a time domain feature analysis is performed on a parallel computing result set, and data feature extraction is achieved by establishing a cross-core pulse signal communication network to obtain a global feature vector, including:
[0105] Perform time series analysis on the parallel computing result set, establish time domain sampling points based on the 4096 state configurations of each neuron, obtain a time feature sequence, and convert the data into pulse signals based on the time feature sequence. Map the time series features into a pulse sequence through a synaptic connection network to obtain a characteristic pulse stream.
[0106] Based on the characteristic pulse flow, cross-core channel configuration is performed, and 128 neuromorphic cores are connected through synapses to build a signal transmission network to obtain a pulse transmission topology map. The bandwidth of the pulse transmission topology map is allocated to obtain a characteristic transmission scheme.
[0107] Data feature extraction is performed according to the feature transmission scheme, and feature recognition operation is performed on the pulse signal using neurons to obtain a distributed feature set, and feature fusion is performed based on the distributed feature set to obtain a fused feature matrix;
[0108] The fused feature matrix is mapped to a dimension reduction mode to obtain a feature coding vector, and global feature synthesis is performed based on the feature coding vector to obtain a global feature vector.
[0109] In this example, the parallel computing result set is analyzed in time series, time domain features are extracted from the calculation results, and these features are mapped into pulse signals that can be processed by the neural network. Each neuron has 4096 state configurations, which represent the state changes of the neuron at different time steps. By establishing time domain sampling points on these state configurations, a time feature sequence is formed to capture the changes of neuron states over time. Assume that the state sequence of each neuron is ,in represents the neuron index, Represents the time step, and by sampling the state, a time domain feature sequence is obtained:
[0110] ;
[0111] in, is a discrete time step in the time domain. The time feature sequence is converted into a pulse sequence emitted by the neuron through pulse signal conversion. The pulse signal conversion maps the characteristic value of each time step into a pulse, and the frequency and intensity of the pulse represent the corresponding state characteristics. The dynamic behavior of neurons is effectively captured and transmitted to other neurons in the network through the frequency and time interval of the pulse signal. Between neurons, pulse signals are transmitted through synaptic connections to form a pulse stream, which contains the timing characteristics from each neuron. Based on the pulse signal stream, cross-core channel configuration is performed. Assume that there are 128 neuromorphic cores, and each core transmits information through synaptic connections. In this process, a signal transmission network is designed for these cores to ensure that each core can receive the required pulse signal in time to achieve parallel computing. In order to optimize the signal transmission efficiency, a pulse transmission topology graph is constructed, in which each node represents a neuromorphic core, and the edge connecting each node represents the signal transmission channel. Assume that the pulse signal transmission bandwidth is ,in and Represents the core index, and through the pulse transmission topology diagram, bandwidth allocation is performed to obtain a characteristic transmission plan. The goal of bandwidth allocation is to ensure that the signal can be transmitted stably between the cores to avoid signal delay or loss due to insufficient bandwidth. Through bandwidth allocation, the data transmission rate of each core is optimized, thereby improving the computing efficiency of the overall system. The process of bandwidth allocation is expressed by the following formula:
[0112]
[0113] in, is the total amount of transmitted signal, is the time required for signal transmission, = is the bandwidth requirement. Data feature extraction is performed according to the feature transmission scheme. After receiving the pulse signal, each neuron converts these signals into meaningful feature information through its synaptic connection. The feature extraction process includes extracting key information features from the pulse signal, such as the frequency, intensity and timing relationship of the pulse. This information is the basis for neuron feature recognition. Assume that the feature vector of the pulse signal is , the neuron calculates the pulse signal and extracts the characteristics of each neuron:
[0114] ;
[0115] in, Represents a neuron-based spike signal sequence After collecting the features of all neurons, a distributed feature set is obtained to represent the state of the entire neural network. Feature fusion is performed based on the distributed feature set. The features from different neurons are merged to obtain a more global feature representation, and a fused feature matrix is obtained, in which each element contains the comprehensive information of multiple neurons. The process of feature fusion is expressed by the following formula:
[0116] ;
[0117] in, is the weight of each neuron feature, indicating its importance in the fusion process, It's a neuron The characteristic vector of is the fused feature matrix. The fused feature matrix is mapped down to reduce the dimension and convert the high-dimensional feature matrix into a low-dimensional feature coding vector to reduce the computational complexity and improve the computational efficiency. The process of dimensionality reduction mapping is performed through principal component analysis, t-SNE and other algorithms to obtain the feature coding vector. Global feature synthesis is performed based on the feature coding vector to obtain the global feature vector. The global feature vector is an efficient expression of the entire neural network state.
[0118] In one example, a dynamic load matrix of S×T is generated according to the global eigenvector, and computing tasks are redistributed to the processing cores based on the neuron state distribution to obtain a parallel task execution scheme, including:
[0119] Perform first-dimension mapping and second-dimension mapping processing on the global feature vector, map the feature data to the S-dimensional task space and the T-dimensional load space respectively, and obtain a two-dimensional feature projection matrix;
[0120] Based on the two-dimensional feature projection matrix, a two-layer neuron encoding process is performed, and the task distribution network is constructed using the 1.15 billion neurons in the first layer. The load mapping network is established through the second layer of neurons to obtain the S×T dynamic load matrix.
[0121] A three-stage computational density analysis process is performed on the dynamic load matrix, which includes coarse-grained task segmentation, medium-grained load evaluation, and fine-grained resource matching, to obtain a multi-level resource allocation strategy.
[0122] According to the multi-level resource allocation strategy, two-way task grouping is performed, and a task association graph is established through forward task dependency analysis and reverse resource demand analysis to obtain a hierarchical computing task set;
[0123] Based on the hierarchical computing task set, adaptive task migration processing is performed, and the master-slave data channel and backup data channel are constructed using the 16PB / s memory bandwidth to obtain a redundant transmission network;
[0124] Multi-dimensional load balancing is performed on redundant transmission networks, combining time dimension scheduling and space dimension allocation, and dynamically adjusting task distribution through 3.5PB / s inter-core communication bandwidth to obtain an optimized load solution;
[0125] Mixed-mode task scheduling is performed according to the optimized load plan, and 380 trillion synaptic operations per second are divided into real-time processing stream and batch processing stream to obtain a dual-channel execution sequence. The dual-channel execution sequence is then asynchronously controlled and processed. The master-slave switching mechanism and data synchronization mechanism are established through the 5TB / s inter-chip communication bandwidth to obtain a parallel task execution plan.
[0126] In this example, the global feature vector is mapped in the first dimension and the second dimension, and the feature data is mapped to the task space and the load space respectively to obtain a two-dimensional feature projection matrix. Assume that there is a global feature vector , which contains various features of the computing task. Map it to the first dimension to obtain the feature representation in the task space , and maps it to the feature representation in the load space The resulting two-dimensional feature projection matrix is:
[0127] ;
[0128] in, and Represent the feature vectors in the task space and load space respectively. In this process, the first dimension mapping represents mapping the feature to the complexity space of the task, and the second dimension mapping maps the feature to the load demand space. Based on the two-dimensional feature projection matrix, a two-layer neuron encoding process is performed to construct a task distribution network and a load mapping network. In the first layer of neurons, 1.15 billion neurons are used to build a task distribution network, through which tasks are allocated to different computing resources. This process is expressed as:
[0129] ;
[0130] in, represents the processing function of the first layer of neurons, is a feature in the task space. Through the neural network, tasks are mapped to various processing units. In the second layer of neurons, the load mapping network maps tasks to load resources to obtain a load distribution solution:
[0131] ;
[0132] in, represents the mapping function of the second layer of neurons, is a feature in the load space. We get Dynamic load matrix of dimensions , where each element Indicates The task is in The load distribution on the processing cores. Based on the dynamic load matrix, the computational density is analyzed and processed in three stages. The first is coarse-grained task segmentation. Through this stage, the task is roughly divided into multiple subtask units. The goal of coarse-grained segmentation is to make a preliminary division of tasks according to their computational complexity so that the load of the computational process is as balanced as possible. Secondly, a medium-grained load assessment is performed to determine the load requirement of each subtask by evaluating the load characteristics of the task. Assume that the load requirement of each subtask is , then the load assessment formula is:
[0133] ;
[0134] in, is the total number of tasks, The task is The load on the processing unit. In the final fine-grained resource matching stage, the resource consumption of each task on each computing unit is accurately calculated to ensure that the resources of each computing unit will not be overloaded. The processing at this stage makes the load distribution more accurate, and finally obtains a multi-level resource allocation strategy. Based on the multi-level resource allocation strategy, two-way task grouping processing is performed. The execution order between tasks is determined through forward task dependency analysis, while computing resources are allocated more accurately through reverse resource requirement analysis. The goal of forward task dependency analysis is to determine which tasks must be completed before other tasks, while reverse resource requirement analysis helps to allocate corresponding computing resources according to the resource requirements of the tasks. Based on these analysis results, a task association graph is established to obtain a hierarchical computing task set. . For hierarchical computing task sets, adaptive task migration processing is performed. According to the real-time load status of each computing unit, the execution position of the task is dynamically adjusted to achieve optimal resource utilization. The master-slave data channel and the backup data channel are built through the 16PB / s memory bandwidth to ensure the smoothness of task migration, and data is backed up between different cores to improve the fault tolerance of the system. Through the redundant transmission network, reliable transmission of tasks is achieved to ensure that the system can continue to run even if some computing units fail. In order to improve performance, multi-dimensional load balancing processing is performed on the redundant transmission network. Combined with time dimension scheduling and space dimension allocation, the data transmission rate between cores is optimized. Through the dynamic adjustment of the 3.5PB / s inter-core communication bandwidth, the distribution of tasks is adjusted in real time to ensure that the load of each computing core is balanced. Mixed mode task scheduling processing is performed according to the optimized load plan. The 380 trillion synaptic operations per second are divided into real-time processing flow and batch processing flow. The real-time processing flow is responsible for processing tasks with high latency requirements, while the batch processing flow is responsible for processing tasks that can be executed with delay. In this way, the execution load of tasks is effectively distributed to different streams, thereby improving the throughput and response speed of the system. The execution sequences of real-time processing streams and batch processing streams are asynchronously coordinated and controlled. Through the 5TB / s inter-chip communication bandwidth, a master-slave switching mechanism and a data synchronization mechanism are established to ensure that data synchronization and backup can be completed in real time during task execution. Asynchronous collaborative control enables the system to run efficiently and stably when facing complex tasks, and ultimately obtains a parallel task execution solution.
[0135] Reference Figure 2 , this embodiment provides a parallel computing chip based on a low-power AI processor, including:
[0136] A partitioning module 1 is used to construct a neuron mapping matrix according to the computing tasks input by the low-power AI processor, and divide the neuron mapping matrix into multiple subtask groups according to the data communication bandwidth threshold of the processing core to obtain a parallel processing unit allocation plan;
[0137] The orchestration module 2 is used to construct a storage allocation matrix according to the parallel processing unit allocation scheme, allocate cache resources and arrange execution sequences for the subtask groups through pulse neural network operations, and obtain a computing core mapping scheme;
[0138] Establish module 3, which is used to build an anomaly monitoring matrix based on the computing core mapping scheme, deploy the anomaly monitoring matrix to the local storage unit of the neuromorphic core, and establish a cross-core data recovery channel to obtain a distributed fault-tolerant solution;
[0139] The computing module 4 is used to perform parallel matrix operations on the subtask groups according to the distributed fault-tolerant scheme, distribute the data blocks to the corresponding neuron processing units according to the storage-computation integrated architecture, and establish an asynchronous data transmission channel to obtain a parallel computing result set;
[0140] Analysis module 5, used for performing time domain feature analysis on the parallel computing result set, extracting data features by establishing a cross-core pulse signal communication network, and obtaining a global feature vector;
[0141] The allocation module 6 is used to generate a dynamic load matrix of S×T according to the global feature vector, and to reallocate computing tasks to the processing cores based on the neuron state distribution to obtain a parallel task execution plan.
[0142] In this embodiment, for the specific implementation of each unit in the above chip embodiment, please refer to the above method embodiment, which will not be described in detail here.
[0143] An embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, the above method is implemented. It can be understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.
[0144] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media provided by the present invention and used in the embodiments may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double-speed data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM.
[0145] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, device, article or method including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, device, article or method. In the absence of further restrictions, an element defined by the sentence "includes a ..." does not exclude the presence of other identical elements in the process, device, article or method including the element.
[0146] The above description is only a preferred embodiment of the present invention, and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A parallel computing method based on a low-power AI processor, characterized in that: The following steps are involved: A neuron mapping matrix is constructed according to the computing tasks input by the low-power AI processor, and the neuron mapping matrix is divided into a plurality of subtask groups according to the data communication bandwidth threshold of the processing core to obtain a parallel processing unit allocation scheme; Constructing a storage allocation matrix according to the parallel processing unit allocation scheme, performing cache resource allocation and execution sequence arrangement on the subtask groups through pulse neural network operation, and obtaining a computing core mapping scheme; An anomaly monitoring matrix is constructed based on the computing core mapping scheme, the anomaly monitoring matrix is deployed to the local storage unit of the neuromorphic core, and a cross-core data recovery channel is established to obtain a distributed fault-tolerant solution; Perform parallel matrix operations on the subtask group according to the distributed fault-tolerant scheme, distribute data blocks to corresponding neuron processing units according to the storage-computation integrated architecture, and establish an asynchronous data transmission channel to obtain a parallel computing result set; Performing time domain feature analysis on the parallel computing result set, extracting data features by establishing a cross-core pulse signal communication network, and obtaining a global feature vector; An S×T dynamic load matrix is generated according to the global feature vector, and computing tasks are redistributed to processing cores based on neuron state distribution to obtain a parallel task execution scheme, wherein S is the dimension of features in task space and T is the dimension of features in load space.
2. The parallel computing method based on a low-power AI processor according to claim 1, characterized in that: The method constructs a neuron mapping matrix according to the computing tasks input by the low-power AI processor, divides the neuron mapping matrix into multiple subtask groups according to the data communication bandwidth threshold of the processing core, and obtains a parallel processing unit allocation scheme, including: Perform data dependency scanning and task node directed graph construction on the computing tasks input by the low-power AI processor to obtain initial task dependency relationships, calculate the data communication volume between task nodes based on the initial task dependency relationships, and generate a task data flow graph; Divide the task data flow graph into neuron units, divide the data into blocks according to the 192KB cache capacity of each neuromorphic core, obtain a neuron mapping matrix, calculate the communication cost based on the neuron mapping matrix, compare the data communication volume with the data communication bandwidth threshold, and obtain an inter-core data transmission cost table; Performing task clustering operations according to the inter-core data transmission cost table, assigning task nodes with similar communication costs to the same neuromorphic core, obtaining an initial subtask group, and performing load balancing analysis on the initial subtask group, calculating the computing density and storage requirements of each subtask group, and obtaining a task load distribution diagram; Boundary tasks are adjusted based on the task load distribution diagram, and subtask groups whose computing load exceeds a preset threshold are split and reorganized to obtain multiple subtask groups. Neuromorphic core resources are allocated according to the multiple subtask groups, and 128 processing cores are constructed into processing unit groups according to the degree of data dependency between tasks to obtain a parallel processing unit allocation plan.
3. The parallel computing method based on a low-power AI processor according to claim 2, characterized in that: The step of constructing a storage allocation matrix according to the parallel processing unit allocation scheme, performing cache resource allocation and execution sequence arrangement on the subtask groups through pulse neural network calculations, and obtaining a computing core mapping scheme includes: Scan the parallel processing unit allocation scheme for storage resources, calculate the ratio of the data volume of each processing unit group to the 192KB cache capacity to obtain a storage requirement matrix, and divide the subtask groups into data blocks based on the storage requirement matrix to obtain a storage allocation matrix; The storage allocation matrix is input into a pulse neural network for data flow analysis, the data transmission path is quantitatively calculated by setting synaptic weights to obtain a data path diagram, and the subtask groups are subjected to timing dependency analysis and data transmission sequence sorting according to the data path diagram to obtain an initial execution timing table; Performing parallelism analysis on the initial execution timing table, dividing computing resources according to synaptic computing capabilities, obtaining a pulse neuron allocation scheme, and mapping the pulse neuron allocation scheme to the neuromorphic core, building a data communication link between cores based on a storage-computing integrated architecture, and obtaining a core connection topology diagram; Based on the core connection topology diagram, task pipeline orchestration is performed, neuron operations are allocated to processing cores in data dependency order, a task execution queue is obtained, and resource binding is performed on the task execution queue. A mapping relationship is established between the task and the corresponding neuromorphic core to obtain a computing core mapping solution.
4. The parallel computing method based on a low-power AI processor according to claim 3, characterized in that: The abnormality monitoring matrix is constructed based on the computing core mapping scheme, the abnormality monitoring matrix is deployed to the local storage unit of the neuromorphic core, and a cross-core data recovery channel is established to obtain a distributed fault-tolerant solution, including: Extracting abnormal features from the computing core mapping scheme to obtain an initial abnormal pattern set, and constructing monitoring rules for the operating state of each neuromorphic core based on the initial abnormal pattern set to obtain an abnormal monitoring matrix; The abnormal monitoring matrix is mapped in blocks, and the monitoring rules are divided into local storage units of each neuromorphic core according to the 192KB storage capacity threshold to obtain a monitoring rule distribution table, and the state of the computing unit in each neuromorphic core is monitored according to the monitoring rule distribution table to obtain a core-level abnormality detection network; Based on the core-level anomaly detection network, a cross-core communication link is constructed, a data backup channel is established through a 3.5PB / s inter-core communication bandwidth, a data transmission link group is obtained, and a fault-tolerant strategy is deployed for the data transmission link group. A data redundant storage structure is constructed using neurons to obtain a distributed backup solution; According to the distributed backup solution, abnormal recovery channel construction is executed, the backup data is mapped to the available computing cores through the storage and computing integrated architecture, a fault switching mechanism is obtained, and a collaborative control strategy is configured for the fault switching mechanism. A cross-core data recovery process is established through an asynchronous event-driven mechanism to obtain a distributed fault-tolerant solution.
5. The parallel computing method based on a low-power AI processor according to claim 4, characterized in that: The method of performing parallel matrix operations on the subtask group according to the distributed fault-tolerant scheme, allocating data blocks to corresponding neuron processing units according to the storage-computation integrated architecture, and establishing an asynchronous data transmission channel to obtain a parallel computing result set includes: The data redundancy storage scheme in the distributed fault-tolerant scheme is divided into blocks, and the calculation data is grouped according to the 192KB cache capacity of the neuromorphic core to obtain a data block allocation table; Building a computation dependency graph based on the data block allocation table, mapping the matrix operation task to the processing network formed by neurons, and obtaining a computation task mapping scheme; Perform parallelism analysis on the computing task mapping scheme, combine data and computing instructions into atomic operation units through a storage-computation integrated architecture, and obtain a computing instruction sequence; Establishing an asynchronous communication link according to the operation instruction sequence, constructing a data transmission channel network, obtaining a data flow channel graph, and performing data pipeline design based on the data flow channel graph, decomposing the continuous calculation process into multi-level parallel execution units, and obtaining a parallel calculation pipeline; The parallel computing pipeline is load controlled, the data transmission rate is dynamically adjusted using the inter-core communication bandwidth to obtain a computing beat sequence, and matrix block operations are performed according to the computing beat sequence, and the calculation results are summarized through the 5TB / s inter-chip communication bandwidth to obtain distributed computing results; The distributed computing results are verified for consistency, and the outputs of each processing unit are merged through a cross-core data verification mechanism to obtain a parallel computing result set.
6. The parallel computing method based on a low-power AI processor according to claim 5, characterized in that: The time domain feature analysis is performed on the parallel computing result set, and data feature extraction is achieved by establishing a cross-core pulse signal communication network to obtain a global feature vector, including: Performing time series analysis on the parallel computing result set, establishing time domain sampling points based on 4096 state configurations of each neuron, obtaining a time feature sequence, performing pulse signal conversion on the data according to the time feature sequence, mapping the time series features into a pulse sequence through a synaptic connection network, and obtaining a characteristic pulse stream; Based on the characteristic pulse flow, cross-core channel configuration is performed, 128 neuromorphic cores are connected through synapses to construct a signal transmission network, a pulse transmission topology diagram is obtained, and bandwidth is allocated to the pulse transmission topology diagram to obtain a characteristic transmission scheme; Performing data feature extraction according to the feature transmission scheme, performing feature recognition operation on the pulse signal using neurons to obtain a distributed feature set, and performing feature fusion based on the distributed feature set to obtain a fused feature matrix; The fusion feature matrix is subjected to dimensionality reduction mapping to obtain a feature coding vector, and global feature synthesis is performed according to the feature coding vector to obtain a global feature vector.
7. The parallel computing method based on a low-power AI processor according to claim 6, characterized in that: The method of generating an S×T dynamic load matrix according to the global feature vector and redistributing computing tasks to processing cores based on neuron state distribution to obtain a parallel task execution scheme includes: Perform first dimension mapping and second dimension mapping processing on the global feature vector, map the feature data to an S-dimensional task space and a T-dimensional load space respectively, and obtain a two-dimensional feature projection matrix; Based on the two-dimensional feature projection matrix, a double-layer neuron encoding process is performed, 1.15 billion neurons in the first layer are used to build a task distribution network, and a load mapping network is established through the second layer of neurons to obtain a dynamic load matrix of S×T; Performing a three-stage computational density analysis process on the dynamic load matrix, sequentially performing coarse-grained task segmentation, medium-grained load evaluation, and fine-grained resource matching, to obtain a multi-level resource allocation strategy; Perform bidirectional task grouping processing according to the multi-level resource allocation strategy, establish a task association graph through forward task dependency analysis and reverse resource demand analysis, and obtain a hierarchical computing task set; Based on the hierarchical computing task set, adaptive task migration processing is performed, and a master-slave data channel and a backup data channel are constructed using 16PB / s memory bandwidth to obtain a redundant transmission network; The redundant transmission network is subjected to multi-dimensional load balancing processing, combined with time dimension scheduling and space dimension allocation, and task distribution is dynamically adjusted through 3.5PB / s inter-core communication bandwidth to obtain an optimized load solution; According to the optimized load scheme, mixed-mode task scheduling is performed, and 380 trillion synaptic operations per second are divided into real-time processing stream and batch processing stream to obtain a dual-channel execution sequence. The dual-channel execution sequence is then asynchronously collaboratively controlled, and a master-slave switching mechanism and a data synchronization mechanism are established through the 5TB / s inter-chip communication bandwidth to obtain a parallel task execution scheme.
8. A parallel computing chip based on a low-power AI processor, characterized in that: For implementing the steps of the method according to any one of claims 1 to 7, the chip comprises: A partitioning module is used to construct a neuron mapping matrix according to the computing tasks input by the low-power AI processor, and divide the neuron mapping matrix into multiple subtask groups according to the data communication bandwidth threshold of the processing core to obtain a parallel processing unit allocation plan; An orchestration module, configured to construct a storage allocation matrix according to the parallel processing unit allocation scheme, and to perform cache resource allocation and execution sequence orchestration on the subtask groups through pulse neural network operations to obtain a computing core mapping scheme; Establishing a module for constructing an anomaly monitoring matrix based on the computing core mapping scheme, deploying the anomaly monitoring matrix to a local storage unit of the neuromorphic core, and establishing a cross-core data recovery channel to obtain a distributed fault-tolerant solution; A computing module, configured to perform parallel matrix operations on the subtask groups according to the distributed fault-tolerant scheme, distribute data blocks to corresponding neuron processing units according to the storage-computation integrated architecture, and establish an asynchronous data transmission channel to obtain a parallel computing result set; An analysis module is used to perform time domain feature analysis on the parallel computing result set, extract data features by establishing a cross-core pulse signal communication network, and obtain a global feature vector; The allocation module is used to generate an S×T dynamic load matrix according to the global feature vector, and redistribute computing tasks to the processing cores based on the neuron state distribution to obtain a parallel task execution plan, wherein S is the dimension of the feature in the task space and T is the dimension of the feature in the load space.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Load balancing method of low-power AI processor, chip and storage medium
CN119271418A
Neural network acceleration apparatus and method, and device and computer storage medium
WO2023116314A1