Distributed Processing Method, Device and Storage Medium of Neuromorphic Computing Chip

Through the combination of graph neural network and deep reinforcement learning network, computing resource allocation is dynamically optimized, and the asynchronous event triggering mechanism and an integrated storage and computing architecture are adopted, the problems of low resource scheduling efficiency and large communication overhead of neuromimicry computing chips in distributed task processing are solved, achieving efficient resource utilization and energy efficiency improvement.

CN119806847BActive Publication Date: 2025-05-27SHENZHEN ITZR TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510294703.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-05-27
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

In distributed task processing scenarios, neuromimicry computing chips have problems such as low resource scheduling efficiency, unbalanced task allocation, and large communication overhead between cores, which affect the overall performance and energy efficiency ratio.

Method used

By introducing graph neural networks to conduct in-depth analysis of task characteristics and dependencies, combined with deep reinforcement learning networks, dynamic optimization allocation of computing resources is achieved, and task parallel processing is adopted to reduce data transmission overhead based on the integrated storage and computing architecture.

Benefits of technology

It significantly improves the resource utilization and processing efficiency of neuromimicry computing chips, reduces energy consumption, enhances the system's adaptability in complex computing environments, and provides an efficient solution for large-scale neural network computing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119806847B_ABST
    Figure CN119806847B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of neuromorphic computing chips, and discloses a distributed processing method, device and storage medium for neuromorphic computing chips. The method includes: decomposing the features of the computing tasks of the neuromorphic computing chip to obtain task encoding vectors; collecting the working states of neuron cores to obtain core state matrices; inputting the task encoding vectors and core state matrices into an initial deep reinforcement learning network to generate resource allocation strategies; reconstructing the sequences of the resource allocation strategies and outputting mapping schemes between neuron cores and computing tasks; configuring synaptic connection weights according to the mapping schemes, converting the computing tasks into pulse signal sequences, and transmitting and executing them between neuron cores through an asynchronous event triggering mechanism; calculating synaptic weight update gradients for parameter optimization to obtain a target deep reinforcement learning network. The present invention improves the resource utilization rate and processing efficiency of neuromorphic computing chips and reduces energy consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of neuromorphic computing chips, and particularly to a distributed processing method, device, and storage medium for neuromorphic computing chips. Background Art

[0002] With the rapid development of artificial intelligence technology, traditional von Neumann architecture computing chips face serious power consumption and performance bottlenecks when dealing with large-scale neural network computations. This is mainly because in the traditional architecture, the computing unit is separated from the storage unit, and data is frequently transferred between the two, resulting in high energy consumption. At the same time, the serial processing method also limits the improvement of computing efficiency. To solve this problem, neuromorphic computing chips have emerged. They process information by mimicking the working mechanism of the human brain, adopting a computing-in-memory architecture and an asynchronous event-driven mechanism. However, current neuromorphic computing chips still have problems such as low resource scheduling efficiency, unbalanced task allocation, and large communication overhead between cores in distributed task processing scenarios, which seriously affect the overall performance and energy efficiency ratio of the chips.

[0003] Especially when dealing with multi-task parallel computing, how to effectively allocate computing resources, balance the load of processing cores, reduce communication overhead, and how to dynamically adjust the processing strategy according to task characteristics are all key technical problems that need to be solved urgently. Existing task scheduling methods often have difficulty in fully utilizing the computing resources of neuromorphic chips and cannot effectively respond to dynamically changing computing requirements. Summary of the Invention

[0004] The present invention provides a distributed processing method, device, and storage medium for neuromorphic computing chips, which improve the resource utilization rate and processing efficiency of neuromorphic computing chips and reduce energy consumption.

[0005] In a first aspect, the present invention provides a distributed processing method for a neuromorphic computing chip. The distributed processing method for the neuromorphic computing chip includes:

[0006] Decompose the features of the computing tasks of the neuromorphic computing chip, establish a task topology graph, and perform message passing aggregation operations on the task topology graph to obtain a task encoding vector;

[0007] Collect the working states of neuron cores in the neuromorphic computing chip to obtain a core state matrix including synaptic weight distribution, neuron activation threshold, and occupancy rate of the computing-in-memory unit;

[0008] Input the task encoding vector and the core state matrix into an initial deep reinforcement learning network, construct an action value function based on the Actor-Critic architecture, and generate a resource allocation strategy;

[0009] Perform sequence reconstruction on the resource allocation strategy, calculate the temporal dependence relationship between tasks, and output the mapping scheme between neuron cores and computing tasks;

[0010] Configure synaptic connection weights according to the mapping scheme, convert the computing tasks into a pulse signal sequence, and transmit and execute them between the neuron cores through an asynchronous event triggering mechanism;

[0011] Collect the processing results of the pulse signal sequence, calculate the synaptic weight update gradient, optimize the parameters of the action value function of the initial deep reinforcement learning network, and obtain the target deep reinforcement learning network.

[0012] In a second aspect, the present invention provides a distributed processing device for a neuromorphic computing chip. The distributed processing device for the neuromorphic computing chip includes:

[0013] A decomposition module, configured to perform feature decomposition on the computing tasks of the neuromorphic computing chip, establish a task topology graph, and perform message passing aggregation operations on the task topology graph to obtain a task encoding vector;

[0014] An acquisition module, configured to collect the working states of neuron cores in the neuromorphic computing chip, and obtain a core state matrix including synaptic weight distribution, neuron activation threshold, and occupancy rate of memory and computing units;

[0015] A construction module, configured to input the task encoding vector and the core state matrix into an initial deep reinforcement learning network, construct an action value function based on the Actor-Critic architecture, and generate a resource allocation strategy;

[0016] A reconstruction module, configured to perform sequence reconstruction on the resource allocation strategy, calculate the temporal dependence relationship between tasks, and output the mapping scheme between neuron cores and computing tasks;

[0017] An execution module, configured to configure synaptic connection weights according to the mapping scheme, convert the computing tasks into a pulse signal sequence, and transmit and execute them between the neuron cores through an asynchronous event triggering mechanism;

[0018] An optimization module, configured to collect the processing results of the pulse signal sequence, calculate the synaptic weight update gradient, optimize the parameters of the action value function of the initial deep reinforcement learning network, and obtain the target deep reinforcement learning network.

[0019] The third aspect of the present invention provides a distributed processing device for a neuromorphic computing chip, including: a memory and at least one processor, wherein instructions are stored in the memory; the at least one processor calls the instructions in the memory to enable the distributed processing device of the neuromorphic computing chip to execute the above-mentioned distributed processing method of the neuromorphic computing chip.

[0020] The fourth aspect of the present invention provides a computer-readable storage medium, in which instructions are stored, and when it runs on a computer, it enables the computer to execute the above-mentioned distributed processing method of the neuromorphic computing chip.

[0021] In the technical solution provided by the present invention, through the introduction of a graph neural network for in-depth analysis of task features and dependencies, combined with a deep reinforcement learning network, dynamic optimization allocation of computing resources is achieved. An asynchronous event-triggered mechanism is adopted for parallel processing of tasks, and the data transmission overhead is reduced based on the in-memory computing architecture. At the same time, through an adaptive parameter optimization method, online learning and optimization of the deep reinforcement learning network are realized, enabling the system to dynamically adjust the processing strategy according to the real-time load situation. This method significantly improves the resource utilization rate and processing efficiency of the neuromorphic computing chip, reduces energy consumption, enhances the adaptability of the system in complex computing environments, and provides an efficient solution for large-scale neural network computing. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0023] Figure 1 It is a schematic flowchart of the distributed processing method of the neuromorphic computing chip provided by the embodiments of the present application;

[0024] Figure 2 It is a schematic block diagram of the structure of the distributed processing device of the neuromorphic computing chip provided by the embodiments of the present application;

[0025] Figure 3 It is a schematic block diagram of the structure of the distributed processing device of the neuromorphic computing chip provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0027] The flowchart shown in the accompanying drawings is only an example illustration, and does not necessarily include all contents and operations / steps, nor does it necessarily need to be executed in the described order. For example, some operations / steps can also be decomposed, combined, or partially merged. Therefore, the actual execution order may change based on the actual situation.

[0028] It should also be understood that the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. As used in the specification of this application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0029] It should be further understood that the term "and / or" used in the specification of this application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0030] Next, some embodiments of this application will be described in detail in conjunction with the accompanying drawings. Without conflict, the features in the following embodiments and the embodiments can be combined with each other.

[0031] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of the distributed processing method of the neuromorphic computing chip provided by the embodiment of this application. As Figure 1 shown, the distributed processing method of the neuromorphic computing chip provided by the embodiment of this application includes steps S100 to S600.

[0032] Step S100: Decompose the computing tasks of the neuromorphic computing chip, establish a task topology graph, and perform message passing aggregation operations on the task topology graph to obtain a task encoding vector;

[0033] It can be understood that the execution subject of the present invention can be a distributed processing device of a neuromorphic computing chip, or a terminal or a server. Specifically, it is not limited here. The embodiment of the present invention takes the server as the execution subject as an example for illustration.

[0034] Specifically, collect the timing characteristics of the computing tasks of the neuromorphic computing chip, and extract task characteristics such as task arrival time, the number of required neurons, the number of synaptic connections, and data flow directions. These characteristics provide the basic information of each computing task and help identify the computing requirements and resource consumption of the tasks. Perform data dependence analysis on the task characteristic set, analyze the data transmission relationship and computing dependence between tasks, and calculate the dependence relationship between tasks. The dependence relationship between tasks is established by determining the input-output data flow and execution order between tasks. Based on the dependence relationship between tasks, construct a task topology graph. The task topology graph is a directed graph structure, where each node represents a computing task, and the edges in the graph represent the dependence relationship between tasks. Map the characteristic set of tasks to the attributes of graph nodes. The attributes of nodes include information such as the computing requirements and required resources of the tasks. At the same time, the dependence strength between tasks is determined by calculating the data transmission volume or transmission frequency between tasks, so as to map the dependence relationship to the edge weights in the graph and obtain a specific task topology graph. Perform multi-head attention calculation on the nodes in the task topology graph. This process is similar to applying the attention mechanism in a graph neural network, enabling the model to focus on the key features of each task node. Through the multi-head attention mechanism, each node can not only rely on its own information but also focus on the information of its neighboring nodes connected to it, thereby dynamically adjusting its representation ability. Input the node attention features into the graph convolutional layer. In the graph convolutional layer, perform a weighted summation operation on the features of adjacent nodes through a message aggregation function to enhance the correlation between tasks and extract the hidden layer representation of the nodes. The hidden layer representation of the nodes is a high-dimensional representation obtained by combining its own and neighborhood information, which can effectively capture the complex relationships and information flows between tasks. Calculate the global features of the graph according to the hidden layer representation of the nodes. Perform a pooling operation on the features of all nodes to reduce the dimension and integrate the important information of the nodes in the graph. At the same time, perform a dimension transformation on the structural information of the graph, which is achieved through a fully connected neural network, and map the features of the tasks to a target space. Through a non-linear transformation, these features are mapped to a low-dimensional space to obtain a compressed feature vector, which effectively represents the overall requirements and behavior patterns of the tasks in the neuromorphic computing chip. Perform feature fusion on the compressed feature vector and the hidden layer representation of the nodes to enhance the expression ability of the task encoding vector. Feature fusion is achieved through a residual connection. The residual connection can retain the original feature information and combine the compressed feature vector with the hidden layer representation of the nodes through an addition operation to obtain the encoding vector of the task.

[0035] Step S200: Collect the working status of the neuron cores in the neuromorphic computing chip to obtain a core status matrix including synaptic weight distribution, neuron activation threshold, and occupancy rate of the memory and computing unit;

[0036] Specifically, parallel scanning is performed on the synaptic connections of the neuron cores in the neuromorphic computing chip to extract information such as the synaptic weights, connection status, and transmission bandwidth of each neuron core. By calculating the weight density of the synaptic connection features, a synaptic weight distribution map is obtained. The synaptic weight distribution map shows the connection strength between neurons and the density of information transmission, which helps analyze the communication load and execution efficiency of neurons. Sampling the neuron states of each neuron core yields a set of activation thresholds for the neurons. The activation threshold is a key parameter for each neuron to determine whether to trigger, and the firing frequency of the neuron is affected by this threshold. Sampling and analyzing these thresholds helps understand the working state of the neurons and their performance in task execution. Based on the set of activation thresholds, state mapping is performed, and through quantization calculations, the pulse triggering conditions and corresponding firing frequencies of the neuron cores are obtained. Quantization calculations can help evaluate the activation state of the neurons and reveal whether the neurons are overloaded or idle. Through the obtained threshold distribution sequence and the memory and computing units of the neuron cores, workload analysis is carried out to calculate the load status value. The load status value reflects the working intensity and resource consumption of the neuron core at a certain moment. After inputting the load status value into the adaptive sampler, the real-time occupancy rate of the memory and computing units is captured by dynamically adjusting the sampling frequency, resulting in an occupancy rate sequence. The adaptive sampler adjusts the sampling frequency according to the real-time changing load conditions, thereby improving the sampling efficiency and making the collected data more conform to the actual working state of the chip. The occupancy rate sequence and the synaptic weight distribution map are integrated in terms of features. Through multi-dimensional matrix operations, these feature data are merged and processed to obtain a state feature matrix, which contains information such as synaptic weight distribution, neuron activation thresholds, and the occupancy rate of the memory and computing units, reflecting the working state of the neuron core. Based on the state feature matrix and the threshold distribution sequence, a data fusion operation is performed. By tensor operations, the state information from multiple data sources is merged, and all the feature data are fused together to obtain a core state matrix.

[0037] Step S300: Input the task encoding vector and the core state matrix into the initial deep reinforcement learning network, construct an action-value function based on the Actor-Critic architecture, and generate a resource allocation strategy;

[0038] Specifically, the dimension alignment is performed on the task encoding vector and the core state matrix, and the features of the task and the core state data of the chip are fused into a unified input format. Through the way of feature recombination, the data from different sources are mapped to a 128-dimensional input feature space to obtain the state input vector. The state input vector is input into the Actor network. The Actor network models the action space according to the computing capabilities of neurons and synapses and outputs the preliminary action distribution. The preliminary action represents different resource allocation schemes that may be taken under the current resource state. To ensure the feasibility of the action, the feasibility constraint is executed on the initial action distribution. Based on the hardware resource constraints of the system, such as the memory bandwidth of 16 PB / s and the inter-core communication bandwidth of 3.5 PB / s, the resource allocation verification is carried out to ensure that the generated action is executable under the actual hardware resource conditions, and the effective action set is obtained. The effective action set contains all resource allocation schemes that meet the hardware constraints, ensuring that the subsequent calculations can be efficiently executed within the limits of the hardware capabilities. The state input vector and the effective action set are input into the Critic network. The task of the Critic network is to calculate the value evaluation of the state. Based on 38 trillion synaptic operations and 24 trillion neuron operations, the Critic network can estimate the value of each state and provide a quantitative evaluation of the effects of different resource allocation strategies. The value evaluation function is calculated according to the mapping of the current state and the action set, providing key feedback for the policy update of the Actor network. The temporal difference operation (TD) is performed on the value evaluation function, and the immediate reward is constructed by calculating the cache utilization rate of 192 KB for each neuron core. The immediate reward reflects the immediate effect of the current state and action combination, helping to evaluate the advantages and disadvantages of different resource allocation schemes. By calculating the TD error, the deviation between the current policy and the expected policy is measured, so as to perform gradient update on the policy distribution of the Actor network. The gradient update is optimized by the asynchronous event-driven method, so as to increase the action selection probability and ensure that the resource allocation policy continuously tends to be optimal. After the policy update is completed, the updated policy is input into the resource scheduler. The resource scheduler performs actual resource allocation according to the policy. In this process, the scheduler performs action mapping based on the configuration space of 4096 neuron states to obtain the scheduling scheme set. The scheduling scheme set contains multiple potential resource allocation schemes, and each scheme corresponds to different computing task allocations and resource utilization methods. To select the best scheme, the performance of these schemes is evaluated, and the evaluation criteria include indicators such as the occupancy rate of the memory-computation unit and the synaptic connection load. By calculating these indicators and sorting the schemes, the optimal resource allocation scheme is finally selected to obtain the effective resource allocation policy.

[0039] Step S400: Reconstruct the sequence of the resource allocation policy, calculate the temporal dependence relationship between tasks, and output the mapping scheme between neuron cores and computing tasks;

[0040] Specifically, sequence encoding is performed on the resource allocation strategy. By using a Seq2Seq encoder, the scheduling requests of multiple concurrent tasks are converted into an input sequence. The Seq2Seq model encodes the scheduling requests from the perspectives of time series and tasks, converting complex task allocation relationships into task time series vectors. The task time series vectors are fed as inputs into the encoder of a recurrent neural network (RNN). The design of the RNN can capture the dependencies between tasks in the time series, and construct the data flow dependencies between tasks through time series feature extraction to obtain a time series feature matrix. Hierarchical clustering is performed on the time series feature matrix. The hierarchical clustering method divides task groups based on the data flow direction and transmission bandwidth requirements between tasks. By analyzing the mutual dependencies and resource requirements between tasks, tasks are effectively organized into a hierarchical structure to obtain a hierarchical dependency tree. The hierarchical dependency tree presents the priorities and dependencies between tasks, which helps with priority assignment and resource optimization in subsequent task scheduling. Based on the hierarchical dependency tree, an attention weight matrix is constructed, and task priority scores are calculated through the attention mechanism in the RNN decoder. The attention mechanism allows the model to automatically focus on the features of other tasks that are most relevant to each task being processed, thus allocating resources more precisely. Through the calculation of task priorities, a scheduling priority sequence is obtained, which determines the execution order of tasks and the priorities of resource allocation. In this process, the generation of the scheduling priority sequence is dynamic and can be adjusted as the dependencies between tasks change. The scheduling priority sequence is matched with the computing resources of the neuron cores. Resource matching is performed based on the synaptic connection status and memory-computation unit occupancy rate of each neuron core to ensure that tasks can be executed on appropriate cores, obtaining an effective core allocation plan. The core allocation plan is input into the time series verification module, and the rationality of the task execution time series is verified through the decoder of the recurrent neural network. The decoder can simulate the execution process of tasks and check whether there are conflicts when tasks are actually executed. By verifying the task execution time series, an execution time series graph is obtained, showing the execution order of tasks at different time nodes. Conflict detection is performed on the execution time series graph. The resource competition situation is analyzed based on the neuron activation threshold and synaptic weight distribution. If resource competition or conflicts are found between tasks, a conflict resolution strategy is generated. The conflict resolution strategy formulates a strategy to avoid resource conflicts by analyzing the conflicts between tasks, ensuring that all tasks can be executed smoothly without task delays caused by resource shortages or conflicts. The conflict resolution strategy includes means such as task rearrangement and resource reallocation to help re-plan the execution order of tasks and make reasonable adjustments according to the current resource situation. The core allocation plan is adjusted through the conflict resolution strategy to form the final mapping plan between neuron cores and computing tasks. This mapping plan ensures the optimal utilization of resources through task rearrangement and resource reallocation, while avoiding conflicts or delays in task execution.

[0041] Step S500: Configure the synaptic connection weights according to the mapping scheme, convert the computing task into a pulse signal sequence, and transmit and execute it among neuron cores through an asynchronous event triggering mechanism;

[0042] Specifically, analyze the connection relationship of the mapping scheme. By analyzing the computing tasks and data transmission requirements of each neuron core, construct a connection topology structure to reflect the relationship between tasks and neuron cores. Based on the connection topology structure, generate a synaptic connection matrix to show the connection situation between different neurons and their interaction paths. After inputting the synaptic connection matrix into the weight configurator, the weight configurator configures the weights of each synapse according to the information transmission intensity of each connection. This process is completed by an event-based spiking neural network, where the weight of each connection reflects the efficiency and priority of information transmission, obtaining the synaptic weight distribution. Perform normalization processing on the synaptic weight distribution to ensure that the weights of all synapses are compared and adjusted on a unified scale. Perform weight scaling based on the neuron activation threshold and the occupancy rate of the memory and computing unit to maintain the balance and stability of computing under different resource usage conditions. Normalize the scaled weights to form standardized weights suitable for neural pulse coding. After inputting the standardized weights into the task coding module, convert them into a timing signal through pulse frequency modulation to generate a pulse coding sequence. The pulse coding sequence is a timing signal containing task load information, and the execution intensity and order of the task are reflected by the pulse frequency. Perform event-triggered mapping on the pulse coding sequence. Due to the certain delay and refractory period in the firing characteristics of neurons, set the trigger threshold and refractory period parameters for each pulse signal to ensure that the pulse signal is triggered at the appropriate time point and avoid repeated triggering. Through these parameters, obtain a set of trigger rule sets to help the system coordinate the activities between different neurons to ensure the synchrony and correctness of the signal. After inputting the trigger rule set into the asynchronous controller, the asynchronous controller can adjust the execution timings of multiple neuron cores in an event-driven manner. In this way, the asynchronous controller can dynamically control the work of different cores, enabling the computing task to be coordinated and executed among each core, avoiding conflicts and resource waste. Propagate and route the pulse signals in the execution control sequence. Determine the signal transmission path and intensity based on the synaptic connection weights. By calculating the weights of synaptic connections, determine which paths require more computing resources, thereby optimizing the signal transmission efficiency and ensuring that the pulse signal can be transmitted to the target neuron quickly and accurately. According to the signal transmission scheme, establish data paths between neuron cores, and these paths can efficiently transmit and process pulse signals. Through an asynchronous parallel mechanism, realize the transmission and calculation of pulse signals among multiple neuron cores to complete the processing of the computing task. The asynchronous parallel mechanism enables the neuromorphic computing chip to maintain high-efficiency resource utilization and computing efficiency even when multiple tasks are executed simultaneously.

[0043] Step S600: Collect the processing results of the pulse signal sequence, calculate the synaptic weight update gradient, optimize the parameters of the action value function of the initial deep reinforcement learning network, and obtain the target deep reinforcement learning network.

[0044] Specifically, index statistics are performed on the processing results of the pulse signal sequence. By statistically analyzing the computing and communication performance of 128 neuron cores in the neuromorphic computing chip, the execution efficiency of the system is evaluated, and a performance sampling matrix is obtained, which includes the performance of each neuron core during the execution of computing tasks, such as computing speed, communication delay and other indicators. A performance loss function is constructed by comparing the target task completion time with the actual execution time to quantify the performance loss of the system during the actual execution process. The calculation of the loss function can reflect the deficiencies of the system when processing tasks, providing a basis for subsequent optimization. Based on the performance loss function, loss calculation is performed on the performance sampling matrix to obtain a loss vector. The loss vector reveals the performance bottlenecks and potential problems of the system through gradient analysis, helping to construct the gradient update direction. Combining the synaptic weight distribution and the neuron activation threshold, the gradient update strategy is optimized to obtain a weight adjustment matrix. Based on the weight adjustment matrix, the state of the memory and computing unit is updated. This process constructs an optimization objective by calculating the workload and resource occupancy rate of each memory and computing unit to obtain a state update vector. The state update vector provides a basis for dynamic adjustment of the resource scheduling and task allocation of the memory and computing unit, ensuring that the system can perform optimal resource allocation and computing task allocation under different workloads. After inputting the state update vector into the Actor network, the Actor network iteratively optimizes the policy parameters of the action value function through the dynamic programming algorithm. The dynamic programming method can continuously adjust the policy through the feedback of multiple iterative calculations, and gradually find the optimal resource allocation and task execution strategy. During the optimization process, the optimization strategy generated by the Actor network needs to be evaluated by the Critic. The Critic network calculates the value evaluation score of the state-action pair according to the current policy and constructs a value network to evaluate the quality of the current policy. Through the calculated value evaluation score, the Critic network can provide more accurate feedback to the Actor network to help it adjust the policy. At this time, the update amount of the value function is the result after Critic evaluation, and this amount represents the deviation of the current policy from the ideal policy. The value function update amount is input into the parameter optimizer, and the weight parameters of the deep reinforcement learning network are updated by the stochastic gradient descent method (SGD). SGD is an optimization algorithm that adjusts the weight parameters in each iteration, gradually minimizes the loss function, and improves the network performance. The updated set of weight parameters is the optimized parameter set, providing the latest resource allocation and computing task scheduling strategy for the deep reinforcement learning network. Through the optimized parameter set, the parameters of the initial deep reinforcement learning network are updated, causing the network structure to gradually converge to the optimal solution. The update of the network parameters is completed through the backpropagation algorithm, which adjusts the weights and biases of each layer of the network according to the gradient of the error to optimize the network structure. After multiple iterations, the target deep reinforcement learning network is finally obtained, which can better adapt to the resource management requirements of the neuromorphic computing chip and achieve efficient task scheduling and resource allocation.

[0045] In an embodiment of the present invention, by introducing a graph neural network to deeply analyze task features and dependencies, combined with a deep reinforcement learning network, dynamic optimization allocation of computing resources is achieved. An asynchronous event-triggered mechanism is adopted for parallel processing of tasks, and the data transmission overhead is reduced based on the in-memory computing architecture. Meanwhile, through an adaptive parameter optimization method, online learning and optimization of the deep reinforcement learning network are realized, enabling the system to dynamically adjust the processing strategy according to the real-time load situation. This method significantly improves the resource utilization rate and processing efficiency of the neuromorphic computing chip, reduces energy consumption, enhances the adaptability of the system in complex computing environments, and provides an efficient solution for large-scale neural network computing.

[0046] In a specific embodiment, the process of executing step S100 may specifically include the following steps:

[0047] Collect the timing features of the computing tasks of the neuromorphic computing chip, extract a task feature set including the task arrival time, the number of required neurons, the number of synaptic connections, and the data flow direction, and perform data dependency analysis on the task feature set to calculate the task dependencies based on the data transmission relationships between tasks;

[0048] Construct a directed graph structure based on the task dependencies, map the task feature set to the graph node attributes, and map the dependency strength matrix to the edge weights to obtain the task topology graph;

[0049] Perform multi-head attention calculation on the nodes in the task topology graph to obtain the node attention features, and input the node attention features into the graph convolutional layer. Through the message aggregation function, perform weighted summation operations on the features of adjacent nodes to obtain the node hidden layer representations;

[0050] Calculate the global features of the graph according to the node hidden layer representations, perform pooling operations and dimension transformation on the node features to obtain the graph structure encoding, and input the graph structure encoding into a fully connected neural network to map the features to the target dimension space through a non-linear transformation to obtain the compressed feature vector;

[0051] Fuse the compressed feature vector and the node hidden layer representations, and obtain the task encoding vector through residual connection.

[0052] Specifically, temporal features of computational tasks are collected, and a set of important task feature sets are extracted. These feature sets include task arrival time, the number of required neurons, the number of synaptic connections, and data flow direction, etc. The task arrival time refers to the timestamp when the task starts to execute, reflecting the time information of task scheduling. The number of required neurons refers to the number of neuron cores required to execute the task, reflecting the demand for task computing resources. The number of synaptic connections represents the connection strength between the task and the neuron cores, while the data flow direction describes the data transmission path in the task. Data dependence analysis is performed on the task feature set, and task dependence relationships are calculated based on the data transmission relationships between tasks. The data transmission relationships between tasks are established by calculating the dependence of the input and output data of the tasks. For example, if task B requires the output of task A as input, then task B depends on task A. These dependence relationships between tasks are represented by a dependence strength matrix, where each element represents the dependence strength between two tasks. The dependence strength is calculated by the following formula:

[0053] ;

[0054] where depend represents the dependence strength of task on task , is the time delay between task and task . The longer the delay time, the lower the dependence strength. A directed graph structure is constructed based on the task dependence relationships. Each node in this graph represents a task, and each edge represents the dependence relationship between tasks. The task feature set is mapped to the graph node attributes. The attributes of each node include information such as task arrival time, the number of required neurons, the number of synaptic connections, and data flow direction, etc., which can comprehensively describe the characteristics of the task. And the dependence strength matrix is mapped to the edge weights, and the edge weights reflect the strength of the dependence relationships between tasks. Multi-head attention calculation is performed on the nodes in the task topology graph. The multi-head attention mechanism is a method that captures various dependence relationships between different nodes in the graph through multiple attention heads. The multi-head attention calculation helps to extract the important features of the nodes, especially to capture the complex dependence relationships between tasks. For each node , its attention features are calculated by the following formula:

[0055] ;

[0056] where, represents the set of neighbor nodes of node , is the attention weight between node and neighbor node , is node The eigenvector. Attention weight is calculated as follows:

[0057] ;

[0058] where, represents the similarity between node and node and is calculated by dot product or other similarity metrics. Through multi-head attention calculation, the attention features of each node are obtained, and these features can reflect the degree of association between the node and other nodes. The attention features of the nodes are input into the graph convolutional layer. The graph convolutional network performs a weighted summation operation on the features of adjacent nodes through the message aggregation function to enhance the expressive ability of the node representation. Suppose the hidden layer representation of node is , and the update rule of the graph convolutional layer is expressed as:

[0059] ;

[0060] where, is the weight matrix of the graph convolutional layer, is the activation function, represents the set of neighbor nodes of node . Through the aggregation process, the node hidden layer representation can fuse the key information from neighbor nodes. Calculate the global features of the graph according to the node hidden layer representation. The global features are obtained through pooling operations (such as max pooling or average pooling) on the node features and dimension transformation. The pooling operation can compress the node features, reduce redundant information, and at the same time extract the global information of the graph. The pooled features are non-linearly transformed through a fully connected neural network to map the features of the graph into the target dimensional space to obtain the compressed feature vector. To enhance the representation ability of the graph, feature fusion is performed on the compressed feature vector and the node hidden layer representation. Through residual connection, the node hidden layer representation is combined with the compressed feature vector to retain the local information of the node and the global information of the graph. The fusion process is achieved through the following formula:

[0061] ;

[0062] where, represents the fused task encoding vector, is the compressed feature vector, is the node hidden layer representation. The introduction of residual connection enables the model to be effectively trained and optimized, and at the same time avoids the loss of information during the transmission process. The task encoding vector obtained after feature fusion is used as the input for task scheduling and resource allocation, providing an effective basis for the distributed processing of the neuromorphic computing chip.

[0063] In a specific embodiment, the process of executing step S200 may specifically include the following steps:

[0064] Parallelly scan the synaptic connections of the neuron cores in the neuromorphic computing chip, extract the synaptic weights, connection states, and transmission bandwidths of each core, and obtain synaptic connection features;

[0065] Calculate the weight density of the synaptic connection features to obtain a weight distribution map, and sample the neuron states of each neuron core to obtain an activation threshold set;

[0066] Based on the activation threshold set, perform state mapping, quantitatively calculate the pulse triggering conditions and firing frequencies of the neuron cores to obtain a threshold distribution sequence, and perform a workload analysis on the threshold distribution sequence and the memory and computing units to obtain a load status value;

[0067] Input the load status value into an adaptive sampler, capture the real-time occupancy rate of the memory and computing units by dynamically adjusting the sampling frequency to obtain an occupancy rate sequence, and perform feature integration on the occupancy rate sequence and the weight distribution map, and construct a state feature matrix through multi-dimensional matrix operations;

[0068] Based on the state feature matrix and the threshold distribution sequence, perform data fusion, and merge multi-source state information through tensor operations to obtain a core state matrix.

[0069] Specifically, parallelly scan the synaptic connections in the neuron cores, and extract key information such as the synaptic weights, connection states, and transmission bandwidths of each neuron core. These information constitute the synaptic connection features of the neuron cores. The synaptic weight refers to the strength between the synapses connecting different neurons, the connection state reflects whether an effective connection is established between neurons, and the transmission bandwidth represents the ability and speed of data transmission between neuron cores. Calculate the weight density of the synaptic connection features to generate a weight distribution map. The weight density reflects the connection strength and resource occupancy between neuron cores by calculating the density of synaptic connections. For example, assume the synaptic weight matrix is , where represents the connection weight between neuron core and core , and the weight density is defined as:

[0070] ;

[0071] where, represents the total number of all connections of neuron core , reflects neuron core The weight density represents the connection strength and data transmission density between neuron cores. Through calculation, a weight distribution map is obtained, showing the distribution characteristics of the connections between neuron cores in the entire neuromorphic computing chip. Sampling the neuron states of each neuron core yields a set of activation thresholds. The activation threshold refers to the minimum activation value at which a neuron core generates a pulse after receiving an input signal and is related to the excitability of the neuron. Assume that the activation threshold of each neuron core is , and the set of activation thresholds is the set of all neuron core activation thresholds, denoted as , where represents the total number of neuron cores. By sampling the thresholds, the activity characteristics of the neuron cores are obtained, providing data support for the quantitative calculation of the pulse triggering conditions and firing frequencies. Based on the set of activation thresholds, state mapping is performed to quantify the pulse triggering conditions and firing frequencies of the neuron cores. The pulse triggering condition refers to when a neuron core will emit a pulse, and the firing frequency refers to the frequency at which a neuron core emits pulses. To quantify these conditions, the following calculation formula is used:

[0072] ;

[0073] where, represents the firing frequency of neuron core , is the activation threshold of this core, represents the synaptic signal strength transmitted from neuron core to core . By quantitatively calculating the pulse triggering conditions and firing frequencies, the activity frequency distribution of each neuron core is obtained. At the same time, the threshold distribution sequence and the memory and computing units are analyzed for workload, and a load status value is obtained. The load status value reflects the workload of the neuron core and takes into account the computing task's demand for the neuron core resources. Assume that the load status value is calculated by the following formula:

[0074] ;

[0075] where, represents the computing demand of task for neuron core , represents neuron core Resource availability. Through analysis, the load status values of each neuron core are obtained, and these values are used to judge the resource occupancy of the neuron core. The load status values are input into the adaptive sampler. The adaptive sampler captures the occupancy rate of the memory and computing unit in real time by dynamically adjusting the sampling frequency, and obtains the occupancy rate sequence. The occupancy rate refers to the degree of resource occupancy of the memory and computing unit when executing a computing task, reflecting the resource consumption of the task. By continuously adjusting the sampling frequency, the adaptive sampler can accurately capture the dynamic resource occupancy of the memory and computing unit, ensuring the accuracy and real-time nature of resource allocation. The occupancy rate sequence reflects the resource occupancy of the memory and computing unit at each moment, which can help adjust the resource allocation strategy to optimize the performance of the chip. Integrate the features of the occupancy rate sequence and the weight distribution map, and construct a state feature matrix through multi-dimensional matrix operations. The state feature matrix contains the load status information of the neuron core and integrates the features of the occupancy rate and the weight distribution. Integrate the occupancy rate sequence and the weight distribution map through the following matrix operations:

[0076] state_matrix = weight_distribution occupancy_rate;

[0077] where state_matrix represents the state feature matrix, weight_distribution is the weight distribution of synaptic connections, and occupancy_rate is the occupancy rate sequence of the memory and computing unit. Through matrix operations, information from different sources is integrated to generate a comprehensive state feature matrix. Perform data fusion based on the state feature matrix and the threshold distribution sequence. Merge multi-source state information through tensor operations to obtain the core state matrix. Tensor operation is an efficient multi-dimensional data processing method that can effectively integrate feature information from different sources. Assume the state feature matrix is , and the threshold distribution sequence is , then the fused core state matrix M is expressed as:

[0078] ;

[0079] In this way, different types of state information are fused, and the generated core state matrix can comprehensively reflect the working state of the neuron core.

[0080] In a specific embodiment, the process of executing step S300 may specifically include the following steps:

[0081] Align the dimensions of the task encoding vector and the core state matrix, and construct a 128-dimensional input feature space through feature recombination to obtain the state input vector;

[0082] Input the state input vector into the Actor network, model the action space based on the computing capabilities of neurons and synapses to obtain the initial action distribution, and perform feasibility constraints on the initial action distribution. Verify the resource allocation based on the 16 PB / s memory bandwidth and 3.5 PB / s inter-core communication bandwidth to obtain the effective action set;

[0083] Input the state input vector and the effective action set into the Critic network, calculate the state value based on 380 trillion synaptic operations and 240 trillion neuron operations to obtain the value evaluation function;

[0084] Perform temporal difference operations on the value evaluation function, construct the immediate reward by calculating the 192 KB cache utilization rate of each neuron core to obtain the TD error, and perform gradient updates on the policy distribution of the Actor network based on the TD error. Optimize the action selection probability through the asynchronous event-driven method to obtain the updated policy;

[0085] Input the updated policy into the resource scheduler, perform action mapping based on the configuration space of 4096 neuron states to obtain the set of scheduling schemes, and perform performance evaluation on the set of scheduling schemes. Sort the schemes by calculating the occupancy rate of the memory-computation unit and the synaptic connection load to obtain the resource allocation strategy.

[0086] Specifically, align the dimensions of the task encoding vector and the core state matrix so that they can meet the input requirements of the deep reinforcement learning model. Assume that the dimension of the task encoding vector is and the dimension of the core state matrix is . Then, merge the two into a 128-dimensional input feature space through feature recombination. Feature recombination is completed by normalizing, standardizing, or linearly mapping each feature. For example, the task encoding vector is linearly transformed and mapped to dimensions, while the core state matrix is mapped to dimensions through appropriate dimension expansion. The information of these two parts is fused into a 128-dimensional input vector through concatenation or weighted summation. This input vector is the state input vector, representing the task requirements and neuron core states of the chip at a specific moment. When the state input vector is input into the Actor network, the goal of the network is to model the action space according to the computing capabilities of neuron cores and synapses. The core task of action space modeling is to allocate resources for each possible computing task to optimize the overall computing efficiency and performance. Assume that the computing capabilities of neuron cores and synapses are represented by vectors and . The action space modeling is as follows:

[0087] ;

[0088] where, Represents the action distribution generated by the model, which describes the resource allocation probability between each task and the neuron core. Based on the memory bandwidth (16 PB / s) and the inter-core communication bandwidth (3.5 PB / s), feasibility constraints are imposed on the initial action distribution. These bandwidth limitations have an important impact on the scheduling of computing tasks. Excessive memory or communication bandwidth requirements lead to inefficient task execution or system overload, and the action distribution is adjusted according to the actual bandwidth. For example, if the memory requirement of task A is 5 PB and the bandwidth limit is 16 PB / s, then the action probability of this task should be affected by the bandwidth limit and adjusted to a reasonable value. Based on this constraint, an effective action set is obtained , which contains all effective actions that can meet the bandwidth requirements. The state input vector and the effective action set are input into the Critic network. The goal of the Critic network is to evaluate the value of the given state and action, and it adopts the temporal difference learning method. The Critic network calculates the state value function based on the computing power of the synapses (380 trillion operations) and the computing power of the neurons (240 trillion operations) , and its formula is expressed as:

[0089] ;

[0090] where, represents the value of the state , is the discount factor, represents the immediate reward at time step , is the maximum time step. The calculation of the immediate reward is closely related to the cache utilization rate of the neuron core. For example, assume that the cache size of each neuron core is 192 KB, and the cache utilization rate is calculated by the following formula:

[0091] ;

[0092] where, represents the amount of cache used by the neuron core , then reflects the cache utilization rate of this core. By calculating the cache utilization rate of each neuron core, an immediate reward is generated for each state, and the TD error is calculated:

[0093] ;

[0094] The TD error represents the deviation of the current policy in the given state and is used to adjust the policy distribution of the Actor network. Based on the TD error, the policy of the Actor network is optimized through gradient update. The update formula is:

[0095] ;

[0096] Among them, represents the parameters of the Actor network, is the learning rate, represents the probability of selecting an action under a given state . Through an asynchronous event-driven manner, the action selection probability of the Actor network is optimized to obtain an updated policy. The updated policy is input into the resource scheduler. The goal of the scheduler is to schedule tasks according to the configuration space of 4096 neuron states. The configuration space represents the state combinations of neuron cores when executing tasks, and each combination corresponds to different resource requirements and computing capabilities. The scheduler generates a corresponding set of scheduling schemes , and each scheduling scheme contains detailed information on how to allocate resources to achieve optimal task execution. After obtaining the set of scheduling schemes, the scheduler performs performance evaluation on these schemes. The evaluation metrics include the occupancy rate of the memory and computing unit and the load of synaptic connections. For example, the occupancy rate of the memory and computing unit is expressed as:

[0097] ;

[0098] Among them, represents the workload of the neuron core , represents the total computing power of the neuron core . By calculating these performance metrics, the scheduling schemes are sorted to obtain the optimal resource allocation strategy.

[0099] In a specific embodiment, the process of executing step S400 may specifically include the following steps:

[0100] Perform sequence encoding on the resource allocation strategy, convert the scheduling requests of multiple concurrent tasks into an input sequence based on the Seq2Seq encoder, and obtain a task time series vector;

[0101] Input the task time series vector into the encoder of the recurrent neural network, construct the data flow dependency relationship between tasks through time series feature extraction to obtain a time series feature matrix, and perform hierarchical clustering on the time series feature matrix. Divide the task group based on the data flow direction and transmission bandwidth requirements of adjacent tasks to obtain a hierarchical dependency tree;

[0102] Construct an attention weight matrix based on the hierarchical dependency tree, calculate the task priority scores through the attention mechanism in the decoder of the recurrent neural network to obtain a scheduling priority sequence, and perform neuron core allocation on the scheduling priority sequence. Perform resource matching based on the synaptic connection state and the occupancy rate of the memory and computing unit of each core to obtain a core allocation scheme;

[0103] Input the core allocation scheme into the timing verification module of the task, verify the rationality of the task execution timing through the decoder of the recurrent neural network, obtain the execution timing diagram, and perform conflict detection on the execution timing diagram. Conduct resource competition analysis based on the neuron activation threshold and synaptic weight distribution to obtain a conflict resolution strategy;

[0104] Adjust the core allocation scheme based on the conflict resolution strategy, and generate a mapping scheme between neuron cores and computing tasks through task rearrangement and resource reallocation.

[0105] Specifically, perform sequence encoding on the resource allocation strategy, and convert the scheduling requests of multiple concurrent tasks into an input sequence based on the Seq2Seq encoder. Convert the scheduling requests of multiple concurrent tasks into an input sequence for processing by a deep learning model. The task scheduling requests in the resource allocation strategy are represented as a set of task collections. Assume these tasks are , where each task has multiple attributes, such as task arrival time, required computing resources, data flow direction, etc. To convert this task information into a form that the model can process, use the Seq2Seq encoder. The role of the Seq2Seq encoder is to map the information sequence of these tasks to a fixed-length task timing vector , which can contain the dependency relationships between tasks and the timing information of the scheduling requests. The Seq25eq encoder embeds the attributes of the tasks into a high-dimensional space, and then performs time series modeling through a recurrent neural network (RNN) to output the task timing vector:

[0106] ;

[0107] where, represents the feature of task , and the Seq2Seq model converts these tasks into a task timing vector through the encoder. Input the task timing vector into the encoder of the recurrent neural network (RNN). The encoder of the RNN extracts the dependency relationships between tasks by capturing the timing features between tasks. These dependency relationships between tasks are represented by calculating the data flow dependencies between tasks. For example, assume that the input of task depends on the output of task , and represent the data dependency strength between task and task through . The RNN encoder extracts these dependency relationships and forms a timing feature matrix , this matrix represents the dependency strength and sequential relationship between tasks. The hierarchical clustering algorithm is used to analyze the time series feature matrix. Hierarchical clustering groups tasks according to the dependency relationship and data flow direction between tasks. Assume that the transmission bandwidth requirement between tasks is , then the hierarchical clustering algorithm divides tasks into different task groups according to and and forms a hierarchical dependency tree , which represents the dependency structure between tasks. An attention weight matrix is constructed based on the hierarchical dependency tree to optimize task scheduling. The attention mechanism helps the model focus on more important tasks during scheduling by assigning weights to tasks. For example, assume that the priority weight of task in the hierarchical dependency tree is , and the priority of the task is calculated through the attention weight matrix:

[0108] ;

[0109] where is the task priority sequence calculated based on the hierarchical dependency tree and the attention mechanism. The priority sequence reflects the importance and urgency of tasks during the calculation process. Neuron core allocation is performed on the scheduling priority sequence. The resource requirements and status of each neuron core, including the synaptic connection status and the occupancy rate of the memory and computing unit, need to be optimized through the task scheduling strategy. For example, the synaptic connection status of the neuron core is evaluated by the workload and bandwidth of the synapses, and the occupancy rate of the memory and computing unit represents the current load of the neuron core. If the computing requirements of task match the resources of the core well, the model assigns it to this core. The core allocation scheme is obtained in this way, including the neuron core to which each task should be assigned and the specific resource usage of this core. To ensure the rationality of resource allocation, the core allocation scheme is input into the time series verification module for verification. The time series verification module ensures that the execution order of tasks conforms to the constraint conditions of system resources by analyzing the temporal rationality of task execution. By inputting the core allocation scheme into the decoder of the recurrent neural network, the model verifies the temporal graph of task execution to ensure that task scheduling does not generate conflicts. For example, if tasks and overlap in time and their resource requirements conflict with each other, the time series verification module identifies this problem through the conflict detection mechanism. The conflict detection algorithm calculates resource competition based on the neuron activation threshold and synaptic weight distribution through the following formula:

[0110] ;

[0111] If a conflict exists, a conflict resolution strategy is designed based on the neuron activation threshold and the synaptic weight distribution. The conflict resolution strategy includes task rearrangement, resource reallocation, or adjustment of the task execution order. If a task and task have a resource conflict, the conflict is eliminated by adjusting their execution order or migrating task to other idle cores. After conflict elimination and resource reallocation, the core allocation scheme is adjusted, and a final mapping scheme between neuron cores and computing tasks is generated. This mapping scheme ensures that the computing resources for each task are reasonably allocated, and conflicts during task execution are effectively avoided.

[0112] In a specific embodiment, the process of executing step S500 may specifically include the following steps:

[0113] Perform connection relationship analysis on the mapping scheme, construct a connection topology structure based on the data transmission requirements of each neuron core, obtain a synaptic connection matrix, and input the synaptic connection matrix into a weight configurator. Calculate the information transmission intensity of each connection through an event-based spiking neural network to obtain the synaptic weight distribution;

[0114] Perform normalization processing on the synaptic weight distribution, perform weight scaling based on the neuron activation threshold and the occupancy rate of the memory and computing unit to obtain the normalized weight, and encode the computing task as neural pulses based on the normalized weight. Convert the task load into a timing signal through pulse frequency modulation to obtain a pulse coding sequence;

[0115] Perform event-triggered mapping on the pulse coding sequence, set the trigger threshold and refractory period parameters based on the firing characteristics of the neurons to obtain a trigger rule set, and input the trigger rule set into an asynchronous controller. Coordinate the execution timings of multiple neuron cores through an event-driven method to obtain an execution control sequence;

[0116] Perform propagation routing on the pulse signals in the execution control sequence, determine the signal transmission path and intensity based on the synaptic connection weights to obtain a signal transmission scheme, and establish a data path between neuron cores according to the signal transmission scheme. Complete the transmission and calculation of pulse signals through an asynchronous parallel mechanism to obtain the processing result of the pulse signal sequence.

[0117] Specifically, analyze the data transmission requirements of each neuron core and construct a connection topology structure. Suppose there are neuron cores numbered , and each neuron core Each has its corresponding task load and resource requirements, and these task loads include computing power, memory bandwidth, occupancy rate of memory-computation units, etc. There are complex data transmission relationships among these tasks, which are realized through synaptic connections. In order to reasonably map computing tasks and neuron cores, a synaptic connection matrix is constructed , where each element represents the connection strength between core and core . The construction of this matrix is based on the data transmission requirements between tasks. Assuming that the output of task needs to be transmitted to the input of task , the bandwidth requirement for data transmission is expressed as , so:

[0118] ;

[0119] where is a mapping function that can calculate the connection strength according to the bandwidth requirement . In this way, the synaptic connection matrix is obtained, which can reflect the data transmission requirements and connection relationships between neuron cores. The synaptic connection matrix is input into the weight configurator. The role of the weight configurator is to calculate the information transmission strength of each connection and calculate the synaptic weight distribution according to these strengths. Based on the event-driven spiking neural network model, these transmission strengths are converted into the weights of synapses. Assuming that the transmission strength of each connection is , the weight of the synapse is calculated according to the following formula:

[0120] ;

[0121] where is a scaling factor, representing the adjustment coefficient of the synaptic weight. Through this step, the synaptic weight distribution matrix is obtained, indicating the connection strength between each neuron core. Normalize the synaptic weight distribution to eliminate the influence brought by the differences in computing resources of different neuron cores, so that the synaptic weights can be adjusted within a reasonable range. Scale the weights proportionally so that the sum of all weights is 1 or within a specified range. For example, for the synaptic weight matrix , the normalized synaptic weight is calculated by the following formula:

[0122] ;

[0123] where represents the The sum of all synaptic weights of a neuron core. In this way, a normalized synaptic weight matrix is obtained . The normalized synaptic weights are used to encode the computational tasks into neural pulses. The encoding of neural pulses adopts the pulse frequency modulation method. Assume the task has a computational load of , then the load of the task is encoded by the frequency of the pulses. The higher the frequency of the pulses, the heavier the computational load of the task. The pulse frequency and the task load are related as expressed by the following formula:

[0124] ;

[0125] where is a constant representing the mapping coefficient from load to pulse frequency. In this way, the computational tasks are encoded into a sequence of pulse signals. The computational load of each task is converted into pulses of corresponding frequencies, thus generating a pulse coding sequence :

[0126] ;

[0127] where represents the pulse signal generated at the th time point. An event-triggered mapping is performed on the pulse coding sequence. The event-triggered mapping adjusts the triggering timing and conditions of the pulses based on the firing characteristics of the neurons. The firing characteristics of each neuron are determined by its activation threshold and refractory period. Assume the activation threshold of neuron is , and the refractory period is , then when the neuron core receives pulses from other cores, only when the intensity of the pulse signal exceeds the activation threshold can the neuron trigger and fire. The trigger rule set R is represented as a set of conditions that describe under what circumstances the pulse signal will be triggered. For example, if the pulse signal of a certain task exceeds the activation threshold of the neuron at time point , then the pulse signal is triggered, and the neuron can respond to new pulses again only after the refractory period . The trigger rule set is input into the asynchronous controller. The asynchronous controller coordinates the execution timings of multiple neuron cores according to the event-driven manner. Through the event-driven manner, the executions of the neuron cores are asynchronous at different time points, thus achieving efficient parallel computing. The asynchronous controller generates an execution control sequence , this sequence describes the execution timing and pulse triggering conditions of each neuron core at different time points. During the transmission of the pulse signal, the pulse signal in the execution control sequence is propagated and routed. The task of propagation routing is to determine the transmission path and transmission intensity of the pulse signal according to the synaptic connection matrix Determine the transmission path and transmission intensity of the pulse signal. For example, assume the pulse signal is transmitted from neuron core to neuron core , then the transmission intensity is expressed as:

[0128] ;

[0129] Among them, is the normalized synaptic weight, is the intensity of the pulse signal emitted from core . In this way, according to the synaptic connection weight, the transmission path and intensity of the signal are determined, and a signal transmission scheme is generated. According to the signal transmission scheme, a data path is established between neuron cores, and the transmission and calculation of the pulse signal are completed through an asynchronous parallel mechanism. The asynchronous parallel mechanism ensures that multiple neuron cores can execute calculation tasks in parallel at different times, improving the overall calculation efficiency. During the whole process, the processing result of the pulse signal sequence reflects the execution effect of the calculation task and provides a basis for further optimization.

[0130] In a specific embodiment, the process of executing step S600 may specifically include the following steps:

[0131] Perform index statistics on the processing result of the pulse signal sequence, calculate the execution efficiency based on the calculation and communication performance of 128 neuron cores in the neuromorphic computing chip, and obtain a performance sampling matrix;

[0132] Construct a performance loss function by comparing the target task completion time with the actual execution time, perform loss calculation on the performance sampling matrix to obtain a loss vector, perform gradient analysis on the loss vector, and construct a gradient update direction based on the synaptic weight distribution and neuron activation threshold to obtain a weight adjustment matrix;

[0133] Update the state of the memory and computing unit based on the weight adjustment matrix, construct an optimization target by calculating the workload and resource occupancy rate of each unit, and obtain a state update vector;

[0134] Input the state update vector into the Actor network, iteratively optimize the policy parameters of the action value function through the dynamic programming algorithm to obtain an optimized policy, perform a critic evaluation on the optimized policy, and construct a value network by calculating the value evaluation score of the state-action pair to obtain a value function update amount;

[0135] Input the value function update amount into the parameter optimizer, update the weight parameters of the deep reinforcement learning network by the stochastic gradient descent method to obtain an optimized parameter set, and update the parameters of the initial deep reinforcement learning network according to the optimized parameter set. Adjust the network structure through the backpropagation algorithm to obtain the target deep reinforcement learning network.

[0136] Specifically, perform index statistics on the processing results of the pulse signal sequence, and calculate the execution efficiency based on the computing and communication performance of 128 neuron cores in the neuromorphic computing chip. Assume that the computing performance and communication performance of the neuron cores are represented by and respectively, where represents the computing power of the th core, and represents the communication bandwidth of this core. By statistically analyzing the performance data of these 128 neuron cores, obtain the performance sampling matrix , where each element represents the joint metric of the communication performance and computing performance between core and core . The construction formula of this matrix is:

[0137] ;

[0138] where is a fusion function that can calculate the joint performance between each neuron core according to its computing power and communication ability. In this way, obtain a matrix that comprehensively reflects the computing and communication performance. Construct a performance loss function by comparing the difference between the completion time of the target task and the actual execution time. The completion time of the target task is usually the theoretically calculated optimal time, while the actual execution time is the time obtained based on the current resource configuration and task execution situation. Assume that the completion time of the target task is , and the actual execution time is . Define a performance loss function to measure the difference between the two. The form of the loss function is as follows:

[0139] ;

[0140] where and represent the theoretical completion time and the actual execution time of the target task respectively. To evaluate the magnitude of the loss, calculate the loss on the performance sampling matrix to obtain a loss vector, where each element represents the loss magnitude of each core during task execution. The calculation formula of the loss vector is:

[0141] ;

[0142] Among them, is the actual execution time of the th core. Through this step, the loss between each neuron core and the target task completion time is calculated, providing a basis for gradient calculation in the subsequent optimization process. Perform gradient analysis on the loss vector to determine how to update the synaptic connection weights and the state of the neuron core. Assume that the synaptic weight matrix is , and the neuron activation threshold is . The gradient of the loss vector is obtained by calculating the synaptic weight distribution and the activation threshold. Assume that the relationship between the loss vector and the synaptic weight matrix and the activation threshold is expressed as:

[0143] ;

[0144] Among them, and respectively represent the gradients of the loss function with respect to the synaptic weights and the neuron activation threshold. Through gradient analysis, the weight adjustment matrix is obtained to guide the subsequent optimization. Based on the weight adjustment matrix , the state of the memory and computing unit is updated. The state update of the memory and computing unit is usually completed by calculating the workload and resource occupancy rate of each unit. Assume that the workload of each memory and computing unit is , and the occupancy rate is . Then the state update of the unit is calculated by the following formula:

[0145] ;

[0146] Among them, represents the state update of the memory and computing unit, obtained through weighted calculation of the load and resource occupancy rate. In this way, the state of the memory and computing unit is dynamically adjusted according to the gradient in the loss vector to make it more in line with the target performance requirements. Input the updated state vector into the Actor network. The Actor network iteratively optimizes the policy parameters of the action value function through a dynamic programming algorithm. Assume that the current policy parameters are , the policy parameters are optimized through a dynamic programming algorithm, with the goal of maximizing the overall performance of the system. The dynamic programming algorithm obtains the optimal policy by gradually calculating the optimal action for each state. The optimized policy gives an updated action selection distribution. The Actor network continuously adjusts the policy during each optimization process to make resource scheduling and task allocation more efficient. After the policy update, the optimized policy is input into the Critic network for evaluation. The Critic network determines the quality of the current policy by calculating the value evaluation score for each state-action pair. Suppose the current state-action pair is , and its value evaluation function is . The Critic network calculates the value of each state-action pair according to the value evaluation function:

[0147] ;

[0148] Among them, represents the immediate reward of the current state-action pair, and is the discount factor, representing the weight of future rewards. The Critic network further optimizes the policy based on the calculated value evaluation score. The value function update amount is input into the parameter optimizer, and the weight parameters of the deep reinforcement learning network are updated by the stochastic gradient descent method. The goal of the stochastic gradient descent method is to minimize the loss function, and the policy is optimized by continuously adjusting the parameters in the network. Suppose the current network weights are , and the updated weights are calculated by the following formula:

[0149] ;

[0150] Among them, represents the learning rate, and is the gradient of the loss function. Through the backpropagation algorithm, the optimizer updates the weight parameters according to the gradient to make the policy of the network closer to the optimal policy. Through this step, the policy is continuously adjusted and optimized in the deep reinforcement learning network, and finally the target deep reinforcement learning network is obtained.

[0151] Please refer to Figure 2 , Figure 2 which is the structural schematic block diagram of the distributed processing device 200 of the neuromorphic computing chip provided by the embodiment of the present application. As Figure 2 shown, the distributed processing device 200 of the neuromorphic computing chip includes:

[0152] A decomposition module 210, configured to perform feature decomposition on the computing tasks of the neuromorphic computing chip, establish a task topology graph, and perform message passing aggregation operations on the task topology graph to obtain a task encoding vector;

[0153] The acquisition module 220 is used to collect the working status of the neuron cores in the neuromorphic computing chip, and obtain a core status matrix including synaptic weight distribution, neuron activation threshold, and occupancy rate of the memory-computation unit;

[0154] The construction module 230 is used to input the task encoding vector and the core status matrix into the initial deep reinforcement learning network, construct an action value function based on the Actor-Critic architecture, and generate a resource allocation strategy;

[0155] The reconstruction module 240 is used to reconstruct the sequence of the resource allocation strategy, calculate the temporal dependence relationship between tasks, and output a mapping scheme between neuron cores and computing tasks;

[0156] The execution module 250 is used to configure the synaptic connection weights according to the mapping scheme, convert the computing tasks into a pulse signal sequence, and transmit and execute them among neuron cores through an asynchronous event triggering mechanism;

[0157] The optimization module 260 is used to collect the processing results of the pulse signal sequence, calculate the synaptic weight update gradient, optimize the parameters of the action value function of the initial deep reinforcement learning network, and obtain a target deep reinforcement learning network.

[0158] Through the collaborative cooperation of the above-mentioned various components, by introducing a graph neural network to deeply analyze task features and dependencies, combined with a deep reinforcement learning network, dynamic optimization allocation of computing resources is achieved. An asynchronous event triggering mechanism is used for parallel processing of tasks, and the data transmission overhead is reduced based on the memory-computation integrated architecture. At the same time, through an adaptive parameter optimization method, online learning and optimization of the deep reinforcement learning network are realized, enabling the system to dynamically adjust the processing strategy according to the real-time load situation. This method significantly improves the resource utilization rate and processing efficiency of the neuromorphic computing chip, reduces energy consumption, enhances the adaptability of the system in complex computing environments, and provides an efficient solution for large-scale neural network computing.

[0159] Please refer to Figure 3 , Figure 3 FIG. is a schematic structural diagram of a distributed processing device 300 of a neuromorphic computing chip provided in an embodiment of the present application. The distributed processing device 300 of the neuromorphic computing chip includes a processor 301 and a memory 302. The processor 301 and the memory 302 are connected through a device bus 303. Among them, the memory 302 may include a non-volatile storage medium and an internal memory.

[0160] The non-volatile storage medium can store a computer program. The computer program includes program instructions. When the program instructions are executed by the processor 301, the processor 301 can be enabled to execute any of the above-mentioned distributed processing methods of the neuromorphic computing chip.

[0161] The processor 301 is used to provide computing and control capabilities to support the operation of the distributed processing device 300 of the entire neuromorphic computing chip.

[0162] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor 301, the processor 301 can be enabled to execute any of the above-mentioned distributed processing methods of the neuromorphic computing chip.

[0163] Those skilled in the art can understand that Figure 3 The structure shown in [the figure] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the distributed processing device 300 of the neuromorphic computing chip involved in the solution of this application. The specific distributed processing device 300 of the neuromorphic computing chip may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0164] It should be understood that the processor 301 may be a central processing unit (CPU), and this processor 301 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.

[0165] It should be noted that those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the above-described distributed processing device 300 of the neuromorphic computing chip can refer to the corresponding process of the foregoing distributed processing method of the neuromorphic computing chip, and will not be elaborated here.

[0166] This application embodiment also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program is executed by one or more processors, the one or more processors are enabled to implement the distributed processing method of the neuromorphic computing chip provided by this application embodiment.

[0167] Among them, the computer-readable storage medium may be an internal storage unit of the distributed processing device 300 of the neuromorphic computing chip in the foregoing embodiments, such as the hard disk or memory of the distributed processing device 300 of the neuromorphic computing chip. The computer-readable storage medium may also be an external storage device of the distributed processing device 300 of the neuromorphic computing chip, such as a plug-in hard disk equipped with the distributed processing device 300 of the neuromorphic computing chip, a SmartMedia Card (SMC), a Secure Digital (SD) card, a Flash Card, etc.

[0168] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described system, system, and unit can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0169] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0170] As described above, the above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of this application.

Claims

1. A distributed processing method for a neuromorphic computing chip, characterized in that: include: The computing tasks of the neuromorphic computing chip are feature decomposed, a task topology graph is established, and a message passing aggregation operation is performed on the task topology graph to obtain a task encoding vector; specifically, the method comprises: collecting timing features of the computing tasks of the neuromorphic computing chip, extracting a task feature set including task arrival time, required number of neurons, number of synaptic connections, and data flow direction, performing data dependency analysis on the task feature set, and calculating task dependencies based on the data transmission relationship between tasks; constructing a directed graph structure based on the task dependencies, mapping the task feature set to graph node attributes, and mapping the dependency strength matrix to edge weights to obtain a task topology graph; Perform multi-head attention calculation on the nodes in the task topology graph to obtain node attention features, and input the node attention features into the graph convolution layer, perform weighted sum operation on the features of adjacent nodes through the message aggregation function to obtain the node hidden layer representation; calculate the global features of the task topology graph according to the node hidden layer representation, perform pooling operation and dimension conversion on the node attention features to obtain the graph structure encoding, and input the graph structure encoding into the fully connected neural network, map the node attention features to the target dimensional space through nonlinear transformation, and obtain a compressed feature vector; perform feature fusion on the compressed feature vector and the node hidden layer representation, and obtain the task encoding vector through residual connection; Collect the working status of the neuron core in the neuromorphic computing chip and obtain the core state matrix including synaptic weight distribution, neuron activation threshold, and storage and computing unit occupancy rate; Input the task encoding vector and the core state matrix into an initial deep reinforcement learning network, construct an action value function based on an Actor-Critic architecture, and generate a resource allocation strategy; Reconstructing the resource allocation strategy in sequence, calculating the temporal dependencies between tasks, and outputting a mapping scheme between neuron cores and computing tasks; Configuring synaptic connection weights according to the mapping scheme, converting the computing task into a pulse signal sequence, and transmitting and executing between the neuron cores through an asynchronous event triggering mechanism; The processing results of the pulse signal sequence are collected, the synaptic weight update gradient is calculated, and the action value function of the initial deep reinforcement learning network is optimized to obtain the target deep reinforcement learning network.

2. The distributed processing method of the neuromorphic computing chip according to claim 1, characterized in that: The working state of the neuron core in the neuromorphic computing chip is collected to obtain a core state matrix including synaptic weight distribution, neuron activation threshold, and storage and computing unit occupancy rate, including: The synaptic connections of the neuron cores in the neuromorphic computing chip are scanned in parallel, and the synaptic weight, connection state and transmission bandwidth of each core are extracted to obtain the synaptic connection characteristics; Performing weight density calculation on the synaptic connection features to obtain synaptic weight distribution, and sampling the neuron state of each neuron core to obtain a neuron activation threshold; Based on the neuron activation threshold execution state mapping, the pulse triggering condition and the discharge frequency of the neuron core are quantitatively calculated to obtain a threshold distribution sequence, and the threshold distribution sequence and the storage and calculation unit are subjected to workload analysis to obtain a load state value; The load state value is input into an adaptive sampler, and the real-time occupancy of the storage and computing unit is captured by dynamically adjusting the sampling frequency to obtain an occupancy sequence, and the occupancy sequence and the synaptic weight distribution are feature integrated to construct a state feature matrix through multi-dimensional matrix operations; Data fusion is performed based on the state feature matrix and the threshold distribution sequence, and multi-source state information is merged through tensor operations to obtain a core state matrix, where the multi-source state information includes the state feature matrix and the threshold distribution sequence.

3. The distributed processing method of the neuromorphic computing chip according to claim 2, characterized in that: The step of inputting the task encoding vector and the core state matrix into an initial deep reinforcement learning network, constructing an action value function based on an Actor-Critic architecture, and generating a resource allocation strategy includes: Dimensionally aligning the task encoding vector and the core state matrix, constructing a 128-dimensional input feature space through feature recombination, and obtaining a state input vector; Input the state input vector into the Actor network, perform action space modeling based on the computing power of neurons and synapses, obtain an initial action distribution, perform feasibility constraints on the initial action distribution, perform resource allocation verification based on 16PB / s memory bandwidth and 3.5PB / s inter-core communication bandwidth, and obtain a valid action set; Input the state input vector and the valid action set into the Critic network, calculate the state value based on 380 trillion synaptic operations and 240 trillion neuron operations, and obtain a value evaluation function; Performing a temporal difference operation on the value evaluation function, constructing an instant reward by calculating the 192KB cache utilization of each neuron core, obtaining a TD error, and performing a gradient update on the policy distribution of the Actor network based on the TD error, optimizing the action selection probability in an asynchronous event-driven manner, and obtaining an updated policy; The update strategy is input into the resource scheduler, action mapping is performed based on the configuration space of 4096 neuron states to obtain a scheduling scheme set, and performance evaluation is performed on the scheduling scheme set. The schemes are sorted by calculating the storage and computing unit occupancy rate and the synaptic connection load to obtain the resource allocation strategy.

4. The distributed processing method of the neuromorphic computing chip according to claim 3, characterized in that: The method of sequentially reconstructing the resource allocation strategy, calculating the temporal dependencies between tasks, and outputting a mapping scheme between neuron cores and computing tasks includes: Sequentially encode the resource allocation strategy, and convert the scheduling requests of multiple concurrent tasks into an input sequence based on a Seq2Seq encoder to obtain a task timing vector; Inputting the task timing vector into the encoder of the recurrent neural network, constructing the data flow dependency relationship between the tasks through timing feature extraction, obtaining a timing feature matrix, performing hierarchical clustering on the timing feature matrix, dividing the task groups based on the data flow direction and transmission bandwidth requirements of adjacent tasks, and obtaining a hierarchical dependency tree; An attention weight matrix is ​​constructed based on the hierarchical dependency tree, and the task priority scores are calculated through the attention mechanism in the decoder of the recurrent neural network to obtain a scheduling priority sequence, and the scheduling priority sequence is allocated to neuron cores, and resource matching is performed based on the synaptic connection state and storage and computing unit occupancy rate of each core to obtain a core allocation plan; The core allocation scheme is input into the timing verification module, and the rationality of the task execution timing is verified by the decoder of the recurrent neural network to obtain an execution timing diagram, and conflict detection is performed on the execution timing diagram, and resource competition analysis is performed based on the neuron activation threshold and synaptic weight distribution to obtain a conflict elimination strategy; The core allocation scheme is adjusted based on the conflict elimination strategy, and a mapping scheme between neuron cores and computing tasks is generated through task rearrangement and resource reallocation.

5. The distributed processing method of the neuromorphic computing chip according to claim 4, characterized in that: The step of configuring synaptic connection weights according to the mapping scheme, converting the computing task into a pulse signal sequence, and transmitting and executing the sequence between the neuron cores through an asynchronous event triggering mechanism includes: Performing a connection relationship analysis on the mapping scheme, constructing a connection topology structure based on the data transmission requirements of each neuron core, obtaining a synaptic connection matrix, and inputting the synaptic connection matrix into a weight configurator, calculating the information transmission strength of each connection through an event-based spiking neural network, and obtaining a synaptic weight distribution; Performing normalization processing on the synaptic weight distribution, performing weight scaling based on the neuron activation threshold and the storage and computing unit occupancy rate to obtain a standardized weight, encoding the computing task into a neural pulse based on the standardized weight, converting the task load into a timing signal through pulse frequency modulation, and obtaining a pulse signal sequence; Performing event trigger mapping on the pulse signal sequence, setting a trigger threshold and a refractory period parameter based on the discharge characteristics of the neuron to obtain a trigger rule set, and inputting the trigger rule set into an asynchronous controller to coordinate the execution timing of multiple neuron cores in an event-driven manner to obtain an execution control sequence; The pulse signals in the execution control sequence are propagated and routed, the signal transmission path and strength are determined based on the synaptic connection weight, a signal transmission scheme is obtained, and a data path is established between neuron cores according to the signal transmission scheme. The transmission and calculation of the pulse signals are completed through an asynchronous event triggering mechanism to obtain the processing result of the pulse signal sequence.

6. The distributed processing method of the neuromorphic computing chip according to claim 5, characterized in that: The processing result of collecting the pulse signal sequence, calculating the synaptic weight update gradient, and optimizing the parameters of the action value function of the initial deep reinforcement learning network to obtain the target deep reinforcement learning network include: Performing index statistics on the processing results of the pulse signal sequence, calculating the execution efficiency based on the computing and communication performance of 128 neuron cores in the neuromorphic computing chip, and obtaining a performance sampling matrix; A performance loss function is constructed by comparing the target task completion time with the actual execution time, a loss calculation is performed on the performance sampling matrix to obtain a loss vector, a gradient analysis is performed on the loss vector, a gradient update direction is constructed based on the synaptic weight distribution and the neuron activation threshold, and a weight adjustment matrix is ​​obtained; Based on the weight adjustment matrix, the state of the storage and computing unit is updated, and the optimization target is constructed by calculating the workload and resource occupancy rate of each unit to obtain a state update vector; The state update vector is input into the Actor network, the policy parameters of the action value function are iteratively optimized by a dynamic programming algorithm to obtain an optimized policy, and the optimized policy is evaluated by a critic, and a value network is constructed by calculating the value evaluation scores of the state-action pairs to obtain the value function update amount; The updated value function is input into the parameter optimizer, and the weight parameters of the deep reinforcement learning network are updated by the stochastic gradient descent method to obtain an optimized parameter set. The parameters of the initial deep reinforcement learning network are updated according to the optimized parameter set, and the network structure is adjusted by the back propagation algorithm to obtain the target deep reinforcement learning network.

7. A distributed processing device for a neuromorphic computing chip, characterized in that: A distributed processing method for executing a neuromorphic computing chip according to any one of claims 1 to 6, comprising: A decomposition module is used to perform feature decomposition on the computing tasks of the neuromorphic computing chip, establish a task topology graph, and perform message passing aggregation operation on the task topology graph to obtain a task encoding vector; The acquisition module is used to collect the working status of the neuron core in the neuromorphic computing chip and obtain the core state matrix including synaptic weight distribution, neuron activation threshold, and storage and computing unit occupancy rate; A construction module, used for inputting the task encoding vector and the core state matrix into an initial deep reinforcement learning network, constructing an action value function based on an Actor-Critic architecture, and generating a resource allocation strategy; A reconstruction module, used to sequentially reconstruct the resource allocation strategy, calculate the temporal dependencies between tasks, and output a mapping scheme between neuron cores and computing tasks; An execution module, configured to configure synaptic connection weights according to the mapping scheme, convert the computing task into a pulse signal sequence, and transmit and execute the sequence between the neuron cores through an asynchronous event triggering mechanism; The optimization module is used to collect the processing results of the pulse signal sequence, calculate the synaptic weight update gradient, and optimize the parameters of the action value function of the initial deep reinforcement learning network to obtain the target deep reinforcement learning network.

8. A distributed processing device for a neuromorphic computing chip, characterized in that: The distributed processing device of the neuromorphic computing chip includes: a memory and at least one processor, wherein instructions are stored in the memory; The at least one processor calls the instruction in the memory so that the distributed processing device of the neuromorphic computing chip executes the distributed processing method of the neuromorphic computing chip as described in any one of claims 1-6.

9. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instruction is executed by the processor, the distributed processing method of the neuromorphic computing chip as described in any one of claims 1-6 is implemented.

Citation Information

Patent Citations

  • Deep learning optimization method for edge computing server

    CN118520936A

  • Load balancing method of low-power AI processor, chip and storage medium

    CN119271418A