A multi-tasking neural network processor based on systolic array

Through a multi-task neural network processor based on pulsating arrays, combined with an adaptive task allocation algorithm and shift addition multiplier, the problems of low resource utilization, high power consumption and large delay in the existing technology are solved, and efficient and flexible multi-task processing and resource optimization are achieved, which is suitable for the computing needs of edge devices and mobile terminals.

CN119783743BActive Publication Date: 2025-05-13UESTC (SHENZHEN) ADVANCED RES INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510277701.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-05-13
Estimated Expiration
2045-03-10

AI Technical Summary

Technical Problem

Existing neural network processors have great limitations in general computing, multitasking and resource optimization, especially on terminal devices, which are characterized by low resource utilization, low flexibility, high power consumption and high latency, making it difficult to meet diversified computing needs.

Method used

The multi-task neural network processor based on pulsating array is adopted, combined with the adaptive task allocation algorithm, dynamically adjust task weights and priorities, and optimize resource allocation and scheduling through shift addition multiplier and data flow-driven multitasking architecture design to achieve efficient parallel processing.

Benefits of technology

It significantly improves resource utilization, reduces hardware resource consumption and system latency, improves computing performance and flexibility, and is suitable for deep learning and artificial intelligence computing in edge devices and mobile terminals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119783743B_ABST
    Figure CN119783743B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-task neural network processor based on a systolic array, which relates to the technical field of neural network processors and solves the technical problem that existing neural network processors have great limitations when performing general computing, multi-task processing and resource optimization. The neural network processor adopts a systolic array design, and dynamically adjusts and allocates multi-task weights through an adaptive task allocation algorithm; the systolic array controls the computing status of different tasks and gives priority to computing tasks with higher priorities; the systolic array transmits an input feature data and a weight data to each PE computing unit during each clock cycle, and gradually increases the number of input feature data and weight data sent to each PE computing unit in subsequent clock cycles. The present invention realizes refined control of PE computing units and efficient scheduling of multiple tasks through the combination of an adaptive task allocation algorithm and a systolic array architecture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of neural network processors, and in particular to a multi-task neural network processor based on a systolic array. Background Art

[0002] Current neural network processor technology is widely used to accelerate neural network reasoning and training tasks, among which representative technologies include Google's Tensor Processing Unit (TPU) and traditional systolic array NPU (Neural Processing Unit) architecture. Google's TPU uses a systolic array architecture to tightly couple a large number of processing elements (PE) to achieve efficient matrix operations and convolution operations. This architecture exhibits extremely high throughput and computing power when executing large neural network models, especially in processing deep learning tasks. However, the design of TPU is too customized, mainly for specific types of deep learning tasks, with poor flexibility and high power consumption. When facing a wider range of application scenarios such as edge devices or mobile terminals, TPU cannot adapt effectively. Its high manufacturing cost and high power consumption make it difficult to promote in terminal devices with limited resources. This overly specialized design makes TPU have great limitations in multi-tasking and general tasks, and it is difficult to meet the diverse needs of terminal devices.

[0003] Traditional systolic array neural network processors also demonstrate efficient performance in accelerating specific tasks. It transmits and processes data in parallel through a two-dimensional processing unit network, which can significantly accelerate numerical computing tasks such as matrix operations. However, the traditional systolic array architecture faces a large bottleneck when processing multiple tasks. First, low resource utilization is one of the core problems of existing neural network processors. Due to the static resource allocation method of the systolic array, the working state of the processing unit cannot be dynamically adjusted according to the priority of different tasks, resulting in a large amount of computing resources being wasted during task switching and scheduling, and failing to fully utilize the computing potential of the hardware. For terminal devices, this inefficient resource utilization directly affects the battery life and real-time response performance of the device. In addition, the existing neural network processor lacks an effective multi-task scheduling mechanism, which makes it difficult to flexibly adjust task priorities when processing multiple tasks. Low-priority tasks are often delayed due to low scheduling efficiency, resulting in a significant decline in overall system performance, especially in scenarios where terminal devices face high concurrent tasks. This lack of flexible scheduling architecture is prone to become a system bottleneck and is difficult to meet the complex computing requirements of terminal scenarios.

[0004] In addition, traditional neural network processors face the problems of high latency and high power consumption in hardware design, which is particularly prominent on terminal devices. The multiplier modules in the systolic array occupy a large amount of hardware resources, which increases the complexity of hardware design and increases power consumption. At the same time, when the multiplier performs multi-task calculations, the excessive number of partial products significantly increases the calculation delay. Especially when the terminal device requires high-real-time task processing, this delay becomes a system performance bottleneck. Existing neural network processors usually adopt a fixed data flow mode, lack fine-grained control of different tasks, and cannot dynamically adjust resource allocation according to task requirements, further exacerbating resource waste and reduced system response speed.

[0005] In the process of implementing the present invention, the inventors found that there are at least the following problems in the prior art:

[0006] Existing neural network processors have great limitations when performing general computing, multi-tasking and resource optimization. There is an urgent need for a neural network processor with high versatility, high resource utilization, easy multi-tasking scheduling of terminal devices, low power consumption and low latency to meet the dual needs of terminal devices for computing performance and energy efficiency. Summary of the invention

[0007] The purpose of the present invention is to provide a multi-task neural network processor based on a systolic array to solve the existing neural network processors in the prior art that have great limitations in general computing, multi-tasking and resource optimization, and urgently need a neural network processor with high versatility, high resource utilization, convenient multi-task scheduling of terminal devices, low power consumption and low latency to meet the technical problem of the dual needs of terminal devices for computing performance and energy efficiency. The many technical effects that can be produced by the preferred technical solutions among the many technical solutions provided by the present invention are described in detail below.

[0008] To achieve the above object, the present invention provides the following technical solutions:

[0009] The present invention provides a multi-task neural network processor based on a systolic array. The multi-task neural network processor adopts a systolic array design and dynamically adjusts and allocates multi-task weights through an adaptive task allocation algorithm; the systolic array processes multiple tasks in parallel and gives priority to computing tasks with higher priorities by controlling the computing states of different tasks; the systolic array transmits an input feature data and a weight data to each PE computing unit during each clock cycle, and gradually increases the number of input feature data and weight data sent to each PE computing unit in subsequent clock cycles.

[0010] Preferably, the multi-task neural network processor includes a multiplier module, a PE module, and a scheduler module; the multiplier module adopts a shift-add multiplier; the PE module adopts a data flow-driven multi-task architecture design, including multiple PE operation units, for processing data transmitted by the previous PE operation unit or the left PE operation unit; the scheduler module performs single-task scheduling and multi-task scheduling, the single-task scheduling optimizes the execution of a single task in the PE module, and the multi-task scheduling manages multiple tasks, allocates computing resources and dynamically adjusts task priorities.

[0011] Preferably, the multiplier module performs multiplication calculation by adding term by term, and performs calculation according to whether each bit of the multiplier is 1, and if it is 1, the multiplicand is shifted and added.

[0012] Preferably, the multiplier module performs a multiplication-addition operation in each clock cycle and simultaneously transmits the calculation result to an adjacent PE operation unit.

[0013] Preferably, it further comprises a global enable terminal START, and when the global enable terminal START signal is activated, the enable signal of the PE module can take effect.

[0014] Preferably, the single-task scheduling can dynamically adjust the working state of the PE computing unit of a specific task to start or stop the PE computing unit; the multi-task scheduling automatically manages the task queue through an adaptive allocation algorithm to enable high-priority tasks to obtain computing resources first.

[0015] Preferably, the adaptive task allocation algorithm comprises the following steps: S100: the scheduling scheme corresponding to each individual i is represented by a weight array as ,in and They are floating point arrays corresponding to high-priority tasks and low-priority tasks, representing individual The two sets of task decision weights are generated to size Population Get the initialized population; S200: based on the logical resource consumption XLRS, resident set size , use the support vector machine model to train the task load prediction module, and calculate the fitness value of each individual through the fitness function; S300: use the fitness change detection calculation formula Determine whether the algorithm convergence criteria are met, where Represents the fitness value of the optimal solution calculated by the fitness function in the current iteration, It represents the fitness value of the optimal solution calculated by the fitness function in the previous iteration. To set the threshold, if the condition is met, the current optimal solution is output, otherwise S400 is executed; S400: k individuals are randomly selected from the population N, and the Bayesian optimization mechanism is used to select the individuals with the highest fitness. individuals; S500: select a crossover point cp for the two weight arrays of the W individuals with the highest fitness to perform a crossover operation and generate a new weight allocation scheme; S600: determine whether to perform weight mutation for each individual based on the mutation probability Pm, and use the formula Perform mutation operation, where and denote the indices of individuals and genes, respectively. For the population Individuals, For this individual genes, is the gene value before mutation, mutate represents the specific implementation function of the mutation operation, and finally a new generation of population is obtained to enter the next round of iteration.

[0016] Preferably, the step S100 specifically includes: S110: for the elements in the first set of weights, their values ​​are given by the formula Randomly generated, where It means to generate a random floating point number in the interval [0,1]. is the number of calculation tasks; S120: the elements in the second group of weights are normalized by satisfying the constraint that the sum of the weights is 1; S130: the above process is repeated until an initialization population of size N is generated.

[0017] Preferably, the fitness function is defined as follows: ,in, Indicates the first The fitness value of each individual, and Respectively represent The predicted values ​​of individual logical resource consumption and resident set size, and is a weight factor used to adjust the relative importance of logic resource consumption and resident set size in the fitness function calculation.

[0018] Implementing one of the above technical solutions of the present invention has the following advantages or beneficial effects:

[0019] The present invention realizes refined control of PE computing units and efficient scheduling of multiple tasks, as well as adaptive optimization of hardware resources in a multi-tasking environment, through the combination of an adaptive task allocation algorithm and a systolic array architecture, significantly improving the computing performance and flexibility of the neural network processor on the terminal device. Compared with the prior art, the present invention can significantly reduce hardware resource consumption, reduce system latency, and effectively improve the utilization rate and overall computing efficiency of the PE computing unit. Therefore, the present invention is suitable for the field of deep learning and artificial intelligence computing in edge devices and mobile terminals, and provides a more flexible and efficient solution for computing-intensive tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. It is obvious that the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work. In the drawings:

[0021] Figure 1 It is a schematic diagram of a systolic array structure of a multi-task neural network processor based on a systolic array according to an embodiment of the present invention;

[0022] Figure 2 It is a schematic diagram of the working principle of a scheduler module of a multi-task neural network processor based on a systolic array according to an embodiment of the present invention;

[0023] Figure 3 is a flow chart of an adaptive task allocation algorithm according to Embodiment 2 of the present invention;

[0024] Figure 4 It is a flowchart of step S100 in an adaptive task allocation algorithm according to the second embodiment of the present invention. DETAILED DESCRIPTION

[0025] In order to make the purpose, technical solutions and advantages of the present invention clearer, the various exemplary embodiments to be described below will refer to the corresponding drawings, which constitute a part of the exemplary embodiments, wherein various exemplary embodiments that may be used to implement the present invention are described. Unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation methods described in the following exemplary embodiments do not represent all implementation methods consistent with the present disclosure. It should be understood that they are only examples of processes, methods, devices, etc. that are consistent with some aspects of the present disclosure as detailed in the attached claims, and other embodiments may also be used, or the embodiments listed herein may be modified in structure and function without departing from the scope and essence of the present invention.

[0026] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", etc. indicate the orientation or positional relationship based on the drawings, which is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the elements referred to must have a specific orientation, be constructed and operated in a specific orientation. The terms "first", "second", etc. are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. The term "multiple" means two or more. The terms "connected" and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, an integral connection, a mechanical connection, an electrical connection, a communication connection, a direct connection, an indirect connection through an intermediate medium, and can be the internal connection of two elements or the interaction relationship between two elements. The term "and / or" includes any and all combinations of one or more related listed items. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to the specific circumstances.

[0027] In order to illustrate the technical solution of the present invention, a specific embodiment is used below for description, and only the parts related to the embodiment of the present invention are shown.

[0028] Embodiment 1:

[0029] The present invention provides a multi-task neural network processor based on a systolic array. The multi-task neural network processor adopts a systolic array design and dynamically adjusts and allocates multi-task weights through an adaptive task allocation algorithm. In a parallel computer architecture, a systolic array is an architecture applied to an on-chip multiprocessor. The architecture of the present invention is optimized based on a traditional systolic array. The systolic array is a homogeneous network composed of tightly coupled identical data processing units called units or nodes, and has the characteristics of parallelism, regularity and local communication. For example, numerical linear algebra algorithms, matrix convolution calculations, etc. can all be implemented using systolic arrays. The adaptive task allocation algorithm can automatically adjust task weights according to the characteristics of different tasks to achieve the best scheduling combination. This dynamic task allocation strategy not only improves the efficiency of multi-task processing, but also maximizes the utilization of PE computing units and reduces the consumption of hardware resources. Through weight-based scheduling, it can be ensured that all tasks have the opportunity to reasonably obtain resources according to their priority, thereby reducing the waiting time of tasks and avoiding the problem that low-priority tasks cannot be processed for a long time. The systolic array processes multiple tasks in parallel and gives priority to the computing tasks with higher priority by controlling the computing status of different tasks; thus, the working status of PE can be dynamically controlled according to the priority and real-time requirements of the tasks, thereby optimizing resource allocation. The systolic array transmits an input feature data and a weight data to each PE computing unit during each clock cycle, and gradually increases the amount of input data transmitted to the PE computing unit in subsequent clock cycles to ensure data processing continuity and improve parallel processing capabilities, such as Figure 1 As shown ( Figure 1 Clk represents the clock cycle, and n clock cycles are shared). The design of the data flow ensures the efficient use of each PE computing unit, so that the entire architecture can instantly output the convolution results of the matrix rows and columns after completing the data processing, thereby improving the data processing speed and optimizing resource allocation, ensuring the efficiency and accuracy of the processing. The present invention realizes the refined control of the PE computing unit and the efficient scheduling of multiple tasks, as well as the adaptive optimization of hardware resources in a multi-tasking environment, which significantly improves the computing performance and flexibility of the neural network processor in complex multi-tasking scenarios. Compared with the prior art, it significantly reduces the consumption of hardware resources, reduces system latency, and effectively improves the utilization rate and overall computing efficiency of the PE computing unit. Therefore, it is suitable for deep learning and artificial intelligence computing fields that require efficient processing of parallel tasks, and provides a more flexible and efficient solution for the needs of computing-intensive tasks.

[0030] As an optional implementation, the multitasking neural network processor includes a multiplier module, a PE module, and a scheduler module. The multiplier module uses a shift-add multiplier. In a traditional systolic array, the multiplier is a basic component of the PE module, and a parallel multiplier is commonly used. Starting from the low bit of the multiplier, one bit is taken each time to be multiplied with the multiplicand, and the product is temporarily stored as a partial product. After all the valid bits of the multiplier are multiplied, the weights of all partial products corresponding to the multiplier digits are staggered and accumulated to obtain the final product. Although this method has a fast operating speed, it leads to higher resource consumption and delay when the bit width increases. The shift-add multiplier in the present invention is implemented by adding term by term, thereby effectively reducing the occupation of hardware resources.

[0031] The PE module adopts the operating system data flow multi-tasking architecture design, including multiple PE computing units, which are used to process data transmitted by the previous PE computing unit or the PE computing unit on the left. In the systolic array, each PE computing unit is directly connected to the on-chip cache to store the current data and perform calculations. The data is transmitted horizontally to the right and vertically downward to the adjacent PE computing unit. This two-dimensional topology can efficiently implement time constraints and data flow control. The SIMD architecture (Single Instruction Multiple Data) in the PE module allows the same operation to be performed on multiple data points, thereby realizing multi-tasking. In the present invention, by introducing a dynamic scheduling mechanism, the PE module can execute different operations in parallel and selectively schedule tasks with higher priority, reducing resource consumption and enhancing SIMD functions. The advantages of this module design are: realizing multi-task data processing and parallel computing, reducing hardware costs; the scheduler flexibly allocates resources and gives priority to high-priority tasks; real-time control of the working status of the PE module, improving PE utilization, such as Figure 2 shown.

[0032] Take two tasks as an example, each task contains nine data calculations, and the priority of the task is represented by the weight. For example, the weight of the first data calculation in the first task is 0.8, which is higher than the weight of the second task, so the convolution calculation of the first task is performed first; in the second data calculation, the second task has a higher weight and is processed first. When some unimportant data appears (such as erroneous data in image processing), the weight can be set to 0, so that the PE module stops calculating, saving resources and improving the accuracy of data processing. This priority model is suitable for scenarios that require fast processing and optimized resource allocation, such as: prioritizing motion frames in real-time video processing, real-time processing of autonomous driving sensor data, and other application scenarios.

[0033] In the hardware design of the PE module, although the multi-task PE module can efficiently complete the priority data processing, the weight-based scheduling requires manual intervention. To this end, the present invention also introduces an adaptive task allocation algorithm to automatically analyze the weight of each task, and realize the fully automatic allocation and priority data processing functions of multiple tasks. The scheduler module performs single-task and multi-task scheduling. Single-task scheduling optimizes the execution efficiency of a single task on the PE module, and multi-task scheduling manages multiple tasks, allocates resources among multiple tasks, and dynamically adjusts priorities. The flexible mechanism of the scheduler ensures that high-priority tasks can obtain processing resources first, reduce the waiting time of low-priority tasks, and significantly improve the overall performance of the system.

[0034] As an optional implementation, the multiplier module implements multiplication calculation by term-by-term addition, and calculates according to whether each bit of the multiplier is 1. If it is 1, the multiplicand is shifted and added. For example, the shift addition design of 3-bit wide data requires at most 2 adders and 3 shift registers, which effectively reduces the hardware resource cost by reducing the number of partial products and the number of adders. The specific implementation is as follows: convert the multiplicand and the multiplier into binary, check each bit of the multiplier from low to high bit to see if it is 1, if it is 1, shift the multiplicand to the left by the corresponding number of bits and add it to the partial product; repeat this step until all binary bits are processed. The multiplier module performs multiplication and addition operations in each clock cycle and passes the calculation results to the adjacent PE operation unit. This design can reduce addition operations, reduce power consumption and improve throughput.

[0035] As an optional implementation, Figure 4 As shown, it also includes the global enable terminal START. When the global enable terminal START signal is activated, the enable signal of the PE module can take effect, ensuring that the system saves resources when not executing tasks. The single-task scheduler dynamically adjusts the working state of the PE module for specific tasks to optimize the efficiency of task execution. Through accurate enable control strategies, the scheduler manages the start and stop of the PE module in real time to ensure that all available PE modules run at the highest efficiency when a single task is executed. The single-task scheduler can not only flexibly allocate computing resources according to the real-time needs of the task, but also save energy when idle, thereby improving the energy efficiency and performance of the system.

[0036] As an optional implementation, single-task scheduling can dynamically adjust the working state of the PE computing unit of a specific task to start or stop the PE computing unit; multi-task scheduling automatically manages the task queue through an adaptive allocation algorithm so that high-priority tasks can obtain computing resources first. Multi-task scheduling manages the task queue through an adaptive task allocation algorithm to ensure that high-priority tasks obtain resources first. The weight-based task scheduling strategy in the scheduler assigns a weight value to each task to reflect the task priority. The algorithm automatically manages the task queue to ensure that urgent tasks can obtain resources first. The scheduler dynamically allocates resources when multiple tasks are concurrent to optimize the system throughput and task completion time and prevent low-priority tasks from waiting for a long time. For example, for two tasks, if the weight of the first task is greater than that of the second task, the first task is processed first; if the weights are equal, the calculation results are output at the same time. This scheduling strategy optimizes resource allocation and improves throughput, adapts to different task requirements and priority changes, and ensures efficient operation of the system in a multi-task environment.

[0037] The embodiment is only a specific example and does not represent only one implementation mode of the present invention.

[0038] Embodiment 2:

[0039] An adaptive task allocation algorithm is implemented based on a multi-task neural network processor based on a systolic array in the first embodiment. The adaptive algorithm solves the optimization problem of multi-task scheduling by simulating the inheritance, mutation, recombination and selection mechanisms in the biological evolution process. The algorithm starts with a population of initial task scheduling schemes, each scheme (individual) represents a possible resource allocation method, and its pros and cons are evaluated by the fitness function. In each generation, the algorithm selects individuals with better fitness as parents, generates new task scheduling scheme populations through crossover (recombining parent characteristics to generate offspring) and mutation (randomly adjusting offspring characteristics) operations, and iterates repeatedly until the termination condition is met, such as reaching a predetermined number of iterations or fitness threshold. The core advantage of the adaptive task allocation algorithm lies in its adaptability, robustness and parallel processing capabilities to the problem, which enables it to effectively handle multi-objective, nonlinear and high-dimensional complex optimization problems. The present invention selects an adaptive task allocation algorithm to handle multi-task scheduling problems based on systolic arrays, mainly for the following reasons: Supporting multi-objective optimization: In the design process of a multi-task neural network processor, the adaptive algorithm can realize automatic scheduling of different task weights, optimize the efficiency of hardware resource utilization, and generate a Pareto front solution set at the same time. Decision makers can select solutions based on task priorities, optimizing execution time while improving resource utilization. Suitable for high-dimensional parameter space optimization: Multi-task neural network processors based on systolic arrays enable scheduling optimization with high-dimensional parameter space. The adaptive algorithm automates task scheduling through global optimization, avoiding local optimal problems, and is particularly suitable for complex multi-task scheduling needs. Adapt to the optimization of discrete variables: Multi-task scheduling combinations are discrete, and the adaptive task allocation algorithm shows strong adaptability to the optimization problem of discrete variables and can efficiently search the scheduling solution space. Improve the diversity of the solution space: By exploring different areas of the solution space, the algorithm provides a diverse set of feasible solutions, which helps to adaptively adjust according to task scheduling requirements and ensure scheduling flexibility and efficiency in a multi-task environment.

[0040] In addition, since there is no fixed formula to drive resource utilization such as XLRS (logic resource consumption) and RSS (resident set size) of FPGA (field programmable gate array), the present invention adopts a task load prediction model. The model combines resource allocation of different scheduling schemes to generate predictions for XLRS and RSS, thereby reducing hardware simulation time and accelerating the scheduling optimization process.

[0041] like Figure 3 As shown, the adaptive task allocation algorithm includes the following steps: S100: The scheduling scheme corresponding to each individual i is represented by a weight array as ,in and They are floating point arrays corresponding to high-priority tasks and low-priority tasks, representing individual The two groups of task decision weights are used to generate a population of size N and obtain an initialized population. S200: Based on the logical resource consumption XLRS and the resident set size , use the support vector machine model to train the task load prediction module, and calculate the fitness value of each individual through the fitness function. The goal of the fitness function is to minimize the logical resource consumption (XLRS) and the resident set size (RSS). The fitness value calculation is based on the actual resource consumption results of XLRS and RSS. The effect of the scheduling scheme is evaluated through the optimal combination of resource consumption and fitness function. Deployed to the FPGA according to the scheduling scheme, the XLRS and RSS consumption of the current scheduling combination are obtained to evaluate the pros and cons of the individual and accelerate the optimization. S300: Detect the calculation formula through fitness change Determine whether the algorithm convergence criteria are met, where Represents the fitness value of the optimal solution calculated by the fitness function in the current iteration, It represents the fitness value of the optimal solution calculated by the fitness function in the previous iteration. To set the threshold, if the condition is met, the current optimal solution is output, otherwise S400 is executed. Fitness is composed of logical resource consumption (XLRS) and resident set size (RSS), which measures the quality of the scheduling scheme in terms of resource utilization and performance optimization, and determines whether the algorithm stops iterating through the convergence criterion (threshold). At the end of each generation, compare the current optimal fitness with the optimal fitness of the previous N generations, and use the fitness change detection formula to determine whether the algorithm convergence criterion is met. If the difference between the best fitness value in the current iteration and the optimal value of the previous iteration is less than the set threshold, the algorithm is considered to have converged; otherwise, continue to iterate. Fitness change detection can effectively monitor the convergence state of the algorithm and ensure that it is terminated when the expected goal is reached, so as to ensure the quality of the solution and improve computing efficiency. S400: Randomly select k individuals from the population, and use the Bayesian optimization mechanism for selection, that is, use Bayesian optimization to screen out candidate individuals with higher fitness, and then randomly select k individuals from them for comparison, and select the one with the highest fitness. individuals, ensuring the efficiency of global search and the flexibility of local selection. Specifically, this embodiment can adopt a tournament selection mechanism to perform selection operations and select W individuals with the highest fitness. The tournament scale is set to k, which means that k individuals are randomly selected from the population N for comparison each time, and the individual with the highest fitness is selected as the winner W, and it is added to the candidate list of the next population. S500: Select a crossover point cp for the two weight arrays of the W individuals with the highest fitness to perform a crossover operation to generate a new weight distribution scheme. In the crossover operation, a crossover point is randomly selected, and part of the genetic information of the parent individual is exchanged to generate a new offspring individual. This process simulates chromosome exchange in biological genetics, aiming to explore the scheduling solution space, generate offspring with new characteristics by inheriting the excellent characteristics of the parent, thereby increasing the diversity of the population and improving the search ability of the algorithm. For example, two parent individuals P1 and P2 are selected, one or more crossover points are randomly selected on the chromosome (weight combination), and their partial weight combinations are exchanged to generate new offspring, and finally a new generation of individuals is generated. S600: Based on the mutation probability Pm, it is determined whether each individual performs weight mutation, and the formula is used Perform mutation operation, where and denote the indices of individuals and genes, respectively. For the population Individuals, For this individual genes, is the gene value before mutation, mutate represents the specific implementation function of the mutation operation, and finally a new generation of population is obtained to enter the next round of iteration. The mutation operation introduces new scheduling characteristics to increase population diversity by randomly adjusting the gene sequence of individuals. The probability of mutation Pm is generally set to 0.01 to 0.1 to ensure the stability of the algorithm and the quality of the solution. For the selected individuals, one or more gene loci are randomly selected for mutation. The mutation operation is executed through a formula to generate random perturbations and introduce new features to explore uncovered solution space to prevent the algorithm from falling into local optimality. In real number coding, the mutation operation can generate random perturbations through Gaussian distribution to increase the diversity of scheduling schemes. Through the above steps, the adaptive task allocation algorithm gradually explores the solution space in scheduling optimization, realizes resource allocation, priority control and minimization of resource consumption for multiple tasks, and effectively improves the scheduling efficiency and system performance of multi-task neural network processors.

[0042] As an optional implementation, Figure 4 As shown, step S100 specifically includes: S110: For the elements in the first set of weights, the value of the element is given by the formula Random generation, where rand(0,1) means generating a random floating point number in the interval [0,1]. is the number of calculation tasks. S120: The elements in the second set of weights are calculated by ensuring that the sum of the weights is 1; through the above steps, the sum of the two sets of weights for each individual is equal to 1, which satisfies a key constraint in the optimization process. S130: Repeat the above process until an initialization population of size N is generated.

[0043] As an optional implementation, in step S200, the fitness function is defined as follows: , where f(i) represents the fitness value of the i-th individual in the population, XLRS(i) and RSS(i) represent the predicted values ​​of the individual's logical resource consumption XLRS and resident set size RSS, respectively. and is a weight factor used to adjust the relative importance of the logical resource consumption XLRS and the resident set size RSS in the fitness function calculation. and The value can be set to 1, but specific performance optimization requirements can also be achieved by assigning different weights. In order to calculate the fitness of each individual, the model is first used to predict the logical resource consumption XLRS and resident set size RSS of the individual. The scheduling scheme is generated using the individual weight, and then the trained model is used to predict the logical resource consumption XLRS and resident set size RSS corresponding to the scheduling scheme.

[0044] The above description is only the preferred embodiment of the present invention. It is known to those skilled in the art that various changes or equivalent substitutions may be made to these features and embodiments without departing from the spirit and scope of the present invention. In addition, under the teachings of the present invention, these features and embodiments may be modified to adapt to specific circumstances and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the protection scope of the present invention.

Claims

1. A multi-task neural network processor based on a systolic array, characterized in that: The multi-task neural network processor adopts a systolic array design, and dynamically adjusts and allocates multi-task weights through an adaptive task allocation algorithm; the systolic array processes multiple tasks in parallel, and gives priority to computing tasks with higher priorities by controlling the computing states of different tasks; the systolic array transmits one input feature data and one weight data to each PE computing unit during each clock cycle, and gradually increases the number of input feature data and weight data sent to each PE computing unit in subsequent clock cycles; The multi-task neural network processor includes a multiplier module, a PE module, and a scheduler module; the multiplier module adopts a shift-add multiplier; the PE module adopts a data flow-driven multi-task architecture design, including multiple PE operation units, which are used to process data transmitted by the previous PE operation unit or the left PE operation unit; the scheduler module performs single-task scheduling and multi-task scheduling, the single-task scheduling optimizes the execution of a single task in the PE module, and the multi-task scheduling manages multiple tasks, allocates computing resources and dynamically adjusts task priorities; The adaptive task allocation algorithm comprises the following steps: S100: the scheduling scheme corresponding to each individual i is represented by a weight array as ,in and They are floating point arrays corresponding to high-priority tasks and low-priority tasks, representing individual The two groups of task decision weights are used to generate a population of size N, and the initialization population is obtained; S200: based on the logical resource consumption XLRS, the resident set size , use the support vector machine model to train the task load prediction module, and calculate the fitness value of each individual through the fitness function; S300: use the fitness change detection calculation formula Determine whether the algorithm convergence criteria are met, where Represents the fitness value of the optimal solution calculated by the fitness function in the current iteration, It represents the fitness value of the optimal solution calculated by the fitness function in the previous iteration. To set the threshold, if the condition is met, the current optimal solution is output, otherwise S400 is executed; S400: k individuals are randomly selected from the population N, and the Bayesian optimization mechanism is used to select the individuals with the highest fitness. Individual; S5 00: Select a crossover point cp for the two weight arrays of the W individuals with the highest fitness to perform a crossover operation and generate a new weight distribution scheme; S600: Determine whether to perform weight mutation for each individual based on the mutation probability Pm, and use the formula Perform mutation operation, where and denote the indices of individuals and genes, respectively. For the population Individuals, For this individual genes, is the gene value before mutation, mutate represents the specific implementation function of the mutation operation, and finally a new generation of population is obtained to enter the next round of iteration.

2. A multi-task neural network processor based on a systolic array according to claim 1, characterized in that: The multiplier module performs multiplication calculation by adding item by item, and performs calculation according to whether each bit of the multiplier is 1. If it is 1, the multiplicand is shifted and added.

3. A multi-task neural network processor based on a systolic array according to claim 2, characterized in that: The multiplier module performs a multiplication-addition operation in each clock cycle and transmits the calculation result to the adjacent PE operation unit at the same time.

4. The multi-task neural network processor based on a systolic array according to claim 1, characterized in that: It also includes a global enable terminal START. When the global enable terminal START signal is activated, the enable signal of the PE module can take effect.

5. A multi-task neural network processor based on a systolic array according to claim 4, characterized in that: The single-task scheduling can dynamically adjust the working state of the PE computing unit of a specific task to start or stop the PE computing unit; the multi-task scheduling automatically manages the task queue through an adaptive allocation algorithm to enable high-priority tasks to obtain computing resources first.

6. The multi-task neural network processor based on a systolic array according to claim 1, characterized in that: The step S100 specifically includes: S110: For the elements in the first set of weights, their values ​​are given by the formula Randomly generated, where It means to generate a random floating point number in the interval [0,1]. To calculate the number of tasks; S120: performing normalization calculation on the elements in the second group of weights by satisfying the constraint that the sum of the weights is 1; S130: Repeat the above process until an initialization population of size N is generated.

7. The multi-task neural network processor based on a systolic array according to claim 1, characterized in that: In step S200, the fitness function is defined as follows: , in, Indicates the first The fitness value of each individual, and Respectively represent The predicted values ​​of individual logical resource consumption and resident set size, and is a weight factor used to adjust the relative importance of logic resource consumption and resident set size in the fitness function calculation.

Citation Information

Patent Citations

  • Dynamic reconfigurable PE unit and PE array for graph neural network reasoning

    CN113705773A

  • TransCNN medical eye fundus image classification algorithm based on hyper-parameter optimization

    CN115965807A