An automated resource allocation method for hosting neural networks

By analyzing the hierarchical structure and computing characteristics of the neural network, combining the load prediction model, automatically adjusting resource allocation, and dynamically adjusting the calculation accuracy mode during the task execution, the problem of insufficient resource allocation and lack of dynamic monitoring in the existing technology is solved, and efficient resource utilization and computational efficiency optimization of neural network tasks is achieved.

CN119356836BActive Publication Date: 2025-05-30NATURAL SEMANTICS (QINGDAO) TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411943814.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-30
Estimated Expiration
2044-12-27

AI Technical Summary

Technical Problem

The existing technology lacks a comprehensive analysis and intelligent prediction mechanism for computing accuracy requirements in neural network tasks, resulting in resource allocation being inefficient enough, unable to flexibly adjust resource allocation, and lacks dynamic monitoring and feedback mechanisms, and fails to respond in real time to the adjustment requirements of hardware resources usage status and computing accuracy.

Method used

By analyzing the hierarchical structure and computing characteristics of the neural network, the computing accuracy requirements of each network layer are determined, and based on the load prediction model and hardware performance, the processing unit computing power and memory bandwidth required for each computing stage are predicted. Then, the resource allocation is automatically adjusted, and the use of hard resources is monitored in real time during the task execution process, and the calculation accuracy mode is dynamically adjusted.

Benefits of technology

It realizes efficient resource scheduling and allocation of neural network tasks, reduces computing complexity, optimizes the utilization rate of hardware resources, avoids resource waste, improves task execution efficiency, and reduces system energy consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119356836B_ABST
    Figure CN119356836B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of computer technology, and particularly relates to an automated resource allocation method for hosting neural networks, comprising the following steps: S1: Analyze the computational precision requirements of each network layer according to the hierarchical structure and computational characteristics of the neural network; S2: Predict the hardware resources of the processing unit computing power and memory bandwidth required for each computational stage; S3: Automatically adjust the resource allocation of the task according to the resource prediction result of S2; S4: Dynamically adjust the computational precision mode according to the resource load; S5: After the task is completed, release the allocated hardware resources and generate a resource usage feedback report. With the present invention, by intelligently analyzing the computational precision requirements and hardware resource status of neural network tasks, dynamically adjusting the computational precision mode and optimizing resource allocation, the utilization efficiency of hardware resources is significantly improved, the computational complexity is reduced, and resource waste and bottleneck problems are effectively avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and particularly to an automated resource allocation method for hosting neural networks. Background Art

[0002] With the rapid development of deep learning technology, neural networks have made remarkable progress in applications in multiple fields such as image recognition, natural language processing, and speech recognition; the computational tasks of neural network models usually require a large amount of hardware resources, especially when dealing with complex computations and large-scale data, the demand for CPU, GPU, and memory bandwidth increases sharply; in order to improve the computational efficiency of neural networks and reduce the consumption of hardware resources, researchers have proposed the concept of mixed-precision computing, that is, different precision modes are adopted in different computational stages according to the computational requirements of tasks; this method can not only effectively reduce the computational complexity but also reduce the use of hardware resources; however, how to intelligently switch between different computational precision modes to achieve optimal resource utilization and computational efficiency is still a technical problem to be solved urgently.

[0003] Currently, although there are certain resource allocation methods for the hardware scheduling of neural networks, the existing technologies mainly focus on static resource allocation or experience-based dynamic adjustment, lacking a comprehensive analysis and intelligent prediction mechanism for the computational precision requirements of tasks; the existing methods have the following deficiencies in the face of the computational intensity differences, uneven utilization of hardware resources, and dynamic adjustment of precision during the execution of neural network tasks: First, they cannot accurately analyze the computational characteristics and precision requirements of each neural network layer, resulting in inefficient resource allocation; second, the existing resource allocation strategies often do not make flexible adjustments based on the real-time load of tasks, causing resource waste or bottlenecks; third, the existing resource allocation methods lack a dynamic monitoring and feedback mechanism and fail to respond in real time to the usage status of hardware resources and the adjustment requirements of computational precision. Summary of the Invention

[0004] Based on the above purposes, the present invention provides an automated resource allocation method for hosting neural networks.

[0005] An automated resource allocation method for hosting neural networks includes the following steps:

[0006] S1: Analyze the computational precision requirements of each network layer according to the hierarchical structure and computational characteristics of the neural network, and select a computational precision mode according to the computational intensity of the task;

[0007] S2: Based on the computational precision requirements determined in S1, use a load prediction model combined with hardware performance to predict the computational capabilities of processing units and the hardware resources of memory bandwidth required for each computational stage.

[0008] S3: Automatically adjust the resource allocation of the task according to the resource prediction result of S2;

[0009] S4: During the execution of the task, monitor the usage of hardware resources in real time, and dynamically adjust the computing precision mode according to the resource load;

[0010] S5: After the task is completed, release the allocated hardware resources and generate a resource usage feedback report.

[0011] Optionally, the S1 specifically includes:

[0012] S11: By parsing the topological structure of the neural network, identify the type of each network layer and its position in the entire network, and determine the function and role of each layer;

[0013] S12: Evaluate the computing characteristics of each network layer, including computing intensity, memory access pattern, and data dependency, and obtain the specific computing requirement parameters of each layer;

[0014] S13: Based on the computing characteristics evaluated in step S12, determine the computing precision required for each network layer, select the corresponding computing precision mode to match the computing requirements of the layer. The computing precision modes include full precision mode FP32, half precision mode FP16, and low precision mode INT8; for each layer, determine the computing precision mode through the following formula:

[0015] ;

[0016] where, represents the selected precision mode, represents the computing intensity of the current network layer, and are the preset high computing intensity and low computing intensity thresholds respectively; according to the different computing intensities, select the corresponding computing precision mode: when the computing intensity is higher than select the full precision mode FP32; when the computing intensity is between and select the half precision mode FP16; when the computing intensity is lower than select the low precision mode INT8.

[0017] Optionally, the S2 specifically includes:

[0018] S21: Real-time collect the performance data of each processing unit in the system, including clock frequency, core utilization rate, and memory bandwidth usage rate, as the basic data for resource demand prediction;

[0019] S22: Use the historical task execution data to train a load prediction model based on the support vector machine algorithm; this load prediction model can learn and identify the resource demand patterns for the processing unit's computing power and memory bandwidth in different computing stages according to the computing accuracy requirements of the task.

[0020] S23: Take the computing accuracy requirements determined in S1 as the input, and combine with the hardware performance data collected in step S21, input them into the load prediction model, and output the specific values of the computing power of the processing unit and the memory bandwidth requirements for each computing stage.

[0021] Optionally, S23 specifically includes:

[0022] S221: Collect execution data from multiple historical tasks, including the computing accuracy requirements of each task, the computing power of the processing unit, the memory bandwidth usage, and the computing load of task execution.

[0023] S222: Extract features from the collected historical task execution data, select features related to resource requirements, including computing accuracy, computing intensity, and memory bandwidth utilization rate; at the same time, normalize the data.

[0024] S223: Use the normalized data as the input to train the load prediction model through the support vector machine algorithm. The support vector machine maximizes the interval between categories by finding the optimal hyperplane and uses the kernel function to map the data into a high-dimensional space to handle non-linear relationships.

[0025] S224: Use the particle swarm optimization method to optimize the parameters of the support vector machine model; the particle swarm optimization simulates the collaborative movement of the particle swarm in the search space and iteratively adjusts the parameters of the support vector machine to minimize the prediction error.

[0026] Optionally, S224 specifically includes:

[0027] S2241: Randomly initialize a group of particles in the parameter space, and each particle represents a set of support vector machine parameters to be optimized.

[0028] S2242: For each particle, train the support vector machine model according to the parameter group it represents, and evaluate the prediction error of the model on the validation set as the fitness value of the particle.

[0029] S2243: According to the personal best position and the global best position of the particle, adjust the speed and position of each particle to push the particle to move towards a better parameter area.

[0030] S2244: Repeat steps S2242 and S2243 until the predetermined stop condition is met, including reaching the maximum number of iterations or the fitness value reaching the preset threshold.

[0031] S2245: Select the best parameter group found in the particle swarm optimization process to construct the final support vector machine load prediction model, and the expression is: , where is the predicted resource demand value, and are the parameters of the support vector machine model, is the input feature vector, is the output of the support vector machine model.

[0032] Optionally, the S3 specifically includes:

[0033] S31: According to the predicted processing unit computing power and memory bandwidth requirements in S2, formulate a resource allocation strategy, including determining the resource allocation ratio, priority, and allocation timing of each processing unit;

[0034] S32: Apply a resource allocation algorithm based on linear programming to optimize the resource allocation of each processing unit. By establishing resource constraint conditions and an objective function, solve the optimal resource allocation plan to maximize resource utilization;

[0035] S33: Automatically adjust the resource allocation of the neural network task on each processing unit according to the resource allocation plan optimized in S32.

[0036] Optionally, the S32 specifically includes:

[0037] S321: Define resource constraint conditions according to the computing tasks of the neural network and the limitations of hardware resources, including the maximum computing power, memory bandwidth, and storage resources of the processing unit; Set the resource constraint conditions as: , , where represents the computing power of the th processing unit, represents the maximum computing power of this processing unit, is the total number of processing units;

[0038] S322: Establish an objective function to maximize resource utilization and computing performance; Set the objective function as: , where represents the resource demand of the th processing unit, represents the computing power of the th processing unit;

[0039] S323: Add a memory bandwidth constraint to the objective function to ensure that the utilization of memory bandwidth is controlled when allocating resources; Set the memory bandwidth constraint condition as: , where For the memory bandwidth requirement of the th processing unit, is the total memory bandwidth available to the system;

[0040] S324: According to the resource constraint conditions and objective function established in S321, S322, and S323, use linear programming to solve the optimal resource allocation scheme; through the result of the optimal solution obtained by linear programming optimization, obtain the optimal resource allocation value for each processing unit, and the solution expression is: , where is the optimal resource allocation scheme obtained through optimal solution.

[0041] Optionally, the S4 specifically includes:

[0042] S41: Real-time collect and record the current load, memory usage rate, bandwidth utilization rate, and temperature performance parameters of each processing unit in the system;

[0043] S42: Analyze the real-time hardware resource data collected in step S41, determine whether the current resource usage status reaches the preset threshold, and identify potential resource bottlenecks or overload situations; the judgment formula is: load ratio , where is the usage rate of the current hardware resource, is the maximum available value of the hardware resource; if the load ratio is greater than or equal to the set threshold of 0.8, it is considered that the current hardware resource is in a state close to overload and there are potential resource bottlenecks;

[0044] S43: Based on the resource load analysis result in step S42, call the calculation precision adjustment algorithm to automatically determine whether it is necessary to adjust the calculation precision mode of the neural network task.

[0045] Optionally, the S43 specifically includes:

[0046] S431: Based on the load ratio calculated in step S42, let the load ratio be , calculate the adjustment decision parameter , and its calculation formula is:

[0047] ,

[0048] where represents the current load ratio, and are the preset high-load and low-load thresholds respectively; automatically determine whether to adjust the calculation precision mode through this formula;

[0049] S432: According to the adjustment decision parameter , automatically switch the computing precision mode of the neural network task;

[0050] If is the reduced precision mode, switch the computing precision of the current network layer from the full precision mode to the low precision mode;

[0051] If is the enhanced precision mode, switch the computing precision of the current network layer from the low precision mode to the full precision mode;

[0052] If is the keep current precision mode, maintain the existing computing precision configuration unchanged.

[0053] Optionally, the S5 specifically includes:

[0054] S51: After the neural network task is completed, mark all the hardware resources allocated to this task, including CPU cores, GPU units, memory bandwidth, and storage resources, and prepare for the release operation;

[0055] S52: Through the resource scheduling controller, sequentially execute the resource recovery operations, specifically including:

[0056] CPU resource release, terminate the CPU cores allocated to the task, and restore them to the idle state;

[0057] GPU resource release, close the GPU units allocated to the task, release the GPU memory bandwidth, and mark the GPU resources as reallocable;

[0058] Memory bandwidth release, release the occupation of the memory bandwidth by the task, and restore the normal use state of the memory bandwidth;

[0059] Storage resource release, unload the storage resources allocated to the task, and clear the cache;

[0060] S53: After the resource release is completed, collect the resource usage data during the task execution, including the usage time of each processing unit, resource occupancy rate, memory bandwidth utilization, and energy consumption data;

[0061] S54: Organize the resource usage data collected in S54 to generate a resource usage feedback report, including resource allocation, resource usage efficiency, energy consumption statistics, and optimization suggestions.

[0062] Advantages of the present invention:

[0063] In the present invention, by introducing an automated resource allocation method based on computational precision requirement analysis and hardware performance prediction, it is possible to intelligently provide efficient resource scheduling and allocation for neural network tasks; by dynamically adjusting the computational precision mode, the system can not only reduce the computational complexity, but also optimize the utilization rate of hardware resources, achieving efficient use of resources during task execution; compared with the prior art, this solution establishes an effective connection between computational precision and hardware resource management, and can perform real-time switching between multiple precision modes according to the actual requirements and computational load of the task, thereby achieving the effect of reducing unnecessary resource waste and optimizing computational efficiency.

[0064] In the present invention, by continuously monitoring the usage of hardware resources and dynamically adjusting the computational precision according to the load analysis results, the efficient utilization of system resources is ensured, and the situation of overload or resource bottleneck can be avoided; after the task is completed, the system automatically releases the allocated hardware resources and generates a resource usage feedback report to provide data support for the optimization of subsequent resource allocation strategies; this intelligent resource allocation and management method greatly improves the execution efficiency of neural network tasks and reduces the system energy consumption. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only those of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0066] Figure 1 Schematic diagram of the automated resource allocation method for carrying neural networks according to an embodiment of the present invention;

[0067] Figure 2 Schematic diagram of the resource allocation process for automatically adjusting tasks according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0068] The present invention will be described in detail below in conjunction with the drawings and specific embodiments. At the same time, it should be noted here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments. For some well-known technologies, those skilled in the art can also adopt other alternative methods for implementation; moreover, the drawing part is only for more specifically describing the embodiments, and is not intended to specifically limit the present invention.

[0069] It should be noted that in the specification, the mention of "an embodiment", "embodiments", "exemplary embodiments", "some embodiments", etc. indicates that the described embodiments may include specific features, structures or characteristics, but not necessarily every embodiment includes such specific features, structures or characteristics. Additionally, when combining an embodiment to describe a specific feature, structure or characteristic, the implementation of such a feature, structure or characteristic in combination with other embodiments (whether explicitly described or not) should be within the knowledge of those skilled in the relevant art.

[0070] Generally, terms can be understood at least in part from their use in context. For example, at least in part depending on the context, the term "one or more" as used herein can be used to describe any feature, structure or characteristic in a singular sense, or can be used to describe a combination of features, structures or characteristics in a plural sense. Additionally, the term "based on" can be understood as not necessarily intended to convey a set of exclusive factors, but rather can alternatively, at least in part depending on the context, allow for the existence of other factors that are not necessarily explicitly described.

[0071] As Figure 1 - Figure 2 shown, an automated resource allocation method for hosting a neural network includes the following steps:

[0072] S1: Analyze the computational precision requirements of each network layer according to the hierarchical structure and computational characteristics of the neural network, and select a computational precision mode according to the computational intensity of the task;

[0073] S2: Based on the computational precision requirements determined in S1, use a load prediction model combined with hardware performance to predict the hardware resources of the processing unit computing power and memory bandwidth required for each computational stage;

[0074] S3: According to the resource prediction results of S2, automatically adjust the resource allocation of the task. If the system resources are insufficient, reduce the computational precision of the task; if the resources are sufficient, increase the computational precision;

[0075] S4: During the execution of the task, monitor the usage of hardware resources in real time, and dynamically adjust the computational precision mode according to the resource load;

[0076] S5: After the task is completed, release the allocated hardware resources and generate a resource usage feedback report for reference when optimizing resource allocation for subsequent tasks.

[0077] S1 specifically includes:

[0078] S11: By parsing the topological structure of the neural network, identify the type of each network layer (such as convolutional layer, fully connected layer, pooling layer, etc.) and its position in the entire network, and determine the function and role of each layer;

[0079] S12: Evaluate the computational characteristics of each network layer, including computational intensity, memory access pattern, and data dependency, and obtain the specific computational requirement parameters for each layer;

[0080] S13: Based on the computational characteristics evaluated in step S12, determine the computational precision required for each network layer, select the corresponding computational precision mode to match the computational requirements of the layer. The computational precision modes include full-precision mode FP32, half-precision mode FP16, and low-precision mode INT8; for each layer, determine the computational precision mode through the following formula:

[0081] ;

[0082] where, represents the selected precision mode, represents the computational intensity of the current network layer, and are the preset high computational intensity and low computational intensity thresholds respectively; according to the different computational intensities, select the corresponding computational precision mode: when the computational intensity is higher than , select the full-precision mode FP32; when the computational intensity is between and , select the half-precision mode FP16; when the computational intensity is lower than , select the low-precision mode INT8; through the above steps S11 to S13, the method dynamically selects an appropriate computational precision mode according to the computational intensity of each network layer, and combines a specific threshold division strategy to ensure that the selection of the precision mode meets the computational characteristic requirements of each layer; using this method can optimize resource allocation, reduce energy consumption, and improve system processing ability while ensuring computational efficiency.

[0083] S2 specifically includes:

[0084] S21: Real-time collect the performance data of each processing unit (including CPU, GPU, etc.) in the system, including clock frequency, core utilization rate, and memory bandwidth utilization rate, as the basic data for resource demand prediction;

[0085] S22: Use historical task execution data to train a load prediction model based on the support vector machine algorithm; this load prediction model can learn and identify the resource demand patterns of different computational stages for the computational ability and memory bandwidth of the processing unit according to the computational precision requirements of the task;

[0086] S23: Use the computing precision requirement determined in S1 as the input, combine it with the hardware performance data collected in step S21, and input them into the load prediction model to output the specific values of the computing power of the processing unit and the memory bandwidth requirement for each computing stage, which serve as the reference basis for subsequent resource allocation; Through the above steps S21 to S23, the method can accurately collect and analyze hardware performance data, and use the load prediction model trained based on the support vector machine algorithm to accurately predict the computing power of the processing unit and the memory bandwidth required for each computing stage. This process ensures the accuracy and reliability of resource requirement prediction and provides a solid technical foundation for subsequent dynamic resource allocation.

[0087] S23 specifically includes:

[0088] S221: Collect execution data from multiple historical tasks, including the computing precision requirement of each task, the computing power of the processing unit (such as CPU / GPU core frequency, core utilization rate, etc.), the memory bandwidth usage, and the computing load of task execution (such as the computing intensity and data transmission volume of each network layer);

[0089] S222: Extract features from the collected historical task execution data, select features related to resource requirements, including computing precision, computing intensity, and memory bandwidth utilization rate; At the same time, perform normalization processing on the data to ensure that the influence of data in different dimensions on the training model is balanced and avoid certain features having too much influence on the model;

[0090] S223: Use the normalized data as the input, and train the load prediction model through the support vector machine algorithm. The support vector machine maximizes the interval between categories by finding the optimal hyperplane and uses the kernel function to map the data into a high-dimensional space to handle non-linear relationships; Its optimization objective is: , where, is the normal vector of the hyperplane, is the slack variable, is the penalty parameter, is the number of samples; This optimization objective is used to find the optimal classification hyperplane so that the model can accurately predict the hardware resources required for each computing stage;

[0091] S224: Use the Particle Swarm Optimization (PSO) method to optimize the parameters of the support vector machine model; The particle swarm optimization iteratively adjusts the parameters of the support vector machine (such as the penalty parameter and the kernel function parameter) by simulating the collaborative movement of the particle swarm in the search space to minimize the prediction error.

[0092] The specific steps of particle swarm optimization in S224 include:

[0093] S2241: Randomly initialize a group of particles in the parameter space, where each particle represents a set of support vector machine parameters to be optimized;

[0094] S2242: For each particle, train a support vector machine model according to the parameter set it represents, and evaluate the prediction error of the model on the validation set as the fitness value of the particle;

[0095] S2243: Adjust the velocity and position of each particle according to the particle's own best position and the global best position, pushing the particle towards a better parameter region;

[0096] S2244: Repeat steps S2242 and S2243 until a predetermined stopping condition is met, including reaching the maximum number of iterations or the fitness value reaching a preset threshold;

[0097] S2245: Select the best parameter set found during the particle swarm optimization process to construct the final support vector machine load prediction model, with the expression: , where is the predicted resource demand value (such as computing power and memory bandwidth), and are the parameters of the support vector machine model, is the input feature vector, is the output of the support vector machine model; this model predicts the computing power and memory bandwidth of the processing unit required for each computing stage by inputting the computing precision requirements and other key features of the task; through the above steps, historical task data can be systematically collected and processed, and the model parameters can be accurately optimized using the support vector machine algorithm combined with the particle swarm optimization method, thereby constructing an efficient and accurate load prediction model; this model can accurately predict the computing power and memory bandwidth of the processing unit required for each computing stage of the neural network, providing a reliable reference basis for subsequent dynamic resource allocation and significantly improving the intelligent level of resource allocation and the overall performance of the system.

[0098] S3 specifically includes:

[0099] S31: According to the predicted computing power and memory bandwidth requirements of the processing unit in S2, formulate a resource allocation strategy, including determining the resource allocation ratio, priority, and allocation timing of each processing unit;

[0100] S32: Apply a resource allocation algorithm based on linear programming (LP) to optimize the resource allocation of each processing unit. By establishing resource constraint conditions and an objective function, solve the optimal resource allocation plan to ensure that the resource requirements of each computing stage are met while maximizing resource utilization;

[0101] S33: According to the resource allocation scheme optimized in S32, automatically adjust the resource allocation of the neural network task on each processing unit, specifically including dynamically adjusting the number of cores of the CPU and GPU, the memory bandwidth allocation, and the storage resource allocation, to ensure that the task can operate efficiently with the allocated resources.

[0102] S32 specifically includes:

[0103] S321: According to the computing tasks of the neural network and the limitations of the hardware resources, define the resource constraint conditions, including the maximum computing power, memory bandwidth, and storage resources of the processing unit; set the resource constraint conditions as: , , where, represents the computing power of the th processing unit, represents the maximum computing power of this processing unit, is the total number of processing units;

[0104] S322: To maximize the resource utilization rate and computing performance, establish an objective function; the design of the objective function aims to minimize resource waste and ensure that the computing requirements of the neural network task are met; set the objective function as: , where, represents the resource requirement of the th processing unit, represents the computing power of the th processing unit. The objective function aims to make the actual resource allocation as close as possible to the required resources to maximize the resource usage efficiency;

[0105] S323: Add a memory bandwidth constraint to the objective function to ensure that the utilization of the memory bandwidth is controlled when allocating resources; set the memory bandwidth constraint condition as: , where, is the memory bandwidth requirement of the th processing unit, is the total memory bandwidth available to the system;

[0106] S324: According to the resource constraint conditions and the objective function established in S321, S322, and S323, use linear programming to solve the optimal resource allocation scheme; through the result of the linear programming optimization solution, obtain the best resource allocation value for each processing unit, and the solution expression is: , where, is the optimal resource allocation scheme obtained through optimized solution; through the above steps S321 to S324, the resource constraint conditions and objective function constructed using the linear programming algorithm effectively optimize the resource allocation; the constraint conditions ensure that the resource allocation of each processing unit does not exceed its hardware capacity, while the objective function maximizes the resource utilization rate and reduces resource waste during the calculation process; through the optimized resource allocation scheme, the system can allocate the computing power and memory bandwidth of the processing unit in a more efficient manner, thereby improving the execution efficiency of the neural network computing task, ensuring the satisfaction of the computing requirements and maximizing the resource utilization rate.

[0107] S4 specifically includes:

[0108] S41: Real-time collect and record the current load, memory usage rate, bandwidth utilization rate, and temperature performance parameters of each processing unit (including CPU, GPU, etc.) in the system to ensure comprehensive monitoring of resource usage;

[0109] S42: Analyze the real-time hardware resource data collected in step S41 to determine whether the current resource usage status reaches a preset threshold and identify potential resource bottlenecks or overload situations; the judgment formula is: load ratio , where is the usage rate of the current hardware resource, is the maximum available value of the hardware resource (such as the maximum frequency of the CPU, the maximum computing power of the GPU, the maximum memory bandwidth, etc.); if the load ratio is greater than or equal to the set threshold of 0.8, it is considered that the current hardware resource is in a state close to overload and there is a potential resource bottleneck; also analyze the correlation between different resources, such as the usage relationship between the CPU and memory. If any resource is detected to exceed the preset threshold, it is considered that there is a potential resource bottleneck or overload situation;

[0110] S43: Based on the resource load analysis result in step S42, call the calculation precision adjustment algorithm to automatically determine whether it is necessary to adjust the calculation precision mode of the neural network task; if the resource load is detected to be too high, select the reduced calculation precision mode; if the resource load is low, select the increased calculation precision mode; through the above steps S41 to S43, the method can, during the execution of the neural network task, real-time monitor and analyze the usage of hardware resources, and dynamically adjust the calculation precision mode according to the change of the resource load; this dynamic adaptation mechanism ensures that the calculation precision can be reduced to maintain the stable operation of the system when resources are tense, and the calculation precision can be increased to optimize the model performance when resources are abundant.

[0111] S43 specifically includes:

[0112] S431: Based on the load ratio calculated in step S42, let the load ratio be , calculate and adjust the decision parameters , and its calculation formula is:

[0113] ,

[0114] wherein, represents the current load ratio, and are respectively the preset high-load and low-load thresholds; whether to adjust the computing precision mode is automatically determined through this formula;

[0115] S432: According to the adjustment decision parameter calculated in S431 , automatically switch the computing precision mode of the neural network task;

[0116] If is the precision reduction mode, switch the computing precision of the current network layer from the full precision mode to the low precision mode;

[0117] If is the precision improvement mode, switch the computing precision of the current network layer from the low precision mode to the full precision mode;

[0118] If is to maintain the current precision mode, keep the existing computing precision configuration unchanged; through the above steps, whether to adjust the computing precision mode of the neural network task is automatically determined, and this method effectively improves the intelligent level of resource allocation and optimizes the computing efficiency.

[0119] S5 specifically includes:

[0120] S51: After the neural network task is completed, mark all the hardware resources allocated to this task, including CPU cores, GPU units, memory bandwidth, and storage resources, and prepare for the release operation;

[0121] S52: Through the resource scheduling controller, sequentially execute the resource recovery operations, which specifically include:

[0122] Release of CPU resources, terminate the CPU cores allocated to the task, and restore them to the idle state;

[0123] Release of GPU resources, turn off the GPU units allocated to the task, release the GPU memory bandwidth, and mark the GPU resources as available for reallocation;

[0124] Release of memory bandwidth, relieve the task's occupancy of the memory bandwidth, and restore the normal use state of the memory bandwidth;

[0125] Release of storage resources, unload the storage resources allocated to the task, and empty the cache;

[0126] S53: After the resource release is completed, collect the resource usage data during task execution, including the usage time of each processing unit, the resource occupancy rate, the memory bandwidth utilization, and the energy consumption data. These data serve as the basic information for resource usage feedback;

[0127] S54: Organize the resource usage data collected in S54 to generate a resource usage feedback report, including resource allocation, resource usage efficiency, energy consumption statistics, and optimization suggestions. Through the above steps S51 to S54, after the neural network task is completed, the allocated hardware resources can be systematically released, and a resource usage feedback report can be accurately generated. The resource release mechanism ensures the timely recycling and reuse of system resources, avoiding resource idleness or conflicts. The comprehensive collection of resource usage data and the generation of the feedback report provide detailed basis for the optimization of subsequent resource allocation strategies.

[0128] This invention covers any substitutions, modifications, equivalent methods, and solutions made within the spirit and scope of this invention. To enable the public to have a thorough understanding of this invention, specific details are described in detail in the following preferred embodiments of this invention. However, those skilled in the art can fully understand this invention without the description of these details. Additionally, to avoid unnecessary confusion to the essence of this invention, well-known methods, processes, procedures, components, and circuits are not described in detail.

[0129] The above description is only a preferred embodiment of this invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of this invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of this invention.

Claims

1. A method for automatic resource allocation for carrying a neural network, characterized in that: The following steps are involved: S1: Analyze the computational accuracy requirements of each network layer based on the hierarchical structure and computational characteristics of the neural network, and select the computational accuracy mode based on the computational intensity of the task; S2: Based on the calculation accuracy requirements determined in S1, the load prediction model is used in combination with hardware performance to predict the hardware resources of the processing unit computing power and memory bandwidth required for each calculation stage, including: S21: collects performance data of each processing unit in the system in real time, including clock frequency, core utilization, and memory bandwidth utilization, as basic data for resource demand prediction; S22: Use historical task execution data to train a load prediction model based on a support vector machine algorithm; the load prediction model can learn and identify resource demand patterns for processing unit computing power and memory bandwidth at different computing stages according to the computing accuracy requirements of the task; S23: Taking the calculation accuracy requirement determined in S1 as input, combined with the hardware performance data collected in step S21, and inputting it into the load prediction model, the specific values ​​of the processing unit computing power and memory bandwidth requirement required for each calculation stage are output; S3: Automatically adjust the resource allocation of tasks based on the resource prediction results of S2, including: S31: formulating a resource allocation strategy based on the processing unit computing power and memory bandwidth requirements predicted in S2, including determining the resource allocation ratio, priority, and allocation sequence of each processing unit; S32: Apply a resource allocation algorithm based on linear programming to optimize resource allocation of each processing unit, and solve the optimal resource allocation solution by establishing resource constraints and objective functions to maximize resource utilization; S33: Automatically adjust resource allocation of the neural network task on each processing unit according to the resource allocation scheme optimized in S32; S4: During the task execution, the usage of hardware resources is monitored in real time, and the calculation accuracy mode is dynamically adjusted according to the resource load, including: S41: collect and record the performance parameters of the current load, memory usage, bandwidth utilization and temperature of each processing unit in the system in real time; S42: Analyze the real-time hardware resource data collected in step S41 to determine whether the current resource usage status reaches a preset threshold, and identify potential resource bottlenecks or overload conditions; the judgment formula is: load ratio ,in, is the current utilization rate of hardware resources, is the maximum available value of the hardware resources; if the load ratio is greater than or equal to the set threshold of 0.8, it is considered that the current hardware resources are close to being overloaded and there is a potential resource bottleneck; S43: Based on the resource load analysis result in step S42, calling the calculation accuracy adjustment algorithm to automatically determine whether the calculation accuracy mode of the neural network task needs to be adjusted; S5: After the task is completed, the allocated hardware resources are released and a resource usage feedback report is generated.

2. The method for automatic resource allocation for carrying a neural network according to claim 1, characterized in that: The S1 specifically includes: S11: By analyzing the topological structure of the neural network, identify the type of each network layer and its position in the entire network, and determine the function and role of each layer; S12: Evaluate the computing characteristics of each network layer, including computing density, memory access pattern, and data dependency, and obtain the specific computing requirement parameters of each layer; S13: Based on the computing characteristics evaluated in step S12, determine the computing accuracy required for each network layer, and select the corresponding computing accuracy mode to match the computing requirements of the layer, wherein the computing accuracy modes include full-precision mode FP32, half-precision mode FP16, and low-precision mode INT8.

3. The method for automatic resource allocation for carrying a neural network according to claim 1, characterized in that: The S22 specifically includes: S221: Collecting execution data from multiple historical tasks, including the computational accuracy requirement of each task, the computational capability of the processing unit, memory bandwidth usage, and computational load of task execution; S222: extracting features from the collected historical task execution data, selecting features related to resource requirements, including computing accuracy, computing density, and memory bandwidth usage; and normalizing the data; S223: using the normalized data as input, training the load forecasting model through a support vector machine algorithm, the support vector machine maximizes the interval between categories by finding the optimal hyperplane, and uses a kernel function to map the data to a high-dimensional space to handle nonlinear relationships; S224: Optimizing the parameters of the support vector machine model using a particle swarm optimization method; the particle swarm optimization iteratively adjusts the parameters of the support vector machine to minimize the prediction error by simulating the coordinated movement of a particle swarm in a search space.

4. The method for automatic resource allocation for carrying a neural network according to claim 3, characterized in that: The S224 specifically includes: S2241: randomly initialize a group of particles in the parameter space, each particle represents a set of support vector machine parameters to be optimized; S2242: For each particle, train a support vector machine model according to the parameter group it represents, and evaluate the prediction error of the model on the validation set as the fitness value of the particle; S2243: According to the particle's own optimal position and the global optimal position, the speed and position of each particle are adjusted to push the particle to move to a more optimal parameter area; S2244: repeating steps S2242 and S2243 until a predetermined stop condition is met, including reaching a maximum number of iterations or a fitness value reaching a preset threshold; S2245: Select the best parameter group found in the particle swarm optimization process to build the final support vector machine load prediction model.

5. The method for automatic resource allocation for carrying a neural network according to claim 1, characterized in that: The S32 specifically includes: S321: According to the computing tasks of the neural network and the limitations of hardware resources, define resource constraints, including the maximum computing power of the processing unit, memory bandwidth and storage resources; set the resource constraints as: ,in, Indicates The computing power of the processing units, Indicates the maximum computing capacity of the processing unit, and N is the total number of processing units; S322: To maximize resource utilization and computing performance, establish an objective function; set the objective function to: ,in, Indicates The resource requirements of each processing unit, Indicates The computing power of each processing unit; S323: Add a memory bandwidth constraint to the objective function to ensure that the memory bandwidth utilization is controlled when allocating resources; set the memory bandwidth constraint condition as: ,in, For the The memory bandwidth requirement of each processing unit, is the total memory bandwidth available to the system; S324: According to the resource constraints and objective function established in S321, S322 and S323, the optimal resource allocation scheme is solved by linear programming; the optimal resource allocation value of each processing unit is obtained through the result of the linear programming optimization solution, and the solution expression is: ,in, It is the optimal resource allocation solution obtained through optimization.

6. The method for automatic resource allocation for carrying a neural network according to claim 1, characterized in that: The S43 specifically includes: S431: Based on the load ratio calculated in step S42, assuming the load ratio is L, calculate the adjustment decision parameter D, and the calculation formula is: , Where L represents the current load ratio, and These are the preset high load and low load thresholds respectively; this formula is used to automatically decide whether to adjust the calculation accuracy mode; S432: Automatically switch the calculation accuracy mode of the neural network task according to the adjustment decision parameter D calculated in S431; If D is in reduced precision mode, the calculation precision of the current network layer is switched from full precision mode to low precision mode; If D is in the improved precision mode, the calculation precision of the current network layer is switched from the low-precision mode to the full-precision mode; If D is to maintain the current precision mode, the existing calculation precision configuration remains unchanged.

7. The method of automatic resource allocation for carrying a neural network according to claim 1, characterized in that: The S5 specifically includes: S51: After the neural network task is completed, all hardware resources allocated to the task, including CPU cores, GPU units, memory bandwidth and storage resources, are marked and prepared for release; S52: The resource scheduling controller sequentially performs resource recovery operations, specifically including: CPU resources are released, the CPU core assigned to the task is terminated, and it is restored to an idle state; GPU resource release, shut down the GPU unit assigned to the task, release GPU memory bandwidth, and mark GPU resources as reallocatable; Memory bandwidth is released, the memory bandwidth occupation by tasks is relieved, and the memory bandwidth is restored to normal usage; Release storage resources, unload storage resources allocated to tasks, and clear cache; S53: After the resource release is completed, the resource usage data during the task execution is collected, including the usage time of each processing unit, resource occupancy rate, memory bandwidth utilization and energy consumption data; S54: Organize the resource usage data collected in S54 and generate a resource usage feedback report, including resource allocation, resource usage efficiency, energy consumption statistics, and optimization suggestions.

Citation Information

Patent Citations

  • Resource scheduling system of AI intelligent computing center

    CN117472587A