Computing resource optimization method and system for analyzing tasks
Through deep reinforcement learning and load prediction models, the problem of uneven allocation of computing resources in the video AI analysis environment is solved, real-time prediction and dynamic adaptive scheduling of computing resources are realized, and efficient stability and resource utilization of the system are improved.
Patent Information
- Application Number
- CN202510195842.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-07-04
AI Technical Summary
In the multi-tasking and multi-scenario video AI analysis environment, the allocation of computing resources lacks real-time monitoring and prediction, resulting in waste of resources and delayed response, making it difficult to meet the computing needs at high loads and over-allocation of resources at low loads.
By obtaining the characteristic information of the analysis task and the system resource status, using the deep Q network model of deep reinforcement learning to determine the task priority, combining historical and real-time data to predict the load, dynamically adjust the resource allocation of edge nodes, establish a unified computing resource pool, and realize flexible task scheduling.
It improves the utilization rate of computing resources and system response speed, optimizes task scheduling, ensures that critical tasks are processed in a timely manner in a high-load environment, avoids resource waste, and improves the overall operating efficiency and stability of the system.
Smart Images

Figure CN120256087A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a method and system for optimizing computing resources for analysis tasks. Background Art
[0002] Currently, with the increasing demand for intelligent video AI analysis in fields such as video surveillance, smart cities, and security, the system needs to process a large amount of video data from different scenarios and different task types simultaneously.
[0003] However, traditional technologies mainly adopt static resource allocation and fixed scheduling strategies, usually allocating computing resources according to pre-set rules, lacking real-time monitoring and prediction of task characteristics and system status. In the face of scenarios with diverse task types, different data scales, and high real-time requirements, this approach often fails to respond to load fluctuations in a timely manner, resulting in critical tasks being unable to quickly obtain the necessary computing resources during high loads, while resources are over-allocated during low loads, causing resource waste. Therefore, there are obvious shortcomings in the existing technologies in terms of ensuring high system responsiveness and resource utilization efficiency. Its static scheduling mechanism is difficult to meet the dynamic changes in resource requirements during the parallel operation of multiple tasks, thus leading to problems such as response delays and uneven resource scheduling.
[0004] Therefore, how to achieve real-time prediction and dynamic adaptive scheduling of computing resources in a multi-task and multi-scenario video AI analysis environment has become a technical problem that urgently needs to be solved. Summary of the Invention
[0005] The present invention provides a method, system, electronic device, and storage medium for optimizing computing resources for analysis tasks to solve the defects in the existing technologies and achieve real-time prediction and dynamic adaptive scheduling of computing resources in a multi-task and multi-scenario video AI analysis environment.
[0006] The present invention provides a method for optimizing computing resources for analysis tasks, including the following steps: Obtain multiple analysis tasks and obtain the feature information of each of the analysis tasks; Determine the priority of each of the analysis tasks according to the feature information of each of the analysis tasks and the current available computing resource status information of the system, and generate a priority queue according to the priorities of all the analysis tasks; Obtain the task load prediction value for a future time period according to historical task load data and real-time system status data; Allocate analysis tasks with real-time requirements higher than a preset standard to edge nodes for execution according to the type, data scale, and data source location of the analysis tasks, and dynamically adjust the resource allocation of the edge nodes according to the task load prediction value; Establish a unified computing resource pool, and dynamically allocate available resources from the computing resource pool according to the priority queue and the predicted task load value to execute multiple parallel analysis tasks.
[0007] According to a computing resource optimization method for analysis tasks provided by the present invention, determining the priority of each analysis task according to the characteristic information of each analysis task and the current available computing resource status information of the system specifically includes: Input the characteristic information of each analysis task and the computing resource status information into a priority calculation model to obtain the priority value of each analysis task output by the priority calculation model; wherein, the priority calculation model is a deep Q-network model trained based on a deep reinforcement learning algorithm; Determine the priority of each analysis task according to the priority value of each analysis task.
[0008] According to a computing resource optimization method for analysis tasks provided by the present invention, it further includes a training method for the priority calculation model: Randomly initialize the initial weights of the deep Q-network model, and set the initial weights of the target network to be the same as the initial weights of the deep Q-network model; Continuously collect training samples, and construct an experience replay pool according to the training samples; Periodically randomly extract training samples from the experience replay pool, perform batch training on the deep Q-network model, and use the target network to stabilize the training process; Repeat the steps of periodically randomly extracting training samples from the experience replay pool, performing batch training on the deep Q-network model, and using the target network to stabilize the training process until the error of the deep Q-network model is lower than a preset threshold or reaches a predetermined maximum number of training rounds; Take the trained deep Q-network model as the final priority calculation model.
[0009] According to a computing resource optimization method for analysis tasks provided by the present invention, the step of periodically randomly extracting training samples from the experience replay pool, performing batch training on the deep Q-network model, and using the target network to stabilize the training process specifically includes: Periodically randomly extract multiple training samples from the experience replay pool to form a mini-batch of data; Input the mini-batch of data into the deep Q-network model to obtain the predicted priority output by the deep Q-network model; Calculate the error between the predicted priority and the target priority label, and update the weights of the deep Q-network model according to the error through the backpropagation algorithm and the gradient descent method. After a preset training step interval, copy the weights of the current deep Q-network model to the target network to reduce training oscillation and accelerate model convergence.
[0010] According to a method for optimizing computing resources for analysis tasks provided by the present invention, obtaining a task load prediction value for a future time period based on historical task load data and real-time system status data specifically includes: Obtain historical task load data and real-time system status data, and perform missing value filling and normalization processing on the historical task load data and the real-time system status data to obtain preprocessed historical task load data and preprocessed real-time system status data; Input the preprocessed historical task load data and the preprocessed real-time system status data into a prediction model to obtain a task load prediction value output by the prediction model; wherein, the prediction model is based on a long short-term memory network or a Transformer network and uses a multi-time step prediction mechanism to generate a task load prediction value for a future time period.
[0011] According to a method for optimizing computing resources for analysis tasks provided by the present invention, allocating analysis tasks with real-time requirements higher than a preset standard to edge nodes for execution according to the type, data scale, and data source location of the analysis tasks, and dynamically adjusting the resource allocation of the edge nodes according to the task load prediction value specifically includes: Based on the type, data scale, and data source location of the analysis task, determine the requirement of the analysis task for response latency, and mark analysis tasks with a real-time requirement higher than a preset real-time threshold as real-time tasks; Allocate the real-time tasks to the edge node closest to the corresponding data source location to shorten the data transmission path and reduce network latency; Combine the CPU, GPU, and memory usage of the edge node and the task load prediction value to dynamically evaluate the available resources of the edge node; According to the evaluation result, allocate or reclaim computing resources for each of the real-time tasks allocated to the edge node.
[0012] According to a method for optimizing computing resources for analysis tasks provided by the present invention, the method further includes: During the execution of the analysis task, continuously monitor the actual task load value; Compare the actual task load value with the task load prediction value to obtain a deviation value; When it is determined that the deviation value is greater than a preset deviation threshold, update the model parameters of the priority calculation model and the prediction model; Adjust the priority or resource allocation strategy of the analysis task according to the updated priority calculation model and the updated prediction model.
[0013] According to an analysis task computing resource optimization method provided by the present invention, the feature information is at least one of task type, data scale, and computing complexity; the computing resource status information includes the utilization rates of CPU, GPU, and memory.
[0014] The present invention also provides an analysis task computing resource optimization system, including the following modules: an acquisition module, configured to acquire multiple analysis tasks and acquire the feature information of each analysis task; A first processing module, configured to determine the priority of each analysis task according to the feature information of each analysis task and the current available computing resource status information of the system, and generate a priority queue according to the priorities of all the analysis tasks; A second processing module, configured to obtain a task load prediction value for a future time period according to historical task load data and real-time system status data; A third processing module, configured to allocate analysis tasks with real-time requirements higher than a preset standard to edge nodes for execution according to the type, data scale, and data source location of the analysis task, and dynamically adjust the resource allocation of the edge nodes according to the task load prediction value; A fourth processing module, configured to establish a unified computing resource pool, and dynamically allocate available resources from the computing resource pool according to the priority queue and the task load prediction value to execute multiple parallel analysis tasks.
[0015] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the computing resource optimization method of any one of the above is implemented.
[0016] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the computing resource optimization method of any one of the above is implemented.
[0017] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the computing resource optimization method of any one of the above is implemented.
[0018] In summary, one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: By obtaining multiple analysis tasks and their characteristic information, a comprehensive understanding of the task types, data scales, and computing requirements is achieved, thereby providing accurate inputs for task scheduling. By determining task priorities based on task characteristics and the current computing resource status and generating a priority queue, key tasks are processed preferentially, resource allocation is optimized, thereby improving the execution efficiency of high-priority tasks and avoiding inefficient occupation of computing resources. By predicting future task loads by combining historical task load data and real-time system status data, resource configuration is adjusted in advance, thereby reducing task latency and improving computing resource utilization. By allocating high-real-time tasks to the optimal edge nodes based on task types, data scales, and data source locations and dynamically adjusting resources, tasks are efficiently executed in a low-latency environment, thereby reducing data transmission overhead, improving system response speed, and avoiding computing bottlenecks at the central node. By establishing a unified computing resource pool and dynamically allocating available resources based on the priority queue and load prediction values, resources are flexibly scheduled and load balancing is optimized, thereby improving computing resource utilization and achieving efficient and stable task scheduling and resource management. Finally, this solution ensures real-time prediction and dynamic adaptive scheduling of computing resources in a video AI analysis environment with multiple tasks and multiple scenarios. Brief Description of the Drawings
[0019] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0020] Figure 1 It is one of the flow diagrams of the method for optimizing computing resources of the analysis task provided by the present invention.
[0021] Figure 2 It is another flow diagram of the method for optimizing computing resources of the analysis task provided by the present invention.
[0022] Figure 3 It is the third flow diagram of the method for optimizing computing resources of the analysis task provided by the present invention.
[0023] Figure 4 It is the fourth flow diagram of the method for optimizing computing resources of the analysis task provided by the present invention.
[0024] Figure 5 It is the fifth flow diagram of the method for optimizing computing resources of the analysis task provided by the present invention.
[0025] Figure 6 It is the sixth flow diagram of the method for optimizing computing resources of the analysis task provided by the present invention.
[0026] Figure 7 It is a schematic structural diagram of a computing resource optimization system for analysis tasks provided by the present invention.
[0027] Figure 8 It is a schematic structural diagram of an electronic device provided by the present invention. Detailed implementation manners
[0028] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without making creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.
[0029] It should be noted that in the description of the present invention, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "including one..." does not exclude the presence of additional identical elements in the process, method, article or device including the element. The orientation or positional relationship indicated by terms such as "upper", "lower", etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the system or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be construed as a limitation to the present invention. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0030] The terms "first", "second", etc. in the present invention are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present invention can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same type and do not limit the number of objects. For example, the first object can be one or multiple. In addition, "and / or" means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.
[0031] The following will be combined with Figures 1-8 to describe the analysis task computing resource optimization method, system, electronic device and storage medium provided by the present invention.
[0032] Figure 1 is one of the schematic flowcharts of the method for optimizing computing resources for analysis tasks provided by the present invention. As Figure 1 shown, it includes but is not limited to the following steps: Step 101: Obtain multiple analysis tasks and obtain the feature information of each analysis task.
[0033] In a computing environment with multi-task parallelism, in order to achieve efficient scheduling of computing resources, the system first needs to obtain multiple analysis tasks and extract the feature information of each analysis task, so as to provide a basis for subsequent priority evaluation and resource allocation. Since the computing requirements of analysis tasks may vary greatly, for example, different task types may involve different AI model inference processes, different data scales may affect storage and bandwidth occupancy, and the computing complexity directly determines the load intensity of the required computing resources. Therefore, relying solely on fixed rules or static scheduling strategies cannot meet the dynamic task scheduling requirements. Therefore, in Step 101, the system receives multiple analysis tasks to be processed through the task management module and extracts their feature information based on the specific attributes of the tasks. Among them, the feature information at least includes key indicators such as task type, data scale, and computing complexity.
[0034] In the specific implementation process, after receiving the task request, the system will parse the basic parameters of the task. For example, the task type may involve AI analysis modes such as object detection, behavior recognition, or image segmentation. The data scale can be measured by information such as the resolution, frame rate, and duration of the input video stream. The computing complexity can be evaluated in combination with indicators such as the depth of the AI inference model and the inference time. At the same time, in order to ensure the accuracy of scheduling, the system will also query the current computing resource status information, including the usage rates of CPU, GPU, and memory, to evaluate the available computing power of the system, so as to provide comprehensive data support for subsequent task priority evaluation.
[0035] In this way, the system can accurately characterize the task attributes at the initial stage of task submission, enabling subsequent scheduling decisions to be dynamically adjusted based on specific task requirements and computing resource status, rather than simply processing according to the order of task arrival.
[0036] Step 102: Determine the priority of each analysis task according to the feature information of each analysis task and the current available computing resource status information of the system, and generate a priority queue according to the priorities of all analysis tasks.
[0037] In a multi-task parallel computing environment, the reasonable allocation of computing resources is crucial for ensuring the overall operation efficiency of the system. Since different analysis tasks have different task types, data scales, and computational complexities, and the computing resource status of the system (such as the utilization rates of CPU, GPU, and memory) is dynamically changing, simply relying on a fixed task scheduling strategy cannot achieve precise resource allocation. To ensure that high-priority tasks can be processed in a timely manner when resources are scarce and to improve the throughput capacity of the system, in step 102, a priority calculation mechanism is constructed to evaluate the priority of each analysis task and generate a task scheduling queue based on the priority results.
[0038] In a possible implementation manner, the feature information is at least one of the task type, data scale, and computational complexity; the computing resource status information includes the utilization rates of CPU, GPU, and memory. Figure 2 It is the second flow diagram of the computing resource optimization method for analysis tasks provided by the present invention. As Figure 2 shown, in step 102, according to the feature information of each analysis task and the current available computing resource status information of the system, the priority of each analysis task is determined, which is specifically implemented through steps 201-202: Step 201: Input the feature information and computing resource status information of each analysis task into the priority calculation model to obtain the priority value of each analysis task output by the priority calculation model; among them, the priority calculation model is a deep Q-network model trained based on the deep reinforcement learning algorithm.
[0039] Step 202: Determine the priority of each analysis task according to the priority value of each analysis task.
[0040] Specifically, the system first obtains the feature information of each analysis task and combines it with the current available computing resource status information, and inputs these data into the priority calculation model. The priority calculation model adopts a deep Q-network model trained based on the deep reinforcement learning algorithm. Its core idea is to enable the system to automatically optimize the task priority allocation strategy according to the characteristics of different tasks and the system resource situation through long-term interactive learning. The specific training method of the priority calculation model will be described in detail in the subsequent embodiments. Through the calculation of the deep Q-network model, each analysis task will obtain a priority value, which represents the processing urgency of the task under the current computing resource conditions. Subsequently, the system sorts the tasks according to the priority values of all analysis tasks and finally forms a priority queue.
[0041] In this process, the deep reinforcement learning ability of the priority calculation model enables the system to adaptively adjust the priority policy. Even in the face of dynamically changing task loads and resource conditions, it can still ensure the priority execution of critical tasks by continuously learning and optimizing the task sorting logic. In addition, due to the introduction of the experience replay and target network mechanisms in the training process of the deep Q-network model, the results of priority calculation are more stable, avoiding frequent adjustments caused by short-term fluctuations in task priority evaluation, thus enhancing the robustness and reliability of the scheduling system.
[0042] In this way, the implementation of steps 201 and 202 effectively solves the problem of uneven distribution of computing resources in a multi-task environment, enabling the system to still give priority to meeting the computing requirements of high-importance or high-real-time tasks in the face of tight computing resources. At the same time, by optimizing the task sorting strategy through deep reinforcement learning, the system can continuously improve the efficiency of resource scheduling after long-term operation, reduce the interference of low-priority tasks on critical tasks, and ultimately improve the utilization rate of overall computing resources and the system response speed.
[0043] In a possible implementation manner, Figure 3 is the third schematic diagram of the process of the computing resource optimization method for analyzing tasks provided by the present invention. As Figure 3 shown, the training method of the priority calculation model is specifically implemented through steps 301-305: Step 301: Randomly initialize the initial weights of the deep Q-network model, and set the initial weights of the target network to the same values as the initial weights of the deep Q-network model.
[0044] In the priority calculation model based on deep reinforcement learning, the performance of the deep Q-network model (Deep Q-Network, DQN) directly affects the accuracy of priority evaluation for analysis tasks, and the setting of the initial weights of the model plays a key role in the stability and convergence speed of training. Since the training of the deep Q-network model involves complex state-action-reward-state (SARS) mapping relationships, if the initial weights are not set reasonably, it may lead to slow learning speed of the model in the initial stage of training, or even problems such as gradient disappearance or gradient explosion. Therefore, in step 301, the system randomly initializes the initial weights of the deep Q-network model and synchronously sets the initial weights of the target network to ensure the stability of training.
[0045] Specifically, when performing step 301, the system first constructs a deep Q-network model, which includes an input layer, multiple hidden layers, and an output layer. The input layer is used to receive the feature information of the task and the current computing resource status information of the system. The hidden layers perform feature extraction and representation learning on the input data through neural networks, and the output layer is used to predict the priority value of the task. In the model initialization stage, the system initializes the neural network weights of the deep Q-network model in a random manner using a uniform distribution or a normal distribution to ensure that the model has sufficient exploration ability in the initial stage of training. At the same time, to further improve the stability of training, the system creates a target network and sets its initial weights to the same value as the initial weights of the deep Q-network model. The role of the target network is to provide a stable learning target and prevent training oscillations or divergence caused by drastic fluctuations in the target value during the Q-value update process.
[0046] Step 302: Continuously collect training samples and construct an experience replay pool based on the training samples.
[0047] During the training process of the deep Q-network model, the quality and diversity of the sample data directly affect the learning effect of the model. Since the decision-making environment of task priorities is highly dynamic, if training samples are directly generated based on the current environment, it may lead to strong sample correlations in the model learning process, thereby affecting the convergence speed and generalization ability of training. Therefore, in step 302, the system continuously collects training samples and constructs an experience replay pool based on these samples to store the training data in different states during the task execution process, enabling the model to learn more stable task scheduling strategies from historical experiences.
[0048] In the specific implementation process, when the system executes an analysis task, it continuously monitors the feature information of the task, the computing resource status information, as well as the actions taken and the corresponding reward values during the task scheduling process, and organizes this information into a quadruple data structure of "state - action - reward - next state". Whenever the system completes a task scheduling decision, it records the current state information, the current action, the reward obtained after executing this action, and the new state information after execution, and stores it in the experience replay pool. To ensure the effectiveness of the experience replay pool, the system adopts a storage strategy with a fixed capacity. When the experience replay pool reaches the capacity limit, it preferentially removes the earliest stored samples to ensure that the model can always learn the latest environmental changes and avoid interference from outdated samples on the training effect.
[0049] In this way, the implementation of step 302 ensures that the deep Q-network model can be trained from rich task scheduling experiences, rather than relying solely on the latest single data point, reducing the correlation between samples and improving the diversity and representativeness of the training data. In addition, the construction of the experience replay pool also enables the model to randomly sample and perform batch training during the training process, thereby avoiding the model's over-reliance on short-term experiences and helping to more stably learn the optimal task priority calculation strategy.
[0050] Step 303: Periodically and randomly draw training samples from the experience replay pool, perform batch training on the deep Q-network model, and use the target network to stabilize the training process.
[0051] During the training process of the deep Q-network model, relying solely on the data generated by the current environment for learning easily leads the model to fall into a local optimum and may result in insufficient generalization ability of the model due to uneven sample distribution. Therefore, in step 303, the system randomly draws training samples from the experience replay pool periodically and performs batch training on the deep Q-network model to reduce sample correlation, improve training stability, and use the target network mechanism to reduce the drastic fluctuations during the training process and accelerate model convergence.
[0052] In a possible implementation manner, Figure 4 is the fourth schematic flow diagram of the computational resource optimization method for analyzing tasks provided by the present invention. As Figure 4 shown, step 303 specifically includes steps 401-404: Step 401: Periodically and randomly draw multiple training samples from the experience replay pool to form a mini-batch of data.
[0053] Step 402: Input the mini-batch of data into the deep Q-network model to obtain the predicted priorities output by the deep Q-network model.
[0054] Step 403: Calculate the error between the predicted priority and the target priority label, and update the weights of the deep Q-network model according to the error through the backpropagation algorithm and the gradient descent method.
[0055] Step 404: After a preset number of training steps interval, copy the weights of the current deep Q-network model to the target network to reduce training oscillations and accelerate model convergence.
[0056] In steps 401 to 404, the system adopts a batch training method, periodically draws samples from the experience replay pool for mini-batch training, and combines error calculation, gradient descent optimization, and target network update strategies to improve the adaptability and stability of the deep Q-network model in the task scheduling environment.
[0057] In the specific implementation process, the system first sets a fixed sampling interval to ensure that the model can evenly cover different stages during task execution in the training process, rather than relying only on task scheduling data in a short period of time. At the beginning of each training cycle, the system randomly extracts a certain number of training samples from the experience replay pool, and each training sample contains a quadruple data structure of "state - action - reward - next state". The system uses a random sampling method without replacement during sampling to ensure the diversity of training data and prevent the model from falling into local optimal solutions. At the same time, to ensure the representativeness of training data, the system limits the maximum capacity of the experience replay pool. When the stored new samples exceed the set capacity, the system automatically removes the earliest stored samples, enabling the model to always learn the latest task scheduling patterns without being interfered by outdated data.
[0058] Next, the small batch of data sampled in step 401 is parsed according to the quadruple structure of "state - action - reward - next state", and the current state information is input into the deep Q - network model. This model is composed of multiple - layer neural networks. The input layer receives the feature information of the task and the state information of the system's computing resources. The hidden layer extracts key features through non - linear transformation, and finally the output layer generates the priority value of the task. The output of the model is the Q - values corresponding to each possible action (i.e., different task scheduling decisions) in the current task state, where the Q - value represents the long - term reward that can be obtained after performing a specific action in the current state. The system selects the Q - value corresponding to the current task scheduling strategy from the set of output Q - values as the predicted task priority value and stores it for subsequent calculation of the target priority and gradient optimization.
[0059] Furthermore, the target priority is calculated. The calculation of the target priority is based on the Bellman equation in deep reinforcement learning, that is, the long - term reward is calculated using the optimal Q - value of the next state and the current reward value. Subsequently, the system compares the predicted priority calculated in step 402 with the target priority and constructs a loss value based on the Mean Squared Error (MSE) or Huber loss function to measure the deviation between the task priority value output by the current model and the ideal task priority value. Next, the system uses the backpropagation algorithm to propagate the loss value along the computational path of the neural network and adjusts the weight parameters of the model through the gradient descent method, making the model more accurately predict the optimal priority of the task in future training iterations. To improve the stability of training, the system can use the Adam optimizer or RMSprop optimizer to dynamically adjust the learning rate, ensuring that the model can make larger weight updates in the early stage of training and reduce the parameter adjustment amplitude when converging in the later stage of training, thereby preventing oscillation or overfitting.
[0060] Finally, set a fixed step interval N. Whenever the number of training iterations of the deep Q-network model reaches N, the system triggers the target network weight update mechanism and copies the weights of the current deep Q-network model to the target network. The role of the target network is to provide a stable reference for Q-value calculation, so that the target Q-value during training will not cause training oscillations due to the frequent changes of the current network parameters. Before updating the target network weights, the system will calculate the loss trend of the model and determine whether the training of the current model is in a stable convergence state, avoiding weight synchronization when the model has not reached a certain stability, thus ensuring the stability of the strategy during training. In addition, in actual implementation, the update of the target network weights can adopt a soft update method, that is, the weights of the target network gradually approach the weights of the current deep Q-network model at a small update rate τ, rather than directly replacing them completely, to further reduce training fluctuations and enable the model to smoothly optimize the task scheduling strategy during long-term training.
[0061] Step 304: Repeat the steps of periodically randomly sampling training samples from the experience replay pool, batch-training the deep Q-network model, and stabilizing the training process using the target network until the error of the deep Q-network model is lower than the preset threshold or reaches the predetermined maximum number of training rounds.
[0062] During the training process of the deep Q-network model, ensuring that the model can effectively converge to the optimal task scheduling strategy is the key to improving the utilization efficiency of the system's computing resources. Since the characteristic information of the analysis tasks and the computing resource status are dynamically changing during training, if the training termination condition is set unreasonably, it may lead to insufficient model training, making the task priority calculation model unable to accurately predict the task scheduling strategy, or lead to overtraining and overfitting, making the model difficult to adapt to the complexity of the actual environment. Therefore, in Step 304, the system ensures that the model terminates training when the error is lower than the preset threshold or reaches the predetermined maximum number of training rounds by repeating the steps of sampling training samples from the experience replay pool, batch-training the deep Q-network model, and stabilizing the training process using the target network, thereby obtaining a stable and highly generalized priority calculation model.
[0063] In the specific implementation process, in each training iteration, the system continuously randomly extracts a small batch of training samples from the experience replay pool and inputs them into the deep Q-network model to calculate the predicted priority in the current state. At the same time, the target priority is calculated through the target network, and gradient optimization is performed according to the error value between the predicted priority and the target priority, enabling the model to continuously adjust the weight parameters to minimize the error. As the number of training rounds increases, the system monitors the change trend of the loss value of the model in real time and determines whether the current error is lower than the set error threshold. If the error has been reduced to the preset range, it is considered that the model has fully learned the task scheduling strategy, and at this time, the training is terminated to avoid overfitting. If the error fails to reach the preset standard, but the number of training rounds has reached the maximum training limit set by the system, the system determines that the model has reached a stable state and can also terminate the training to ensure a balance between the computational complexity and the generalization ability of the model.
[0064] Step 305: Use the trained deep Q-network model as the final priority calculation model.
[0065] Specifically, after the training termination condition in Step 304 is met, the finally optimized deep Q-network model is obtained. This model has learned the mapping relationship between task feature information, computing resource status information, and the optimal scheduling strategy through multiple rounds of iteration and can adaptively predict task priorities under different task loads. Subsequently, the system formats the model and stores it in the model storage unit of the computing resource management module to ensure that the model can be quickly loaded and called during subsequent task scheduling. In addition, to improve the robustness of the model, the system can adopt a model version management strategy, that is, while storing the final model, retain some historical training versions and evaluate the model effect during the system operation. If a performance decline of the task priority calculation model is detected, it can be rolled back to a better historical model version to ensure the stability of the system during long-term operation.
[0066] Step 103: Obtain the predicted value of the task load for the future time period based on historical task load data and real-time system status data; In the process of dynamic scheduling of computing resources, accurately predicting future task loads is crucial for optimizing resource allocation. If the system schedules only based on the computing resource requirements of the current task without anticipating future task load changes, it may lead to a shortage of computing resources during the task peak period, affecting the execution of high-priority tasks, or resource idleness during low load, reducing the overall system utilization rate. Therefore, in Step 103, the system calculates the predicted value of the task load for the future time period by combining historical task load data and real-time system status data, providing data support for subsequent task scheduling and resource optimization.
[0067] In a possible implementation manner,Figure 5 This is the fifth schematic flowchart of the method for optimizing computing resources for analysis tasks provided by the present invention. As Figure 5 shown, step 103 specifically includes steps 501-502: Step 501: Obtain historical task load data and real-time system state data, and perform missing value filling and normalization processing on the historical task load data and real-time system state data to obtain preprocessed historical task load data and preprocessed real-time system state data.
[0068] During the task load prediction process, the integrity and quality of the input data have a decisive impact on the prediction accuracy. If the original historical task load data and real-time system state data are directly used for prediction, the generalization ability of the prediction model may decrease due to problems such as data missing, outliers, or inconsistent scales, thereby affecting the dynamic allocation efficiency of computing resources. Therefore, in step 501, the system first obtains historical task load data and real-time system state data, and performs missing value filling and normalization processing on them to ensure the integrity and consistency of the input data, thereby improving the accuracy and stability of the prediction model.
[0069] In the specific implementation process, the system extracts historical task load data from the task scheduling log and monitors the current computing resource status information in real time, including key metrics such as the utilization rates of CPU, GPU, and memory, the task queue length, and the task execution duration. Since the task execution environment is highly dynamic, some historical data may be missing or contain outliers. Therefore, the system first performs missing value filling in the data preprocessing stage. For a small number of missing data points, the system uses methods such as linear interpolation, mean filling, or moving window smoothing to complete the filling, while for a large range of missing data, the system uses a time series-based interpolation algorithm for reconstruction to reduce the impact of data incompleteness on the prediction results. In addition, to ensure that different feature data are input into the prediction model within the same scale range, the system performs normalization processing on the task load data and computing resource status data. For example, Z-score normalization or Min-Max normalization is used to convert all data to the same numerical range to avoid bias in the learning process of the prediction model caused by features with large values.
[0070] Step 502: Input the preprocessed historical task load data and preprocessed real-time system state data into the prediction model to obtain the task load prediction value output by the prediction model; among them, the prediction model is based on a long short-term memory network or a Transformer network, and uses a multi-time step prediction mechanism to generate the task load prediction value for a future time period.
[0071] In computing resource scheduling, accurately predicting future task loads is crucial for optimizing resource allocation and enhancing the system's response capabilities. If the system relies solely on the current computing resource status for task scheduling without being able to anticipate future changes in task loads, it may lead to resource shortages during high loads or over-allocation during low loads, thereby reducing the overall operating efficiency of the system. Therefore, in step 502, the system inputs the preprocessed historical task load data and real-time system status data into a prediction model. Using a Long Short-Term Memory (LSTM) network or a Transformer network, through a multi-time-step prediction mechanism, it generates task load prediction values for future time periods to ensure that the task scheduling strategy can adapt to future resource requirements.
[0072] In the specific implementation process, the system first invokes the prediction model. The input layer of this model receives the standardized historical task load data and the current system status data, including characteristic information such as task execution frequency, computing resource utilization rates (CPU, GPU, memory), and task queue length. Subsequently, the model extracts features from the input data based on time series analysis methods and performs prediction tasks through different deep learning architectures. If a Long Short-Term Memory (LSTM) network is used, the model, through its gating mechanism and long-term dependency learning ability, captures the changing trends of task loads between different time steps and uses the hidden state to transmit information in the time dimension, enabling the model to effectively predict task loads for multiple future time steps. If a Transformer network is used, the model calculates the weights of key features in the time series through the Self-Attention mechanism, enabling it to efficiently model the global dependencies of task loads and improve the accuracy and stability of predictions.
[0073] Step 104: According to the type, data scale, and data source location of the analysis task, allocate analysis tasks with real-time requirements higher than the preset standard to edge nodes for execution, and dynamically adjust the resource allocation of edge nodes according to the task load prediction value; In a multi-task parallel video AI analysis environment, the reasonable allocation of computing resources directly affects the system's processing efficiency and response speed. If all tasks are executed on the central node, it will not only increase data transmission latency but may also cause bottlenecks in computing resources during peak loads, thereby reducing the task execution efficiency. Therefore, in step 104, the system determines tasks with real-time requirements higher than the preset standard based on the type, data scale, and data source location of the task, allocates these tasks to edge nodes for execution, and dynamically adjusts the computing resource allocation of edge nodes in combination with the task load prediction value to enhance the flexibility of task scheduling and the overall computing efficiency of the system.
[0074] In a possible implementation manner, Figure 6It is the sixth schematic flowchart of the method for optimizing computing resources for analysis tasks provided by the present invention. As Figure 6 shown, step 104 specifically includes steps 601-604: Step 601: Based on the type, data scale, and data source location of the analysis task, determine the response latency requirement of the analysis task, and mark the analysis tasks with a response latency higher than the preset real-time threshold as real-time tasks.
[0075] In a video AI analysis environment with multi-task parallelism, the requirements for computing resources and response time vary greatly among different tasks. If all tasks are processed according to the same scheduling strategy, it may cause high-real-time tasks to be unable to execute in a timely manner due to insufficient computing resources or data transmission latency, thus affecting the overall performance of the system. Therefore, in step 601, the system determines the response latency requirement of the task based on the type, data scale, and data source location of the analysis task, and marks the tasks with a response latency higher than the preset real-time threshold as real-time tasks, so as to ensure that critical tasks can obtain resources preferentially and be executed at the optimal computing location, thereby improving the task processing efficiency and overall response speed of the system.
[0076] In the specific implementation process, the system first analyzes the characteristic information of each task. Among them, the task type is used to judge the real-time processing requirement of the task. For example, applications such as object detection and behavior analysis usually have high real-time requirements, while tasks such as offline data mining and historical data analysis are less sensitive to latency. The data scale determines the computing intensity and network transmission overhead of the task. Larger-scale video stream tasks require stronger computing capabilities and may cause additional latency due to transmission bandwidth limitations, while tasks with smaller data scales are relatively flexible. The data source location affects the selection of task execution nodes. Generally, the closer the data source is to the computing node, the lower the task processing latency. Therefore, the system comprehensively considers the three factors of task type, data scale, and data source location, sets a real-time threshold, scores the real-time performance of all tasks, and screens out the tasks higher than the threshold and marks them as real-time tasks.
[0077] Step 602: Allocate the real-time tasks to the edge node closest to the corresponding data source location to shorten the data transmission path and reduce network latency.
[0078] In the specific implementation process, the system first obtains the data source location corresponding to each task based on the real-time task list marked in step 601, and queries the edge computing node topology structure in the current system. The topology information of edge computing nodes includes physical location, network connection status, computing resource usage (CPU, GPU, memory, etc.), network latency, and task load. The system calculates the network transmission latency between the task and each edge node based on the data source location, and filters out the most suitable edge computing node for the task execution according to the principle of minimum transmission latency. Subsequently, the system combines the computing resource status of this edge node to evaluate whether it has sufficient computing power to execute the selected task, and allocates the task to this edge node for execution on the premise that resources permit.
[0079] During the task scheduling process, if the computing resources of the currently optimal edge node are insufficient, the system will select the sub-optimal edge node, and combine the real-time threshold of the task to determine whether it is necessary to adjust the computing resources or fallback to the central computing node for execution in necessary cases. In addition, the system will dynamically monitor the execution situation of tasks on edge nodes to ensure that tasks can run smoothly on the target edge nodes, and release resources in a timely manner after the tasks are completed, so as to avoid edge computing nodes occupying too much computing power for a long time and affecting the execution of subsequent tasks.
[0080] Step 603: Dynamically evaluate the available resources of the edge node by combining the CPU, GPU, and memory usage of the edge node and the predicted task load value.
[0081] In the specific implementation process, the system first calls the edge computing management module to monitor the computing resource occupancy of the currently assigned tasks in real time, including key metrics such as the CPU utilization rate, GPU computing load, memory occupancy ratio, and I / O throughput of each edge node. At the same time, the system obtains the task load prediction value generated in step 502, which reflects the possible changes in computing resource requirements in the future time period, enabling the system to perceive possible computing bottlenecks in advance and make corresponding adjustments. Based on these data, the system calculates the resource utilization rate of each edge node, and combines the task priority and the predicted computing resource requirements to evaluate whether the current task scheduling needs to be adjusted.
[0082] If the evaluation result indicates that the computing resource utilization rate of a certain edge node is about to exceed the set load threshold, the system will trigger a resource adjustment mechanism, including task transfer, resource expansion, or task priority adjustment. For example, when the load prediction value of a certain edge node indicates that there may be a shortage of computing resources in the future, the system can pre-migrate some low-priority tasks to other edge nodes with lower loads, or combine with the resource pool sharing module to dynamically allocate more computing resources to this edge node to ensure that critical tasks can be executed smoothly. In addition, the system can also continuously monitor the resource utilization status of edge nodes based on the real-time feedback of task execution, and dynamically optimize the computing resource allocation strategy when necessary.
[0083] Step 604: According to the evaluation result, allocate or reclaim computing resources for each real-time task allocated to the edge node.
[0084] In the specific implementation process, the system first calls the computing resource management module to monitor the usage rates of computing resources such as CPU, GPU, and memory of all edge nodes in real time, and combines with the task load prediction value to evaluate whether the computing resource requirements of the current task match the computing capabilities of the edge nodes. For edge nodes with computing resource utilization rates approaching the upper limit, the system adopts a dynamic resource expansion strategy and preferentially allocates more computing resources, such as increasing the computing allocation quota of the GPU, or calling the resource pool sharing module to schedule additional computing resources from the global computing resource pool to this edge node to ensure that critical tasks can still be executed efficiently under resource constraints. If the resource requirements of a certain task decrease during execution or the task is about to be completed, the system then triggers a resource reclaim mechanism to release the idle computing resources of this task back to the resource pool or reallocate them to other high-priority tasks to maximize the utilization rate of computing resources.
[0085] Step 105: Establish a unified computing resource pool, and dynamically allocate available resources from the computing resource pool according to the priority queue and task load prediction value to execute multiple parallel analysis tasks.
[0086] In Step 105, the system establishes a unified computing resource pool and dynamically allocates available resources from the computing resource pool according to the priority queue and task load prediction value to ensure that multiple parallel analysis tasks can obtain reasonable resource support, thereby optimizing the utilization rate of computing resources and improving the task execution efficiency.
[0087] In the specific implementation process, the system first integrates all available resources in the computing resource management module, including computing resources such as CPUs, GPUs, and memories of the central node and edge nodes, and constructs a unified computing resource pool. This computing resource pool not only contains the currently available physical computing resources, but also can work in coordination with the resource pool sharing module to achieve cross-node resource scheduling, ensuring that tasks can dynamically obtain additional resources when computing resources are scarce, and improving the overall elasticity and load balancing ability of the system. Subsequently, the system allocates resources in sequence according to the priority queue generated in step 102, ensuring that high-priority tasks obtain computing resources first, while low-priority tasks are scheduled for execution when resources are sufficient. In addition, the system combines the task load prediction value generated in step 103 to estimate future computing resource requirements, and adjusts the resource allocation strategy in advance before task execution. For example, when predicting a task load peak, the computing resource pool is expanded in advance to reduce delays caused by task backlogs or computing resource shortages.
[0088] In the process of dynamic allocation of computing resources, the system adopts an adaptive resource management mechanism to continuously monitor the resource occupancy of each task and dynamically adjust resource allocation during task execution. For example, when the actual resource requirements of a task during execution exceed expectations, the system can allocate additional computing resources for it through the resource pool sharing module to ensure that the task can be completed on time. On the contrary, if the resource requirements of a task are lower than the initially allocated computing resources, the system automatically reclaims the excess resources and reallocates them to other high-priority tasks to improve the overall utilization rate of computing resources. In addition, the system also has the ability of intelligent resource migration, that is, when the load of a computing node is too high, the system can migrate some tasks to a computing node with a lower load to balance the computing load and improve the stability of the system.
[0089] In a possible implementation manner, the method further includes the following steps: Continuously monitor the actual task load value during the execution of the analysis task; Compare the actual task load value with the task load prediction value to obtain a deviation value; When it is determined that the deviation value is greater than the preset deviation threshold, update the model parameters of the priority calculation model and the prediction model; Adjust the priority or resource allocation strategy of the analysis task according to the updated priority calculation model and the updated prediction model.
[0090] In a computing resource dynamic scheduling system, the task load is constantly changing. If only relying on static task priorities and load predictions for computing resource allocation without continuously monitoring and adjusting the actual execution situation, it may lead to a lag in the task scheduling strategy, making the computing resources unable to adapt to the real-time task requirements. Therefore, in the above steps, the system continuously monitors the actual task load value during task execution, compares it with the task load prediction value, and calculates the deviation value. If the deviation value exceeds the preset threshold, the system will update the priority calculation model and the task load prediction model, and adjust the priority of the analysis task or the resource allocation strategy to ensure that the task scheduling can adapt to the dynamic changes of the system operating environment.
[0091] In the specific implementation process, the system first continuously collects the actual task load data during task execution, including indicators such as CPU, GPU, memory usage, task execution duration, and task completion status, and stores this data in the real-time load monitoring module for subsequent analysis. Subsequently, the system calls the task load prediction module, obtains the task load prediction value for the current time period, compares the prediction value with the actual task load data, and calculates the deviation between the actual load and the predicted load. If the deviation value is within a reasonable range, it indicates that the prediction model is still effective, and the system continues to operate according to the original scheduling strategy; but if the deviation value exceeds the preset deviation threshold, it indicates that the current task load status has deviated from the expectation of the prediction model, and the scheduling strategy needs to be dynamically adjusted.
[0092] Furthermore, the system updates the priority calculation model and the task load prediction model based on the deviation analysis results. Specifically, the system first collects the recent actual load data and inputs it into the priority calculation model, and retrains the model through deep reinforcement learning methods to enable it to more accurately evaluate task priorities and ensure that the task scheduling strategy can adapt to the latest computing resource status. At the same time, the system adjusts the parameters of the task load prediction model or retrains it to optimize the time series learning ability of the prediction model and improve its prediction accuracy for future task loads. During the model update process, the system adopts an incremental learning mechanism and optimizes only based on the latest task data to reduce the computational overhead and ensure that the model can quickly adapt to environmental changes.
[0093] After the model update is completed, the optimized priority calculation model and the task load prediction model are used to dynamically adjust the task priority and the computing resource allocation strategy. For example, if the updated priority calculation model determines that the priorities of certain tasks need to be increased, the system will adjust the task scheduling order and give priority to critical tasks; if the task load prediction model indicates that the future load of certain computing nodes will exceed their computing capabilities, the system will adjust the task allocation plan in advance and migrate some tasks to computing nodes with lower loads. In addition, the system can also dynamically increase or reclaim computing resources in combination with the real-time load situation to ensure that the system can still operate efficiently and stably under resource constraints or load fluctuations.
[0094] In this way, the implementation of the above steps ensures that the computing resource scheduling system can adapt to the changes in real-time task loads, and through actual load monitoring, prediction error analysis, model update, and dynamic resource adjustment, the system can continuously optimize the task scheduling strategy. Compared with traditional static scheduling methods, this method enables the system to quickly adjust computing resources when the task load environment changes through a closed-loop optimization mechanism, reducing task execution delays and improving the overall utilization rate of computing resources. Ultimately, this optimization strategy enables the system to maintain an efficient and stable operating state in a complex, multi-task environment, realizing the intelligent and adaptive allocation of computing resources, thereby enhancing the system's task scheduling ability and computing performance.
[0095] Refer to Figure 7 , Figure 7 is a schematic structural diagram of the computing resource optimization system for analyzing tasks provided by the present invention. The system includes: An acquisition module for acquiring a plurality of analysis tasks and acquiring the characteristic information of each analysis task; A first processing module for determining the priority of each analysis task according to the characteristic information of each analysis task and the current available computing resource status information of the system, and generating a priority queue according to the priorities of all analysis tasks; A second processing module for obtaining the task load prediction value for a future time period based on historical task load data and real-time system status data; A third processing module for allocating analysis tasks with real-time requirements higher than a preset standard to edge nodes for execution according to the type, data scale, and data source location of the analysis tasks, and dynamically adjusting the resource allocation of the edge nodes according to the task load prediction value; A fourth processing module for establishing a unified computing resource pool and dynamically allocating available resources from the computing resource pool according to the priority queue and the task load prediction value to execute multiple parallel analysis tasks.
[0096] In a possible implementation manner, the first processing module is further configured to: Input the feature information and computing resource status information of each analysis task into the priority calculation model to obtain the priority value of each analysis task output by the priority calculation model; among them, the priority calculation model is a deep Q-network model trained based on the deep reinforcement learning algorithm; Determine the priority of each analysis task according to the priority value of each analysis task.
[0097] In a possible implementation manner, the system further includes a training module, which is used for: Randomly initialize the initial weights of the deep Q-network model, and set the initial weights of the target network to the same value as the initial weights of the deep Q-network model; Continuously collect training samples, and construct an experience replay pool according to the training samples; Periodically randomly extract training samples from the experience replay pool, perform batch training on the deep Q-network model, and use the target network to stabilize the training process; Repeat the steps of periodically randomly extracting training samples from the experience replay pool, performing batch training on the deep Q-network model, and using the target network to stabilize the training process until the error of the deep Q-network model is lower than the preset threshold or reaches the predetermined maximum number of training rounds; Take the trained deep Q-network model as the final priority calculation model.
[0098] In a possible implementation manner, the training module is further used for: Periodically randomly extract multiple training samples from the experience replay pool to form a small batch of data; Input the small batch of data into the deep Q-network model to obtain the predicted priority output by the deep Q-network model; Calculate the error between the predicted priority and the target priority label, and update the weights of the deep Q-network model according to the error through the backpropagation algorithm and the gradient descent method; After a preset training step interval, copy the weights of the current deep Q-network model to the target network to reduce training oscillation and accelerate model convergence.
[0099] In a possible implementation manner, the second processing module is further used for: Obtain historical task load data and real-time system status data, and perform missing value filling and normalization processing on the historical task load data and the real-time system status data to obtain the preprocessed historical task load data and the preprocessed real-time system status data; Input the preprocessed historical task load data and the preprocessed real-time system state data into the prediction model to obtain the task load prediction value output by the prediction model; wherein, the prediction model is based on a long short-term memory network or a Transformer network and uses a multi-time step prediction mechanism to generate the task load prediction value for a future time period.
[0100] In a possible implementation manner, the third processing module is further configured to: Based on the type, data scale, and data source location of the analysis task, determine the response latency requirement of the analysis task, and mark the analysis tasks with a response latency higher than the preset real-time threshold as real-time tasks; Allocate the real-time tasks to the edge node closest to the corresponding data source location to shorten the data transmission path and reduce network latency; Combine the CPU, GPU, memory usage of the edge node, and the task load prediction value to dynamically evaluate the available resources of the edge node; According to the evaluation result, allocate or reclaim computing resources for each real-time task allocated to the edge node.
[0101] In a possible implementation manner, the system further includes a fifth processing module, configured to: During the execution of the analysis task, continuously monitor the actual task load value; Compare the actual task load value with the task load prediction value to obtain a deviation value; When it is determined that the deviation value is greater than the preset deviation threshold, update the model parameters of the priority calculation model and the prediction model; According to the updated priority calculation model and the updated prediction model, adjust the priority or resource allocation strategy of the analysis task.
[0102] It should be noted that the computing resource optimization system for analysis tasks provided by the present invention can execute the computing resource optimization method for analysis tasks in any of the above embodiments during specific operation, and this embodiment will not be elaborated here.
[0103] Figure 8 is a schematic structural diagram of an electronic device provided by the present invention, as Figure 8As shown in the figure, the electronic device may include: a processor 810, a communications interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communications interface 820, and the memory 830 complete communication with each other through the communication bus 840. The processor 810 may call the logical instructions in the memory 830 to execute the calculation resource optimization method for analysis tasks. The method includes: obtaining a plurality of analysis tasks and obtaining the characteristic information of each analysis task; determining the priority of each analysis task according to the characteristic information of each analysis task and the current available calculation resource status information of the system, and generating a priority queue according to the priorities of all the analysis tasks; obtaining the predicted task load value for a future time period according to the historical task load data and the real-time system status data; allocating the analysis tasks with real-time requirements higher than the preset standard to the edge nodes for execution according to the type, data scale, and data source location of the analysis tasks, and dynamically adjusting the resource allocation of the edge nodes according to the predicted task load value; establishing a unified calculation resource pool, and dynamically allocating available resources from the calculation resource pool according to the priority queue and the predicted task load value to execute multiple parallel analysis tasks.
[0104] In addition, when the logical instructions in the above-mentioned memory 830 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0105] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer is capable of executing the computational resource optimization method for the analysis tasks provided in the above-mentioned embodiments. The method includes: obtaining a plurality of analysis tasks and obtaining the characteristic information of each analysis task; determining the priority of each analysis task according to the characteristic information of each analysis task and the current available computational resource status information of the system, and generating a priority queue according to the priorities of all the analysis tasks; obtaining the task load prediction value for a future time period according to the historical task load data and the real-time system status data; allocating the analysis tasks with real-time requirements higher than a preset standard to edge nodes for execution according to the type, data scale, and data source location of the analysis tasks, and dynamically adjusting the resource allocation of the edge nodes according to the task load prediction value; establishing a unified computational resource pool, and dynamically allocating available resources from the computational resource pool according to the priority queue and the task load prediction value to execute a plurality of parallel analysis tasks.
[0106] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the computational resource optimization method for the analysis tasks provided in the above-mentioned embodiments. The method includes: obtaining a plurality of analysis tasks and obtaining the characteristic information of each analysis task; determining the priority of each analysis task according to the characteristic information of each analysis task and the current available computational resource status information of the system, and generating a priority queue according to the priorities of all the analysis tasks; obtaining the task load prediction value for a future time period according to the historical task load data and the real-time system status data; allocating the analysis tasks with real-time requirements higher than a preset standard to edge nodes for execution according to the type, data scale, and data source location of the analysis tasks, and dynamically adjusting the resource allocation of the edge nodes according to the task load prediction value; establishing a unified computational resource pool, and dynamically allocating available resources from the computational resource pool according to the priority queue and the task load prediction value to execute a plurality of parallel analysis tasks.
[0107] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative effort.
[0108] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.
[0109] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for optimizing computing resources for an analysis task, characterized in that, Including: Obtain multiple analysis tasks, and obtain the characteristic information of each of the analysis tasks; Determine the priority of each analysis task according to the characteristic information of each analysis task and the current available computing resource status information of the system, and generate a priority queue according to the priorities of all the analysis tasks; Obtain the task load prediction value for a future time period according to historical task load data and real-time system status data; According to the type, data scale, and data source location of the analysis task, allocate the analysis tasks with real-time requirements higher than the preset standard to the edge nodes for execution, and dynamically adjust the resource allocation of the edge nodes according to the task load prediction value; Establish a unified computing resource pool, and dynamically allocate available resources from the computing resource pool according to the priority queue and the task load prediction value to execute multiple parallel analysis tasks.
2. The computational resource optimization method for the analysis task according to claim 1, wherein The step of determining the priority of each analysis task according to the characteristic information of each analysis task and the current available computing resource status information of the system specifically includes: Input the characteristic information of each analysis task and the computing resource status information into a priority calculation model to obtain the priority value of each analysis task output by the priority calculation model; wherein, the priority calculation model is a deep Q-network model trained based on a deep reinforcement learning algorithm; Determine the priority of each analysis task according to the priority value of each analysis task.
3. The computational resource optimization method for the analysis task according to claim 2, characterized in that It also includes the training method of the priority calculation model: Randomly initialize the initial weights of the deep Q-network model, and set the initial weights of the target network to the same values as the initial weights of the deep Q-network model; Continuously collect training samples, and construct an experience replay pool according to the training samples; Periodically randomly extract training samples from the experience replay pool, perform batch training on the deep Q-network model, and use the target network to stabilize the training process; Repeat the step of periodically randomly extracting training samples from the experience replay pool, performing batch training on the deep Q-network model, and using the target network to stabilize the training process until the error of the deep Q-network model is lower than the preset threshold or reaches the predetermined maximum number of training rounds; Use the trained deep Q-network model as the final priority calculation model.
4. The method for optimizing computing resources of an analysis task according to claim 3, characterized in that, The step of periodically randomly extracting training samples from the experience replay pool, performing batch training on the deep Q-network model, and using the target network to stabilize the training process specifically includes: Periodically randomly extract multiple training samples from the experience replay pool to form a mini-batch of data; Input the mini-batch of data into the deep Q-network model to obtain the predicted priority output by the deep Q-network model; Calculate the error between the predicted priority and the target priority label, and update the weights of the deep Q-network model according to the error through the backpropagation algorithm and the gradient descent method; After a preset training step interval, copy the weights of the current deep Q-network model to the target network to reduce training oscillation and accelerate model convergence.
5. The method for optimizing computing resources of the analysis task according to claim 1, characterized in that Obtaining a task load prediction value for a future time period based on historical task load data and real-time system state data specifically includes: Obtaining historical task load data and real-time system state data, and performing missing value filling and normalization processing on the historical task load data and the real-time system state data to obtain preprocessed historical task load data and preprocessed real-time system state data; Inputting the preprocessed historical task load data and the preprocessed real-time system state data into a prediction model to obtain a task load prediction value output by the prediction model; wherein, the prediction model is based on a long short-term memory network or a Transformer network and uses a multi-time step prediction mechanism to generate a task load prediction value for a future time period.
6. The computing resource optimization method for the analysis task according to claim 1, characterized in that According to the type, data scale, and data source location of the analysis task, allocating analysis tasks with real-time requirements higher than a preset standard to edge nodes for execution, and dynamically adjusting the resource allocation of edge nodes according to the task load prediction value specifically includes: Based on the type, data scale, and data source location of the analysis task, determining the response delay requirement of the analysis task, and marking analysis tasks with a real-time requirement higher than a preset real-time threshold as real-time tasks; Allocating the real-time tasks to the edge node closest to the corresponding data source location to shorten the data transmission path and reduce network latency; Combining the CPU, GPU, and memory usage of the edge node and the task load prediction value to dynamically evaluate the available resources of the edge node; According to the evaluation result, allocating or reclaiming computing resources for each of the real-time tasks allocated to the edge node.
7. The method for optimizing computing resources of an analysis task according to claim 1, wherein The method further includes: During the execution of the analysis task, continuously monitoring the actual task load value; Comparing the actual task load value with the task load prediction value to obtain a deviation value; When it is determined that the deviation value is greater than a preset deviation threshold, updating the model parameters of the priority calculation model and the prediction model; According to the updated priority calculation model and the updated prediction model, adjusting the priority or resource allocation strategy of the analysis task.
8. The computational resource optimization method for the analysis task according to claim 1, characterized in that, The feature information is at least one of task type, data scale, and computational complexity; the computing resource status information includes the usage rates of CPU, GPU, and memory.
9. A computing resource optimization system for an analysis task, characterized in that, It includes: An acquisition module for acquiring a plurality of analysis tasks and acquiring the feature information of each analysis task; A first processing module for determining the priority of each analysis task according to the feature information of each analysis task and the current available computing resource status information of the system, and generating a priority queue according to the priorities of all the analysis tasks; A second processing module for obtaining a task load prediction value for a future time period based on historical task load data and real-time system state data; A third processing module for allocating analysis tasks with real-time requirements higher than a preset standard to edge nodes for execution according to the type, data scale, and data source location of the analysis task, and dynamically adjusting the resource allocation of edge nodes according to the task load prediction value; A fourth processing module, configured to establish a unified computing resource pool, and dynamically allocate available resources from the computing resource pool according to the priority queue and the task load prediction value to execute a plurality of parallel analysis tasks.
10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, wherein, When the processor executes the computer program, it implements the computing resource optimization method for the analysis task according to any one of claims 1-8.
11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the computing resource optimization method for the analysis task according to any one of claims 1-8.
12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the computing resource optimization method for the analysis task according to any one of claims 1-8.
Citation Information
Cited By
Intelligent task scheduling method and system combined with artificial intelligence optimization
CN120448077A
Heterogeneous multitask computing power dynamic scheduling method and system
CN120469784A
Transmission scheduling method and system based on intelligent storage
CN120509685A
Industrial defect detection task scheduling method and device, equipment and medium
CN120780489A
A method, apparatus, equipment and medium for scheduling industrial defect detection tasks
CN120780489B