Virtualized resource scheduling method and device oriented to high-concurrency AI tasks and storage medium

By building a minimal containerized sandbox and graph neural network scheduling, the resource occupation and collaboration problems in high-concurrency AI tasks are solved, efficient resource scheduling and error recovery are achieved, and system performance and user experience are improved.

CN120276871AActive Publication Date: 2025-07-08董喆
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510766082.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-08
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

The existing technology has problems such as excessive virtualization technology resource utilization, large task startup delay, difficulty in heterogeneous tool collaboration, and shortcomings in operation and maintenance monitoring in high-concurrency AI tasks, which are difficult to meet real-time requirements and efficient resource scheduling.

Method used

Build a minimum containerized sandbox, generate dependency tables based on source code analysis, use graph neural network for intelligent scheduling, monitor and optimize resource allocation in real time, and realize dynamic dependency awareness and error recovery.

Benefits of technology

It realizes the ultra-density scheduling capability of a single physical node, reduces the memory occupancy rate of the task environment, improves system performance, and improves task execution efficiency and error detection recovery speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276871A_ABST
    Figure CN120276871A_ABST
Patent Text Reader

Abstract

The invention provides a high-concurrency AI task oriented virtual resource scheduling method and device and a storage medium. Source code scanning analysis is performed on an application program called by a server side to generate an operation dependency table, and a minimum containerization sandbox for application program operation is constructed based on the dependency table; receiving an AI task input by a user at the cloud end, and inputting parameters of the AI task into the trained task intelligent scheduling model to determine a server for executing the task; and on the determined server for executing the AI task, according to the priority of the AI task, calling a minimum containerization sandbox for executing an application program corresponding to the AI task, and based on the AI task, allocating server resources and loading other components of the application program, the high-concurrency AI task refers to that more than 500 AI tasks are executed on the server at the same time. According to the method, the super-density scheduling capability of the server is realized, the memory occupancy rate of a task environment is greatly reduced, and the system performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of the integration of artificial intelligence and distributed computing, and particularly relates to a virtualized resource scheduling method, device, and storage medium for high-concurrency AI tasks. Background Art

[0002] With the acceleration of the enterprise intelligentization process, AI-driven automated tasks have shown explosive growth. Existing technical solutions face significant challenges in dealing with high-concurrency and cross-platform tasks: 1. Limitations of traditional virtualization technologies: In containerization solutions represented by Docker and Kubernetes, each container needs to load a complete runtime environment, resulting in excessive memory occupation (average ≥ 500MB / container), and it is impossible to support the concurrency of large-scale AI tasks (the concurrency on a single node is usually ≤ 20); the resource isolation mechanism of virtual machine technology generates additional overhead, and the task startup latency reaches the second level (the average startup time in the measured ESXi environment is 3.2 seconds), making it difficult to meet scenarios with high real-time requirements.

[0003] 2. Difficulties in heterogeneous tool collaboration: Cross-platform tasks fail to collaborate due to differences in data interfaces. Industry research shows that the failure rate of cross-system tasks is as high as 32%; existing scheduling systems lack the ability to dynamically perceive tool dependencies, and fixed priority strategies cause resource contention; Shortcomings in operation and maintenance monitoring: Traditional logging systems are difficult to accurately associate fault nodes in distributed task chains, and the average fault location time exceeds 45 minutes; the retry mechanism is simple and crude (full-process rollback), resulting in redundant calculations as high as 73% (data from MIT research). Summary of the Invention

[0004] In view of one or more of the above technical defects in the prior art, the present invention proposes the following technical solutions.

[0005] A virtualized resource scheduling method for high-concurrency AI tasks, the method comprising: A construction step of performing source code scanning and analysis on an application program called by a server end to generate a running dependency table, and constructing a minimum containerized sandbox for the operation of the application program based on the dependency table; A scheduling step of receiving an AI task input by a user in the cloud, and inputting parameters of the AI task into a trained task intelligent scheduling model to determine a server for executing the task; An allocation step of, on the determined server for executing the AI task, calling the minimum containerized sandbox of the application program corresponding to the execution of the AI task according to the priority of the AI task, and allocating server resources based on the AI task and loading other components of the application program, wherein a high-concurrency AI task refers to simultaneously executing more than 500 AI tasks on the server.

[0006] Furthermore, the operation of the building step is as follows: statically analyze the source code of the application to identify the dependencies of the application during runtime, where the dependencies include dynamic link libraries, configuration files, and third-party plugins; execute the application in a simulated sandbox environment and perform dynamic tracing on the application to obtain the resource call path during the execution of the application, where the resource call path includes system calls, network requests, and file reads and writes; generate a dependency table of the application based on the dependencies and the resource call path, and mark the core dependencies and optional dependencies of the application in the dependency table; strip non-essential components during runtime from the base image of the application based on the dependency table; separate the core running components and non-essential components of the application based on a hierarchical building method, optimize the memory occupancy of the core running components based on memory compression technology to build the smallest containerized sandbox of the application, and when the AI task calls the application, by default, only load the core running components of the application in the smallest containerized sandbox.

[0007] Furthermore, the training method of the intelligent scheduling model is as follows: construct a dependency graph of the dependencies between AI tasks, statistically analyze the historical execution logs of the application execution to determine the efficiency of combined execution between applications, construct a graph of a graph neural network based on the dependency graph, where each AI task serves as a node in the graph of the graph neural network, and the set of execution parameters of all applications used to execute one AI task serves as the feature vector of the corresponding node of the AI task. If there is a dependency relationship between two AI tasks, there is an edge between the corresponding nodes of these two AI tasks; perform cleaning processing on the historical data of AI task execution collected and use it as a training sample to train the graph neural network. Stop training after a certain number of training rounds or when the loss function is less than a threshold, and use the trained graph neural network as the intelligent scheduling model, where the feature vector includes: the size of the smallest containerized sandbox of each application, the CPU occupancy rate during application execution, and the execution time.

[0008] Furthermore, the eigenvalue of each node in the graph of the graph neural network is calculated based on the feature vector of the node, and the calculation method is as follows: + + ; Among them, represents the eigenvalue of the i-th node, represents the number of applications available for executing the AI task corresponding to the i-th node, j represents the j-th application for executing the AI, 、 、 respectively represent the minimum sandbox size of the j-th application, the CPU occupancy rate during application execution, and the execution time, 、 、 respectively represent the average minimum sandbox size for the execution of AI tasks by i applications, the average CPU occupancy rate during application execution, and the average execution time, where i ≥ 2, ≥ j ≥ 1.

[0009] Furthermore, the weight of the edge in the graph of the graph neural network is defined as follows: If there is a dependency relationship between the AI tasks corresponding to two nodes, the weight of the edge between these two nodes is 1, otherwise it is 0.

[0010] Furthermore, the operation of allocating server resources based on the AI tasks is as follows: Install a lightweight monitoring program on each server node to monitor the resource usage data of the server in real time. The resource usage data includes the CPU, memory, and network data used by each AI task. Obtain the priority of the AI task, and allocate server resources to the AI task based on the priority.

[0011] Furthermore, the execution of the AI task is monitored in real time, and recovery is performed when an error occurs during execution.

[0012] Furthermore, the specific operations include: Generate a globally unique AI task ID for each AI, perform encryption processing based on the AI task ID, application identifier, and timestamp to generate an operation identifier with a fixed length; Embed the producer AI task ID and operation identifier in the metadata of the data unit; Perform operation traceability when an error occurs during execution: Construct a task DAG through the input-output relationship called by the distributed link tracing record tool; Classify the error: Perform instantaneous error detection through abnormal pattern matching and time window statistics, and mark the detected error as a retryable error. If the error involves data format or permission issues, it is a persistent error and is marked as an error that requires manual intervention. For retryable errors, parse the task DAG, identify the direct upstream and downstream nodes of the faulty node, and only reschedule the affected subgraph in the task DAG, and perform incremental retry by checking point to restore the output data of the upstream node.

[0013] The present invention also proposes a virtualized resource scheduling device for high-concurrency AI tasks. The device includes: A construction unit that performs source code scanning and analysis on the application programs called by the server side to generate a runtime dependency table, and constructs the minimum containerized sandbox for the operation of the application program based on the dependency table; The scheduling unit receives the input AI tasks of the user in the cloud, and inputs the parameters of the AI tasks into the trained task intelligent scheduling model to determine the server for executing the tasks; The allocation unit, on the determined server for executing the AI tasks, calls the minimized containerized sandbox of the application corresponding to the execution of the AI tasks according to the priority of the AI tasks, and allocates server resources and loads other components of the application based on the AI tasks. Among them, high-concurrency AI tasks refer to the situation where more than 500 AI tasks are executed simultaneously on the server.

[0014] Furthermore, the operation of the construction unit is as follows: perform static analysis on the source code of the application to identify the runtime dependencies of the application, where the dependencies include dynamic link libraries, configuration files, and third-party plugins; execute the application in a simulated sandbox environment and perform dynamic tracing on the application to obtain the resource call path during the execution of the application, and the resource call path includes system calls, network requests, and file reads and writes; generate a dependency table of the application based on the dependencies and the resource call path, and mark the core dependencies and optional dependencies of the application in the dependency table; strip the unnecessary components at runtime from the base image of the application based on the dependency table; separate the core running components and unnecessary components of the application based on the hierarchical construction method, optimize the memory occupancy of the core running components based on the memory compression technology to construct the minimized containerized sandbox of the application, and by default, only load the core running components of the application in the minimized containerized sandbox when the AI task calls the application.

[0015] Furthermore, the training method of the intelligent scheduling model is as follows: construct a dependency graph of the dependency relationships between AI tasks, count the historical execution logs of the application executions, determine the efficiency of the combined execution between applications, construct a graph of the graph neural network based on the dependency graph, where each AI task is a node in the graph of the graph neural network, and the set of execution parameters of all applications used to execute one AI task is used as the feature vector of the corresponding node of the AI task. If there is a dependency relationship between two AI tasks, there is an edge between the corresponding nodes of the two AI tasks; perform cleaning processing on the collected historical data of AI task executions and use it as training samples, train the graph neural network, stop training after reaching a certain number of rounds or when the loss function is less than a threshold, and use the trained graph neural network as the intelligent scheduling model, where the feature vector includes: the size of the minimized containerized sandbox of each application, the CPU occupancy rate during the execution of the application, and the execution time.

[0016] Further, the eigenvalue of each node in the graph of the graph neural network is calculated based on the eigenvector of the node, and the calculation method is as follows: + + ; Wherein, represents the eigenvalue of the i-th node, represents the number of application programs available for performing the AI task corresponding to the i-th node, j represents the j-th application program for performing the AI, 、 、 respectively represent the minimized sandbox size, CPU occupancy rate during application execution, and execution time of the j-th application program, 、 、 respectively represent the average minimized sandbox size, average CPU occupancy rate during application execution, and average execution time of application programs for performing the AI task, where i≥2,

[0017] ≥j≥1.

[0017] Further, the weight of the edge in the graph of the graph neural network is defined as: if there is a dependency relationship between the AI tasks corresponding to two nodes, the weight of the edge between these two nodes is 1, otherwise it is 0.

[0018] Further, the operation of allocating server resources based on the AI task is as follows: Install a lightweight monitoring program on each server node to monitor the resource usage data of the server in real time. The resource usage data includes the CPU, memory, and network data used by each AI task, obtain the priority of the AI task, and allocate server resources to the AI task based on the priority.

[0019] Further, the execution of the AI task is monitored in real time, and recovery is performed when an error occurs during execution.

[0020] Furthermore, the specific operations include: generating a globally unique AI task ID for each AI, performing encryption processing based on the AI task ID, application identification, and timestamp to generate an operation identification of a fixed length; embedding the producer AI task ID and the operation identification in the metadata of the data unit; performing operation traceability when an error occurs during execution: constructing a task DAG through the input-output relationships called by the distributed link tracing recording tool; classifying the error: performing instantaneous error detection through abnormal pattern matching and time window statistics, and marking the detected error as a retryable error. If the error involves data format or permission issues, it is a persistent error and is marked as an error requiring manual intervention. For retryable errors, parse the task DAG, identify the direct upstream and downstream nodes of the faulty node, and only reschedule the affected subgraph in the task DAG, and perform incremental retry by restoring the output data of the upstream node through checkpoints.

[0021] The present invention also provides a computer-readable storage medium, on which computer program code is stored, and when the computer program code is executed by a computer, the method described above is executed.

[0022] The technical effect of the present invention is as follows: A virtualized resource scheduling method, device, and storage medium for high-concurrency AI tasks of the present invention. In the construction step S101, source code scanning and analysis are performed on the application program called by the server side to generate a runtime dependency table, and a minimum containerized sandbox for running the application program is constructed based on the dependency table; in the scheduling step S102, an AI task input by a user is received in the cloud, and the parameters of the AI task are input into the trained task intelligent scheduling model to determine the server for executing the task; in the allocation step S103, on the determined server for executing the AI task, the minimum containerized sandbox of the application program corresponding to the execution of the AI task is called according to the priority of the AI task, and server resources are allocated based on the AI task and other components of the application program are loaded. Herein, a high-concurrency AI task refers to the simultaneous execution of more than 500 AI tasks on the server. In the present invention, first, a minimum containerized sandbox is constructed based on the analysis of the source code of the application program, then the server for executing the task is determined based on the parameters of the AI task, then the corresponding sandbox is called on the server, resources are allocated to the AI task based on the priority, and other components of the application program can be loaded according to the needs of the AI task, thereby realizing the ultra-dense scheduling ability of a single physical node (server node), greatly reducing the memory occupancy rate of the task environment, and improving the system performance. Description of the Drawings

[0023] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objects, and advantages of the present application will become more apparent.

[0024] Figure 1 It is a flowchart of a virtualized resource scheduling method for high-concurrency AI tasks according to an embodiment of the present invention.

[0025] Figure 2 It is a structural diagram of a virtualized resource scheduling device for high-concurrency AI tasks according to an embodiment of the present invention. Detailed implementation manners

[0026] The present application will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention, rather than limiting the invention. Additionally, it should be noted that for the sake of description, only parts related to the relevant invention are shown in the drawings.

[0027] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and embodiments.

[0028] Figure 1 A virtualized resource scheduling method for high-concurrency AI tasks of the present invention is shown, and the method includes: A construction step S101, performing source code scanning and analysis on the application program called by the server side to generate a running dependency table, and constructing a minimum containerized sandbox for running the application program based on the dependency table; A scheduling step S102, receiving the input AI task of the user in the cloud, and inputting the parameters of the AI task into the trained task intelligent scheduling model to determine the server for executing the task; An allocation step S103, on the determined server for executing the AI task, calling the minimum containerized sandbox of the application program corresponding to the execution of the AI task according to the priority of the AI task, and allocating server resources based on the AI task and loading other components of the application program, where a high-concurrency AI task refers to simultaneously executing more than 500 AI tasks on the server.

[0029] In the present invention, in order to solve the limitations of traditional virtualization technology and the problem of cooperation with heterogeneous tools, first, source code scanning and analysis are performed on the application programs called by the server side to generate a runtime dependency table. Based on the dependency table, a minimum containerized sandbox for running the application programs is constructed. Among them, the application programs called by the server side are application programs of the local system, also known as heterogeneous tools. A server is connected to multiple local terminals, and one or more local programs are running on the local terminals. The images of the local programs are encapsulated in the sandbox for the server to call. Then, the cloud receives the AI tasks input by the user, and inputs the parameters of the AI tasks into the trained task intelligent scheduling model to determine the server for executing the tasks. The parameters of the AI tasks include execution time requirements, functions to be completed, etc. Finally, on the determined server for executing the AI tasks, the minimum containerized sandbox of the application program corresponding to the execution of the AI tasks is called according to the priority of the AI tasks, and server resources are allocated based on the AI tasks and other components of the application program are loaded. That is, in the present invention, first, a minimum containerized sandbox is constructed based on the analysis of the source code of the application program, then the server for executing the tasks is determined based on the parameters of the AI tasks, then the corresponding sandbox is called on the server, resources are allocated to the AI tasks based on the priority, and other components of the application program can be loaded according to the needs of the AI tasks, thus realizing the ultra-dense scheduling ability of a single physical node (server node) to carry more than 500 concurrent AI tasks, and reducing the memory occupancy of the task environment to less than 10% of the traditional solution. This is an important inventive concept of the present invention.

[0030] In one embodiment, the operation of the construction step S101 is as follows: perform static analysis on the source code of the application program to identify the runtime dependencies of the application program, where the dependencies include dynamic link libraries, configuration files, and third-party plugins; execute the application program in a simulated sandbox environment and perform dynamic tracing on the application program to obtain the resource call path during the execution of the application program, where the resource call path includes system calls, network requests, and file reads and writes; generate a dependency table of the application program based on the dependencies and the resource call path, and mark the core dependencies (must be loaded) and optional dependencies (loaded as needed) of the application program in the dependency table; strip the unnecessary components during runtime from the base image of the application program based on the dependency table; separate the core running components and unnecessary components of the application program based on the hierarchical construction method, and optimize the memory occupancy of the core running components based on the memory compression technology to construct the minimum containerized sandbox of the application program. When the AI task calls the application program, by default, only the core running components of the application program are loaded in the minimum containerized sandbox.

[0031] In the present invention, by performing static analysis on the source code of the application to identify the runtime dependencies of the application and executing the application in a simulated sandbox environment and dynamically tracing the application to obtain the resource call path during the execution of the application, and then generating a dependency table of the application based on the dependencies and the resource call path, the core dependencies of the application, i.e., those that must be loaded, and the optional dependencies, i.e., those loaded on demand, are marked in the dependency table; based on the dependency table, unnecessary components during runtime, such as GUI libraries and debugging tools, are stripped from the base image of the application, the core runtime components of the application are separated from the unnecessary components based on a layered construction method, and the memory occupancy of the core runtime components is optimized based on memory compression technology to build the smallest containerized sandbox for the application, achieving a memory of ≤5 - 100 MB for a single sandbox environment, thus solving the defect in the background art that a single container needs to load a complete runtime environment, resulting in excessive memory occupancy and being unable to support large-scale AI task concurrency (the single-node concurrency is usually ≤20), and enabling the concurrent execution of more than 500 AI tasks under the same hardware, improving the internal performance of the computer server. This is another important inventive point of the present invention.

[0032] In one embodiment, another important inventive concept of the present invention is that after the application is minimized and containerized into a sandbox, the best execution server needs to be determined according to the requirements of the AI task. The present invention determines it based on an intelligent scheduling model. The intelligent scheduling model needs to be trained before use. The training method of the intelligent scheduling model is as follows: construct a dependency graph of the dependency relationships between AI tasks, statistically analyze the historical execution logs of the application execution, determine the efficiency of combined execution between applications, construct a graph of a graph neural network based on the dependency graph, where each AI task is a node in the graph of the graph neural network, and the set of execution parameters of all applications used to execute an AI task is used as the feature vector of the corresponding node of the AI task. If there is a dependency relationship between two AI tasks, there is an edge between the corresponding nodes of the two AI tasks; perform cleaning processing on the collected historical data of AI task execution and use it as a training sample to train the graph neural network. Stop training after a certain number of training rounds or when the loss function is less than a threshold, and use the trained graph neural network as the intelligent scheduling model, where the feature vector includes: the size of the smallest containerized sandbox of each application, the CPU occupancy rate during application execution, and the execution time.

[0033] The present invention trains the intelligent scheduling model based on a graph neural network, constructs a graph of the graph neural network for AI tasks, and during training, cleans the historical data of AI task execution and uses it as training samples for training. The training samples include the efficiency of combined execution between applications, the real-time performance data of the server, the applications that the server can call, the size of the minimized container sandbox of the application, the average time for the application to execute an AI task, and the CPU occupancy rate during application execution. The present invention creatively adds the efficiency of combined execution between applications to the training samples, thereby ensuring that the selected server can maximally improve the execution efficiency of AI tasks. This is another inventive concept of the present invention.

[0034] In one embodiment, the eigenvalue of each node in the graph of the graph neural network is calculated based on the feature vector of the node, and the calculation method is: + + ; Wherein, represents the eigenvalue of the i-th node, represents the number of applications available for executing the AI task corresponding to the i-th node, j represents the j-th application for executing the AI, 、 、 respectively represent the size of the minimized sandbox of the j-th application, the CPU occupancy rate during application execution, and the execution time, 、 、 respectively represent the average minimized sandbox size of the applications for executing the AI task, the average CPU occupancy rate during application execution, and the average execution time. Among them, i≥2, ≥j≥1.

[0035] In one embodiment, the weight of the edge in the graph of the graph neural network is defined as: if there is a dependency relationship between the AI tasks corresponding to two nodes, the weight of the edge between these two nodes is 1, otherwise it is 0.

[0036] In order to train the graph neural network, the present invention proposes a specific setting method for the node eigenvalues and edge weights in the graph. The eigenvalue of the node mainly considers the number of applications for executing the AI task corresponding to the node, and determines the eigenvalue based on the deviation degree between the size of the minimized sandbox of the application, the CPU occupancy rate during application execution, the execution time and the corresponding average values, so that the trained model can select a server suitable for executing the current AI task. This is another important inventive point of the present invention.

[0037] In addition, the model is continuously adjusted during actual operation. For example, if it is found that a certain type of task often fails, relevant resource allocations will be automatically avoided.

[0038] In one embodiment, the operation of allocating server resources based on the AI task is as follows: install a lightweight monitoring program on each server node to monitor the resource usage data of the server in real time. The resource usage data includes the CPU, memory, and network data used by each AI task. Obtain the priority of the AI task, and allocate server resources to the AI task based on the priority. Allocating resources according to priority can be as follows: high-priority tasks (such as payment interfaces): ensure the lowest resource quota, similar to a VIP channel; low-priority tasks (such as data analysis): allow dynamic reduction of resources and even suspension. A pause and resume mechanism is also set up. Low-priority tasks that occupy resources for a long time will be temporarily "frozen" (save the current state to the hard disk) to release resources for emergency tasks; when recovery is required, continue running from the freeze point without starting from scratch. Thus, it solves the defect that the existing scheduling system in the background technology lacks the dynamic perception ability of tool dependency relationships and causes resource contention by adopting a fixed priority strategy. This is another important inventive concept of the present invention.

[0039] In one embodiment, the execution of the AI task is monitored in real time, and recovery is performed when an error occurs during execution. The specific operations include: generating a globally unique AI task ID for each AI, performing encryption processing based on the AI task ID, application program identifier, and timestamp to generate an operation identifier with a fixed length; embedding the producer AI task ID and operation identifier in the metadata of the data unit; performing operation traceability when an error occurs during execution: constructing a task DAG by recording the input-output relationships of tool calls through distributed link tracing; classifying the error: performing instantaneous error detection through abnormal pattern matching and time window statistics, and marking the detected error as a retryable error. If the error involves data format or permission issues, it is a persistent error and is marked as an error that requires manual intervention. For retryable errors, parse the task DAG, identify the direct upstream and downstream nodes of the faulty node, and only reschedule the affected subgraph in the task DAG. Incrementally retry by restoring the output data of the upstream node through checkpoints to avoid full-process rollback. This method is called DNA-style traceability tracking.

[0040] In the present invention, by invoking the input-output relationship through a distributed link tracing recording tool, a task DAG is constructed. Then, the monitored errors are classified into two categories: retryable errors and errors requiring manual intervention. For retryable errors, the task DAG is parsed to identify the direct upstream and downstream nodes of the faulty node. In the task DAG, only the affected subgraph is rescheduled, and the output data of the upstream node is restored through checkpoints for incremental retry, avoiding full-process rollback. This solves the defects in the background art that the traditional logging system is difficult to accurately associate the faulty nodes in the distributed task chain and the full-process rollback is time-consuming, improves the error detection efficiency and recovery speed, and further improves the user experience. This is another important inventive concept of the present invention.

[0041] Figure 2 Fig. shows a virtualized resource scheduling device for high-concurrency AI tasks according to the present invention. The device includes: A construction unit 201 scans and analyzes the source code of the application program called by the server side to generate a running dependency table, and constructs a minimum containerized sandbox for running the application program based on the dependency table; A scheduling unit 202 receives the input AI task of the user in the cloud, and inputs the parameters of the AI task into the trained task intelligent scheduling model to determine the server for executing the task; An allocation unit 203, on the determined server for executing the AI task, calls the minimum containerized sandbox of the application program corresponding to the execution of the AI task according to the priority of the AI task, and allocates server resources based on the AI task and loads other components of the application program. Herein, a high-concurrency AI task refers to the simultaneous execution of more than 500 AI tasks on the server.

[0042] In the present invention, to solve the limitations of traditional virtualization technology and the problem of cooperation with heterogeneous tools, first, source code scanning and analysis are performed on the application programs called by the server side to generate a runtime dependency table, and a minimum containerized sandbox for running the application programs is constructed based on the dependency table. Here, the application programs called by the server side are application programs of the local system, also known as heterogeneous tools. A server is connected to multiple local terminals, and one or more local programs are running on the local terminals. The images of the local programs are encapsulated in the sandbox and can be called by the server. Then, the cloud receives the AI tasks input by the user, and inputs the parameters of the AI tasks into the trained task intelligent scheduling model to determine the server for executing the tasks. The parameters of the AI tasks include execution time requirements, functions to be completed, etc. Finally, on the determined server for executing the AI tasks, the minimum containerized sandbox of the application program corresponding to the execution of the AI tasks is called according to the priority of the AI tasks, and server resources are allocated based on the AI tasks and other components of the application program are loaded. That is, in the present invention, first, a minimum containerized sandbox is constructed based on the analysis of the source code of the application program, then the server for executing the tasks is determined based on the parameters of the AI tasks, then the corresponding sandbox is called on the server, resources are allocated to the AI tasks based on the priority, and other components of the application program can be loaded according to the needs of the AI tasks, thus realizing the ultra-dense scheduling ability of a single physical node (server node) to carry more than 500 concurrent AI tasks, and reducing the memory occupancy of the task environment to less than 10% of the traditional solution. This is an important inventive concept of the present invention.

[0043] In one embodiment, the operation of the construction unit 201 is as follows: perform static analysis on the source code of the application program to identify the runtime dependencies of the application program, where the dependencies include dynamic link libraries, configuration files, and third-party plugins; execute the application program in a simulated sandbox environment and perform dynamic tracing on the application program to obtain the resource call path during the execution of the application program, where the resource call path includes system calls, network requests, and file reads and writes; generate a dependency table of the application program based on the dependencies and the resource call path, and mark the core dependencies (must be loaded) and optional dependencies (loaded as needed) of the application program in the dependency table; strip the unnecessary components during runtime from the base image of the application program based on the dependency table; separate the core runtime components and the unnecessary components of the application program based on a hierarchical construction method, optimize the memory occupancy of the core runtime components based on memory compression technology to construct the minimum containerized sandbox of the application program, and by default, only load the core runtime components of the application program in the minimum containerized sandbox when the AI task calls the application program.

[0044] In the present invention, by performing static analysis on the source code of the application program to identify the runtime dependencies of the application program and executing the application program in a simulated sandbox environment and dynamically tracing the application program to obtain the resource call path during the execution of the application program, and then generating a dependency table of the application program based on the dependencies and the resource call path, in which the core dependencies of the application program, that is, must be loaded, and the optional dependencies, that is, loaded on demand, are marked; based on the dependency table, unnecessary components during runtime, such as GUI libraries and debugging tools, are stripped from the base image of the application program, the core runtime components and the unnecessary components of the application program are separated based on the hierarchical construction method, and the memory occupancy of the core runtime components is optimized based on the memory compression technology to build the smallest containerized sandbox of the application program, so that the memory of a single sandbox environment ≤ 5 - 100MB, thus solving the defect in the background technology that a single container needs to load a complete runtime environment, resulting in too high memory occupancy to support large-scale AI task concurrency (the single-node concurrency is usually ≤ 20), and realizing the concurrent execution of more than 500 AI tasks under the same hardware, improving the internal performance of the computer server, which is another important inventive point of the present invention.

[0045] In one embodiment, another important inventive concept of the present invention is that after the application program is minimized and containerized into a sandbox, it is necessary to determine the optimal execution server according to the requirements of the AI task. The present invention is determined based on an intelligent scheduling model, and the intelligent scheduling model needs to be trained before use. The training method of the intelligent scheduling model is as follows: constructing a dependency graph of the dependency relationships between AI tasks, statistically analyzing the historical execution logs of the application program execution, determining the efficiency of combined execution between application programs, constructing a graph of a graph neural network based on the dependency graph, each AI task being a node in the graph of the graph neural network, and the set of execution parameters of all application programs used to execute one AI task being the feature vector of the corresponding node of the AI task. If there is a dependency relationship between two AI tasks, there is an edge between the corresponding nodes of these two AI tasks; based on the historical data of AI task execution collected and processed after cleaning as training samples, training the graph neural network, stopping training after a certain number of rounds of training or when the loss function is less than a threshold, and using the trained graph neural network as the intelligent scheduling model, where the feature vector includes: the size of the smallest containerized sandbox of each application program, the CPU occupancy rate during application program execution, and the execution time.

[0046] The present invention trains the intelligent scheduling model based on a graph neural network, constructs a graph of the graph neural network for AI tasks, and during training, cleans the historical data of AI task execution as training samples for training. The training samples include the efficiency of combined execution between applications, the real-time performance data of the server, the applications that can be called by the server, the size of the minimized container sandbox of the application, the average time for the application to execute an AI task, and the CPU occupancy rate during application execution. The present invention creatively adds the efficiency of combined execution between applications to the training samples, thereby ensuring that the selected server can maximize the execution efficiency of AI tasks. This is another inventive concept of the present invention.

[0047] In one embodiment, the eigenvalue of each node in the graph of the graph neural network is calculated based on the eigenvector of the node, and the calculation method is as follows: + + ; Wherein, represents the eigenvalue of the i-th node, represents the number of applications available for executing the AI task corresponding to the i-th node, j represents the j-th application for executing the AI, 、 、 respectively represent the size of the minimized sandbox of the j-th application, the CPU occupancy rate during application execution, and the execution time, 、 、 respectively represent the average minimized sandbox size for the applications to execute the AI task, the average CPU occupancy rate during application execution, and the average execution time. Among them, i≥2, ≥j≥1.

[0048] In one embodiment, the weight of the edge in the graph of the graph neural network is defined as follows: If there is a dependency relationship between the AI tasks corresponding to two nodes, the weight of the edge between these two nodes is 1, otherwise it is 0.

[0049] In order to train the graph neural network, the present invention proposes a specific setting method for the node eigenvalues and edge weights in the graph. The eigenvalue of the node mainly considers the number of applications for executing the AI task corresponding to the node, and determines the eigenvalue based on the deviation degree of the minimized sandbox size, CPU occupancy rate, and execution time of the application from the corresponding average values, so that the trained model can select a server suitable for executing the current AI task. This is another important inventive point of the present invention.

[0050] In addition, the model is continuously adjusted during actual operation. For example, if it is found that a certain type of task often fails, relevant resource allocation will be automatically avoided.

[0051] In one embodiment, the operation of allocating server resources based on the AI task is as follows: Install a lightweight monitoring program on each server node to monitor the resource usage data of the server in real time. The resource usage data includes the CPU, memory, and network data used by each AI task. Obtain the priority of the AI task, and allocate server resources to the AI task based on the priority. Allocating resources according to priority can be as follows: High-priority tasks (such as payment interfaces): Ensure the lowest resource quota, similar to a VIP channel; Low-priority tasks (such as data analysis): Allow dynamic reduction of resources or even suspension. A pause and resume mechanism is also set up. Low-priority tasks that occupy resources for a long time will be temporarily "frozen" (save the current state to the hard disk) to release resources for emergency tasks; when recovery is needed, continue running from the freeze point without starting from scratch. This solves the defect in the prior art scheduling system in the background technology that lacks the dynamic perception ability of tool dependencies and causes resource contention by adopting a fixed priority strategy. This is another important inventive concept of the present invention.

[0052] In one embodiment, the execution of the AI task is monitored in real time and recovered when an error occurs during execution. The specific operations include: Generate a globally unique AI task ID for each AI, and perform encryption processing based on the AI task ID, application program identifier, and timestamp to generate an operation identifier with a fixed length; Embed the producer AI task ID and the operation identifier in the metadata of the data unit; Perform operation traceability when an error occurs during execution: Construct a task DAG by recording the input-output relationship of tool calls through a distributed link tracing record; Classify the error: Perform instantaneous error detection through abnormal pattern matching and time window statistics, and mark the detected error as a retryable error. If the error involves data format or permission issues, it is a persistent error and is marked as an error that requires manual intervention. For retryable errors, parse the task DAG, identify the direct upstream and downstream nodes of the faulty node, and only reschedule the affected subgraph in the task DAG. Incrementally retry by restoring the output data of the upstream node through a checkpoint to avoid full-process rollback. This method is called DNA-style traceability tracking.

[0053] In the present invention, the input-output relationship called by the distributed link tracing record tool is used to construct a task DAG. Then, the monitored errors are classified into two categories: retryable errors and errors requiring manual intervention. For retryable errors, the task DAG is parsed to identify the direct upstream and downstream nodes of the faulty node. In the task DAG, only the affected subgraph is rescheduled, and the output data of the upstream node is restored through checkpoints for incremental retry, avoiding full-process rollback. This solves the defects in the background technology that the traditional logging system is difficult to accurately associate the faulty nodes in the distributed task chain and the full-process rollback is time-consuming, improves the error detection efficiency and recovery speed, and further improves the user experience. This is another important inventive concept of the present invention.

[0054] In one embodiment of the present invention, a computer storage medium is proposed. A computer program is stored on the computer storage medium, and when the computer program on the computer storage medium is executed by a processor, the above method is implemented. The computer storage medium can be a hard disk, DVD, CD, flash memory, etc.

[0055] For the convenience of description, the above device is described by dividing it into various units according to functions. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware. From the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute the device described in each embodiment or some parts of the embodiments of the present application.

[0056] Finally, it should be noted that the above embodiments are only used to illustrate rather than limit the technical solution of the present invention. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that the present invention can still be modified or equivalently replaced. Any modification or partial replacement without departing from the spirit and scope of the present invention shall be covered by the scope of the claims of the present invention.

Claims

1. A virtualized resource scheduling method for high-concurrency AI tasks, characterized in that, The method includes: A construction step of performing source code scanning and analysis on an application for server - side call to generate a runtime dependency table, and constructing a minimum containerized sandbox for the application to run based on the dependency table; A scheduling step of receiving an AI task input by a user in the cloud, and inputting the parameters of the AI task into a trained task intelligent scheduling model to determine the server for executing the task; An allocation step of, on the determined server for executing the AI task, calling the minimum containerized sandbox of the application corresponding to the AI task according to the priority of the AI task, and allocating server resources and loading other components of the application based on the AI task, where a high - concurrency AI task refers to executing more than 500 AI tasks simultaneously on the server.

2. The method according to claim 1, wherein The operation of the construction step is: performing static analysis on the source code of the application to identify the runtime dependencies of the application, where the dependencies include dynamic link libraries, configuration files, and third - party plugins; executing the application in a simulated sandbox environment and dynamically tracing the application to obtain the resource call path during the execution of the application, where the resource call path includes system calls, network requests, and file reads and writes; generating a dependency table of the application based on the dependencies and the resource call path, and marking the core dependencies and optional dependencies of the application in the dependency table; Stripping unnecessary runtime components from the base image of the application based on the dependency table; Separating the core running components and non - essential components of the application based on a hierarchical construction method, optimizing the memory occupancy of the core running components based on memory compression technology to construct the minimum containerized sandbox of the application. When an AI task calls the application, by default, only the core running components of the application are loaded in the minimum containerized sandbox.

3. The method according to claim 2, characterized in that, The training method of the intelligent scheduling model is: constructing a dependency graph of the dependency relationships between AI tasks, statistically analyzing the historical execution logs of the application executions to determine the efficiency of combined executions between applications, constructing a graph of a graph neural network based on the dependency graph, with each AI task as a node in the graph of the graph neural network, and the set of execution parameters of all applications for executing one AI task as the feature vector of the corresponding node of the AI task. If there is a dependency relationship between two AI tasks, there is an edge between the corresponding nodes of the two AI tasks; performing cleaning processing on the collected historical data of AI task executions and using it as training samples to train the graph neural network. Stop training after a certain number of training rounds or when the loss function is less than a threshold, and use the trained graph neural network as the intelligent scheduling model, where the feature vector includes: the size of the minimum containerized sandbox of each application, the CPU occupancy rate during application execution, and the execution time.

4. The method according to claim 3, characterized in that, The operation of allocating server resources based on the AI task is as follows: install a lightweight monitoring program on each server node to monitor the resource usage data of the server in real time. The resource usage data includes the CPU, memory, and network data used by each AI task. Obtain the priority of the AI task, and allocate server resources to the AI task based on the priority.

5. The method according to claim 4, characterized in that, Monitor the execution of the AI task in real time and perform recovery when an error occurs. Specifically, it includes: generating a globally unique AI task ID for each AI, performing encryption processing based on the AI task ID, application identifier, and timestamp to generate an operation identifier with a fixed length; embedding the producer AI task ID and operation identifier in the metadata of the data unit; performing operation traceability when an error occurs: constructing a task DAG through the input-output relationship called by the distributed link tracing record tool; classifying the error: performing instantaneous error detection through abnormal pattern matching and time window statistics, and marking the detected error as a retryable error. If the error involves data format or permission issues, it is a persistent error and is marked as an error requiring manual intervention. For retryable errors, parse the task DAG, identify the direct upstream and downstream nodes of the faulty node, and only reschedule the affected subgraph in the task DAG. Perform incremental retry by checking the output data of the upstream node through checkpoint recovery.

6. A virtualized resource scheduling device for high-concurrency AI tasks, characterized in that, The device includes: A construction unit that performs source code scanning and analysis on the application program called by the server side to generate a runtime dependency table, and constructs a minimum containerized sandbox for the application program to run based on the dependency table. A scheduling unit that receives the AI task input by the user in the cloud, and inputs the parameters of the AI task into the trained task intelligent scheduling model to determine the server that executes the task. An allocation unit that, on the server determined to execute the AI task, calls the minimum containerized sandbox of the application program corresponding to the execution of the AI task according to the priority of the AI task, allocates server resources based on the AI task, and loads other components of the application program. Among them, a high-concurrency AI task refers to the simultaneous execution of more than 500 AI tasks on the server.

7. The device according to claim 6, characterized in that, The operation of the construction unit is as follows: perform static analysis on the source code of the application program to identify the runtime dependencies of the application program. The dependencies include dynamic link libraries, configuration files, and third-party plugins; execute the application program in a simulated sandbox environment and perform dynamic tracing on the application program to obtain the resource call path during the execution of the application program. The resource call path includes system calls, network requests, and file reads and writes; generate a dependency table for the application program based on the dependencies and resource call paths, and mark the core dependencies and optional dependencies of the application program in the dependency table. Strip unnecessary components during runtime from the base image of the application program based on the dependency table. Separate the core running components of the application from the non-essential components based on a hierarchical construction method, optimize the memory occupancy of the core running components based on memory compression technology to build the smallest containerized sandbox of the application, and when the AI task calls the application, by default, only load the core running components of the application in the smallest containerized sandbox.

8. The device according to claim 7, characterized in that, The training method of the intelligent scheduling model is as follows: construct a dependency graph of the dependencies between AI tasks, count the historical execution logs of the application to determine the efficiency of combined execution between applications, construct a graph of a graph neural network based on the dependency graph, where each AI task is a node in the graph of the graph neural network, and the set of execution parameters of all applications used to execute an AI task is used as the feature vector of the corresponding node of the AI task. If there is a dependency relationship between two AI tasks, there is an edge between the corresponding nodes of these two AI tasks; perform cleaning processing on the collected historical data of AI task execution and use it as training samples to train the graph neural network. Stop training after a certain number of training rounds or when the loss function is less than a threshold, and use the trained graph neural network as the intelligent scheduling model, where the feature vector includes: the size of the smallest containerized sandbox of each application, the CPU occupancy rate during application execution, and the execution time.

9. The device according to claim 8, wherein Perform real-time monitoring on the execution of the AI task and perform recovery when an error occurs during execution.

10. A computer storage medium, characterized in that, A computer program is stored on the computer storage medium, and when the computer program on the computer storage medium is executed by a processor, the method according to any one of claims 1-5 is implemented.

Citation Information

Patent Citations

  • Scheduling task execution method and device, electronic equipment and storage medium

    CN111104212A

  • Method and system for realizing containerization of application program

    CN113821219A

  • Task processing method of AI task and distributed system

    CN113961353A

  • Containerized resource management method based on dynamic graph network

    CN118550652A

  • Enterprise task allocation method and device

    CN119151190A