Dynamic allocation system for computing power resources of AI (Artificial Intelligence) server
By designing the AI server's computing power resource dynamic allocation system, including resource data platform and multi-module working together, the problem that traditional computing power cannot meet the dynamic task resource scheduling of AI servers is solved, and more efficient resource dynamic scheduling capabilities are achieved.
Patent Information
- Application Number
- CN202510715749.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-06-27
AI Technical Summary
Traditional computing power cannot effectively meet the dynamic analysis and processing of AI servers for a variety of complex dynamic tasks. There is resource scheduling delay and lack of scheduling flexibility, and the inability to accurately schedule AI server resources, and the ability to dynamic scheduling is lacking.
A dynamic allocation system for computing power resources for AI servers is designed, including resource data platform, task data acquisition module, resource dynamic division module and resource allocation module. By generating resource load views in real time, dynamically aligning resource priorities, building a dynamic resource matching model, and obtaining decision resources for allocation based on this model.
It effectively improves the ability of AI servers to dynamically schedule resources, can schedule resources more accurately, meet the needs of diverse and complex dynamic tasks, reduce resource scheduling delays, and improve scheduling flexibility.
Smart Images

Figure CN120216150A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computing power scheduling, and specifically to a dynamic allocation system for computing power resources of an AI server. Background Art
[0002] An AI server is a server specifically designed for artificial intelligence applications. It adopts a heterogeneous hardware architecture, usually equipped with acceleration chips such as GPUs, FPGAs, and ASICs, and uses the combination of CPUs and acceleration chips to meet the requirements of high-throughput interconnection, providing powerful computing power support for artificial intelligence application scenarios such as natural language processing, computer vision, and machine learning, and supporting the training and inference processes of AI algorithms. With the development of today's digital age, various information technologies have developed rapidly, and the amount of data has shown an explosive growth; whether it is a research institution conducting complex data analysis, an enterprise carrying out large-scale business operations, or emerging artificial intelligence applications emerging continuously; at this time, the AI server that provides powerful computing power support for artificial intelligence application scenarios such as natural language processing, computer vision, and machine learning needs to exert its powerful computing power to meet diverse and complex dynamic task processing. However, traditional computing power cannot meet the requirements of the AI server for dynamically analyzing and processing resources for diverse and complex dynamic tasks; there may be resource scheduling delays and a lack of scheduling flexibility for diverse and complex tasks; furthermore, it may not be able to meet the dynamically changing requirements, so that it is impossible to accurately schedule the resources of the AI server and lacks the ability of dynamic scheduling; therefore, in order to solve the above technical problems, the present invention provides a dynamic allocation system for computing power resources of an AI server. Summary of the Invention
[0003] In order to solve the above technical problems, the present invention provides a dynamic allocation system for computing power resources of an AI server. The object of the present invention can be achieved by the following technical solutions: A dynamic allocation system for computing power resources of an AI server, including a resource data platform, which is connected with a task data acquisition module, a resource dynamic division module, and a resource allocation module. The resource data platform is used to map the AI server resources corresponding to the AI server and generate a resource load view in real time. The task data acquisition module is provided with a hierarchical position area, which is used to obtain the historical tenant level according to the hierarchical position area, set a distributed tenant network according to the historical tenant level, and collect the task data corresponding to each tenant in the distributed tenant network. The resource dynamic partitioning module is provided with a resource priority unit and a decision-making generation unit; the resource priority unit is used to dynamically arrange the resource priorities of the AI server according to the resource load view to construct a resource dynamic matching model; the decision-making generation unit is used to obtain the latest location node processing queue corresponding to the distributed tenant network in real time; and obtain the decision-making resources corresponding to each task data based on the resource dynamic matching model. The resource allocation module is used to allocate the AI server resources according to the optimal decision-making resources.
[0004] Furthermore, the process of the resource data platform generating the resource load view in real time includes: Obtain the resource load corresponding to the AI server resources, and obtain the resource load data corresponding to the resource load. Set the dynamic resource load data threshold corresponding to the resource load data, and then obtain the resource load data ratio of each resource load data to the dynamic resource load data threshold, and generate the corresponding resource load data pie chart. Set the data priority weights of each resource load data, and calculate the weighted average value of the resource load data ratio corresponding to the resource load through the resource load data pie chart and the corresponding data priority weights, which is recorded as the resource load rate corresponding to the resource load. Mutually map the AI server, AI server resources, resource load, resource load data, resource load data pie chart, and resource load rate to generate a resource load view.
[0005] Furthermore, the process of the task data collection module setting the hierarchical position area includes: Set the AI server as the central node, and set the hierarchical position area corresponding to the central node; the hierarchical position area includes the first hierarchical position area, the second hierarchical position area, and the third hierarchical position area; set the corresponding tenant location information according to the hierarchical position area, which are the central tenant, the edge tenant, and the remote tenant respectively; and set a location collection node at the central node for collecting the tenant location information of the task data corresponding to the AI server.
[0006] Furthermore, the process of obtaining the historical tenant level according to the hierarchical position area includes: Schedule the historical task data corresponding to the AI server, and the historical task data includes the historical task type, the historical task execution duration, and the historical resource load data set. Collect the historical tenant location information corresponding to each historical task data through the collection node, and generate a historical tenant location node; then generate the corresponding historical central tenant location node, historical edge tenant location node, and historical remote tenant location node according to the historical tenant location information for collecting the corresponding task data. Set the position influence coefficient corresponding to the level position area and set the level weight corresponding to the task type; and then obtain the historical tenant level corresponding to the historical tenant location node according to the position influence coefficient, level weight, and historical task execution duration. That is, the specific formula is: ; Among them, R i represents the historical tenant level corresponding to the i-th historical tenant location node; w iL and t iL respectively represent the level weight and historical task execution duration of the task type L corresponding to the i-th historical tenant location node; a i represents the position influence coefficient corresponding to the i-th historical tenant location node.
[0007] Furthermore, the process of setting the distributed tenant network according to the historical tenant level includes: Obtain the maximum value of the frequency of the historical resource load data set corresponding to each historical tenant location node, and mark it as the important resource load data set corresponding to the historical tenant location node; and integrate it with the historical tenant level and map it to the corresponding historical tenant location node; then, the historical tenant location node is connected to the corresponding central node in terms of position to generate a historical tenant location information network. Generate existing tenant location nodes for the existing tenants corresponding to the AI server, and generate level tenant location nodes according to the level position area, namely the central existing tenant location node, the edge existing tenant location node, and the remote existing tenant location node, and connect them to the central node in terms of position respectively to generate an existing tenant location information network. The historical tenant location information network and the existing tenant location information network are merged to generate a distributed tenant network.
[0008] Furthermore, the process of the resource priority unit constructing a resource dynamic matching model includes: According to the resource load data in the resource load data pie chart corresponding to the resource load view, perform a dynamic arrangement of the resource load data from small to large to obtain the resource priority usage queue of the resource load data corresponding to the AI server. Divide the resource priority usage queue into three sub-queues on average, and mark the three sub-queues as the optimal usage sub-queue, the available usage sub-queue, and the waiting usage sub-queue respectively according to the order of the resource priority usage queue. The sub - queues are first self - paired to respectively generate the optimal resource combination group, the available resource combination group, and the waiting - to - use resource combination group; then, according to different sub - queues, two or more resource - load data are freely combined to generate a free resource combination group, and further, a resource dynamic combination model is generated according to the arrangement order of the optimal resource combination group, the available resource combination group, the waiting - to - use resource combination group, the free resource combination group, and the important resource - load data set.
[0009] Further, the process by which the decision - making generation unit obtains the processing queue of the latest position node in real - time includes: Obtain the task data of the corresponding position node of the distributed tenant network at the same time point in real - time. The task data includes the task type and the task quantity; set the resource demand coefficient corresponding to the task quantity, and then calculate the product of the corresponding level weight and the resource demand coefficient of the task data to obtain the task level of the corresponding task data; then, arrange the task data in descending order of the task level to obtain the task data processing queue. Obtain the sum of the task levels corresponding to the task data processing queue, and mark it as the position node task level corresponding to the position node; then, arrange the position nodes in descending order of the position node task level to obtain the position node processing queue. Judge the types of the tenants corresponding to each position node in the position node processing queue. If it is a historical tenant position node, obtain the historical tenant level and multiply it by the position node task level to obtain the task scheduling level of the corresponding position node. If it is an existing tenant position node, obtain the position influence coefficient corresponding to the level tenant position node and multiply it by the position node task level to obtain the task scheduling level of the corresponding position node. Schedule the corresponding position nodes in the position node processing queue according to the task scheduling level to obtain the latest position node processing queue.
[0010] Further, the process of obtaining the optimal decision - making resources based on the resource dynamic combination model includes: According to the task types of the task data corresponding to the task data processing queue of the corresponding position node in the latest position node processing queue, schedule the corresponding combination groups one by one according to the arrangement order of the resource dynamic combination model, and mark them as the optimal decision - making resources.
[0011] Further, the process by which the resource allocation module allocates the AI server resources according to the optimal decision - making resources includes: Send the optimal decision - making resources to the central node to allocate the AI server resources and then process the task data.
[0012] Compared with the prior art, the beneficial effects of the present invention are: 1. The present invention maps the AI server resources corresponding to the AI server and generates a resource load view in real time; a hierarchical position area is set, the historical tenant level is obtained according to the hierarchical position area, a distributed tenant network is set according to the historical tenant level, and the task data corresponding to each tenant in the distributed tenant network is collected; the historical tenant location information network corresponding to the AI server is obtained, which is beneficial to predicting future demands based on historical tenant behaviors; and the existing tenant location information network is obtained, which is beneficial to comprehensively covering the tenants corresponding to the AI server.
[0013] 2. The present invention dynamically arranges the priorities of the AI server resources according to the resource load view to construct a resource dynamic matching model; and obtains the latest position node processing queue corresponding to the distributed tenant network in real time; furthermore, based on the resource dynamic matching model, the decision-making resources corresponding to each task data are obtained; the AI server resources are allocated according to the optimal decision-making resources; effectively improving the ability of the AI server for dynamic resource scheduling. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings.
[0015] Figure 1 It is the schematic diagram of the present invention.
[0016] Figure 2 It is the schematic diagram of the resource load view of the present invention.
[0017] Figure 3 It is the schematic diagram of the resource load data pie chart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0019] As Figure 1 shown, a dynamic computing power resource allocation system for an AI server includes a resource data platform, and the resource data platform is connected to a task data collection module, a resource dynamic division module, and a resource allocation module; The resource data platform is used to map the AI server resources corresponding to the AI server and generate a resource load view in real time; The task data acquisition module is provided with a hierarchical position area, which is used to obtain the historical tenant level according to the hierarchical position area, set a distributed tenant network according to the historical tenant level, and collect the task data corresponding to each tenant in the distributed tenant network; The resource dynamic partitioning module is provided with a resource priority unit and a decision generation unit; the resource priority unit is used to dynamically arrange the priorities of AI server resources according to the resource load view to construct a resource dynamic matching model; the decision generation unit is used to obtain the latest position node processing queue corresponding to the distributed tenant network in real time; and obtain the decision resources corresponding to each task data based on the resource dynamic matching model; The resource allocation module is used to allocate the AI server resources according to the optimal decision resources.
[0020] As Figure 2 shown, it should be further noted that the process of the resource data platform mapping the AI server resources corresponding to the AI server and generating the resource load view in real time includes: Obtain the computing resources, storage resources, and network resources corresponding to the AI server resources; and obtain the resource loads corresponding to the AI server resources, which are CPU / GPU load, memory load, and transmission load respectively; Obtain the resource load data corresponding to the resource load; the resource load data corresponding to the CPU / GPU load includes CPU / GPU utilization rate, computing frequency, and the number of waiting computing tasks; the resource load data corresponding to the memory load includes memory usage rate, memory bandwidth utilization rate, and hard disk utilization rate; the resource load data corresponding to the transmission load includes bandwidth utilization rate and link utilization rate; Set the dynamic resource load data threshold corresponding to the resource load data, and then obtain the resource load data ratio of each resource load data to the dynamic resource load data threshold, and generate a corresponding resource load data pie chart, as Figure 3 shown; judge the load level of the resource load data ratio, and obtain the corresponding warning load color; If the resource load data ratio is less than or equal to 30% of the dynamic resource load data threshold, the load level is marked as the first-level load, and the corresponding warning load color is marked as green; If the resource load data ratio is greater than 80% of the dynamic resource load data threshold, the load level is marked as the third-level load, and the corresponding warning load color is marked as red; otherwise, the load level is marked as the second-level load, and the corresponding warning load color is marked as yellow; Set the data priority weights of each resource load data, and calculate the weighted average value of the resource load data ratio corresponding to the resource load through the resource load data pie chart and the corresponding data priority weights, which is recorded as the resource load rate corresponding to the resource load; Map the AI server, AI server resources, resource load, resource load data, resource load data pie chart, and resource load rate to each other to generate a resource load view.
[0021] In the above embodiments, it should be further noted that the proportion of the resource load data is the resource load data / dynamic resource load data threshold; for example, for CPU / GPU load, the CPU / GPU utilization rate, computing frequency, and the number of waiting computing tasks corresponding to the resource load data are 42%, 40% of the lowest computing frequency, and 5 respectively; the corresponding dynamic resource load data thresholds are 84%, 80% of the lowest computing frequency, and 20 respectively; then the proportions of the resource load data are 1 / 2, 1 / 2, and 1 / 4 respectively; if the corresponding data priority weights are 0.6, 0.3, and 0.1 respectively; then the corresponding resource load rate is 0.6×42% + 0.3×40% + 0.1×5 = 0.872; the dynamic resource load data threshold can be automatically adjusted according to historical data; for example, if it is the peak period of the training task, the CPU / GPU utilization rate is temporarily increased from 42% to 68% to avoid false warnings; where the historical data is not specifically described here.
[0022] It should be further noted that the process of the task data acquisition module for obtaining the historical tenant level includes: Generate a central node with the AI server, and set the level position area corresponding to the central node; the level position area includes the first level position area, the second level position area, and the third level position area; set the corresponding tenant location information according to the level position area, which are the central tenant, the edge tenant, and the remote tenant respectively; and set a location acquisition node at the central node for acquiring the tenant location information of the task data corresponding to the AI server. In the above embodiments, it should be further noted that the task types include but are not limited to AI deep learning training, video, computing, large-scale processing of data sets, databases, logs, storage and retrieval, data transmission, etc.; the historical task execution duration is used to represent all the historical durations when the AI server executes historical tasks; where the first level position area, the second level position area, and the third level position area are used to represent the area deployed within the AI server, the edge area, and the remote area respectively. Schedule the historical task data corresponding to the AI server, where the historical task data includes the historical task type, the historical task execution duration, and the historical resource load data set. Collect the historical tenant location information corresponding to each historical task data through the collection nodes, and generate historical tenant location nodes; then generate corresponding historical central tenant location nodes, historical edge tenant location nodes, and historical remote tenant location nodes according to the historical tenant location information for collecting corresponding task data. It should be further noted that the historical resource load data set is used to represent the cluster of resource load data of the AI server used when the AI server executes historical tasks; for example, CPU / GPU utilization rate, computing frequency, and memory usage rate; then the corresponding resource load data group is (CPU / GPU utilization rate, computing frequency, memory usage rate). Set the location influence coefficient corresponding to the level location area and set the level weight corresponding to the task type; then obtain the historical tenant level corresponding to the historical tenant location node according to the location influence coefficient, level weight, and historical task execution duration. That is, the specific formula is: ; Among them, Ri represents the historical tenant level corresponding to the i-th historical tenant location node; wiL and tiL respectively represent the level weight and historical task execution duration of the task type L corresponding to the i-th historical tenant location node; ai represents the location influence coefficient corresponding to the i-th historical tenant location node.
[0023] It should be further noted that the process of setting the distributed tenant network according to the historical tenant level includes: Obtain the maximum value of the frequency of the historical resource load data set corresponding to each historical tenant location node, and mark it as the important resource load data set of the corresponding historical tenant location node; and integrate it with the historical tenant level and map it to the corresponding historical tenant location node; then connect the historical tenant location node with the corresponding central node in terms of location to generate a historical tenant location information network. Generate existing tenant location nodes for the existing tenants corresponding to the AI server, and generate level tenant location nodes according to the level location area, namely central existing tenant location nodes, edge existing tenant location nodes, and remote existing tenant location nodes, and connect them with the central node in terms of location respectively to generate an existing tenant location information network. The historical tenant location information network and the existing tenant location information network are merged to generate a distributed tenant network.
[0024] In the above embodiment, it should be further noted that the existing tenants are used to represent the tenants that the AI server is currently executing; obtaining the historical tenant location information network corresponding to the AI server is beneficial to predicting future demands based on historical tenant behaviors; and obtaining the existing tenant location information network is beneficial to comprehensively managing the tenants corresponding to the AI server.
[0025] It should be further noted that the process of the resource priority unit dynamically arranging the resource priorities of the AI server resources according to the resource load view to construct a resource dynamic matching model includes: Dynamically arranging the resource load data in the resource load data pie chart corresponding to the resource load view from small to large to obtain the resource priority usage queue of the AI server corresponding to the resource load data; Dividing the resource priority usage queue into three sub-queues on average, and respectively marking the three sub-queues as the optimal usage sub-queue, the available usage sub-queue, and the waiting usage sub-queue according to the order of the resource priority usage queue; Let the sub-queues be self-matched first to generate the optimal resource matching group, the available resource matching group, and the waiting resource matching group respectively; then freely combine two or more resource load data according to different sub-queues to generate a free resource matching group, and then arrange them in the order of the optimal resource matching group, the available resource matching group, the waiting resource matching group, the free resource matching group, and the important resource load data set to generate a resource dynamic matching model.
[0026] In the above embodiment, it should be further noted that the resources corresponding to the resource load data are preferentially arranged according to the resource load data ratio of the AI server corresponding to the resource load data; the larger the resource load data ratio, the smaller the corresponding resource usage. Therefore, arranging the resource load data ratio corresponding to the resource load data from small to large, the priority of the corresponding resource usage can be obtained, and then the resource priority usage queue can be obtained; the resource priority usage queue corresponding to the resource dynamic matching model changes dynamically according to the resource load data ratio corresponding to the resource load view, and then the corresponding matching groups change dynamically; this process better divides the priority levels of the resources corresponding to the AI server and better prepares for resource scheduling.
[0027] It should be further noted that the process of the decision generation unit obtaining the latest location node processing queue corresponding to the distributed tenant network in real time includes: Obtaining the task data of the location nodes corresponding to the distributed tenant network at the same time point in real time, where the task data includes the task type and the task quantity; setting a resource demand coefficient corresponding to the task quantity, and then calculating the product of the level weight corresponding to the task data and the resource demand coefficient to obtain the task level corresponding to the task data; then arranging the task data in descending order of the task level to obtain the task data processing queue; Obtaining the sum of the task levels corresponding to the task data processing queue, and marking it as the location node task level corresponding to the location node; then arranging the location nodes in descending order of the location node task level to obtain the location node processing queue; Judge the types of tenants corresponding to each position node in the position node processing queue; If it is a historical tenant position node, obtain the historical tenant level, and multiply it by the position node task level to obtain the task scheduling level of the corresponding position node; If it is an existing tenant position node, obtain the position influence coefficient corresponding to the level tenant position node, and multiply it by the position node task level to obtain the task scheduling level of the corresponding position node; Schedule the corresponding position nodes in the position node processing queue according to the task scheduling level to obtain the latest position node processing queue.
[0028] In the above embodiment, it should be further noted that if the position node task levels corresponding to the position node processing queue are 1, 0.8, 0.72, 0.52, 0.3; and the task scheduling level is 0.83, then the latest position node processing queue is 1, 0.83, 0.8, 0.72, 0.52, 0.3.
[0029] It should be further noted that the process of obtaining the optimal decision-making resources corresponding to each task data based on the resource dynamic matching model includes: Schedule the corresponding matching groups one by one according to the task types of the task data corresponding to the task data processing queue of the position node corresponding to the latest position node processing queue in the arrangement order of the resource dynamic matching model, and mark them as the optimal decision-making resources.
[0030] In the above embodiment, it should be further noted that if the data type of the task data corresponding to the task data processing queue corresponding to the position node is AI deep learning training, the scheduled matching group should be the one with the smallest average of the resource load data corresponding to the CPU / GPU load, memory load, and transmission load; if it is a log, the scheduled matching group should be the one with the smallest resource load data corresponding to the memory load.
[0031] It should be further noted that the process of the resource allocation module allocating AI server resources according to the optimal decision-making resources includes: Send the optimal decision-making resources to the central node to allocate AI server resources and then process the task data.
[0032] If the warning load color mark is red, generate a warning signal.
[0033] The features and exemplary embodiments of various aspects of the present application will be described in detail below. For the purpose of making the objectives, technical solutions and advantages of the present application more clearly understood, the present application will be further described in detail below in combination with the accompanying drawings and specific embodiments; it should be understood that the specific embodiments described herein are only intended to explain the present application, rather than limiting the present application; for those skilled in the art, the present application can be implemented without some of these specific details; the above description of the embodiments is only to provide a better understanding of the present application by showing examples of the present application.
[0034] The above embodiments are only used to illustrate the technical methods of the present invention rather than to limit. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical methods of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical methods of the present invention.
Claims
1. A computing power resource dynamic allocation system for an AI server, including a resource data platform, characterized in that, The resource data platform is connected to a task data collection module, a resource dynamic partitioning module, and a resource allocation module; The resource data platform is used to map the AI server resources corresponding to the AI server and generate a resource load view in real time; The task data collection module is provided with a hierarchical location area, which is used to obtain the historical tenant level according to the hierarchical location area, set a distributed tenant network according to the historical tenant level, and collect the task data corresponding to each tenant in the distributed tenant network; The resource dynamic partitioning module is provided with a resource priority unit and a decision generation unit; the resource priority unit is used to dynamically arrange the priorities of the AI server resources according to the resource load view to construct a resource dynamic matching model; the decision generation unit is used to obtain the latest location node processing queue corresponding to the distributed tenant network in real time; And obtain the optimal decision-making resources corresponding to each task data based on the resource dynamic matching model; The resource allocation module is used to allocate the AI server resources according to the optimal decision-making resources.
2. The computing power resource dynamic allocation system of an AI server according to claim 1, characterized in that The process of the resource data platform generating a resource load view in real time includes: Obtain the resource load corresponding to the AI server resources, and obtain the resource load data corresponding to the resource load; Set the dynamic resource load data threshold corresponding to the resource load data, and then obtain the resource load data ratio of each resource load data to the dynamic resource load data threshold, and generate a corresponding resource load data pie chart; Set the data priority weight of each resource load data, and calculate the weighted average value of the resource load data ratio corresponding to the resource load through the resource load data pie chart and the corresponding data priority weight, which is recorded as the resource load rate corresponding to the resource load; Mutually map the AI server, AI server resources, resource load, resource load data, resource load data pie chart, and resource load rate to generate a resource load view.
3. The dynamic allocation system for computing power resources of an AI server according to claim 2, characterized in that, The process of the task data collection module setting the hierarchical location area includes: Generate the central node by the AI server, and set the hierarchical location area corresponding to the central node; the hierarchical location area includes the first hierarchical location area, the second hierarchical location area, and the third hierarchical location area; set the corresponding tenant location information according to the hierarchical location area, which are the central tenant, the edge tenant, and the remote tenant respectively; and set a location collection node at the central node, which is used to collect the tenant location information of the task data corresponding to the AI server.
4. The computing power resource dynamic allocation system of an AI server according to claim 3, characterized in that, The process of obtaining the historical tenant level according to the hierarchical location area includes: Schedule the historical task data corresponding to the AI server, and the historical task data includes the historical task type, the historical task execution duration, and the historical resource load data set; Collect the historical tenant location information corresponding to each historical task data through the collection node, and generate a historical tenant location node; and then generate the corresponding historical central tenant location node, historical edge tenant location node, and historical remote tenant location node according to the historical tenant location information, which are used to collect the corresponding task data; Set the position influence coefficient corresponding to the level position area and set the level weight corresponding to the task type; then obtain the historical tenant level corresponding to the historical tenant location node according to the position influence coefficient, level weight, and historical task execution duration.
5. The computing power resource dynamic allocation system of an AI server according to claim 4, characterized in that, The process of setting up the distributed tenant network according to the historical tenant level includes: Obtain the maximum value of the frequency of the historical resource load data set corresponding to each historical tenant location node, and mark it as the important resource load data set corresponding to the historical tenant location node; integrate it with the historical tenant level and map it to the corresponding historical tenant location node; then connect the historical tenant location node with the corresponding central node in terms of location to generate the historical tenant location information network; Generate the existing tenant location nodes corresponding to the existing tenants of the AI server, and generate the level tenant location nodes according to the level position area, namely the central existing tenant location node, the edge existing tenant location node, and the remote existing tenant location node, and connect them with the central node in terms of location respectively to generate the existing tenant location information network; Merge the historical tenant location information network and the existing tenant location information network to generate the distributed tenant network.
6. The computing power resource dynamic allocation system of an AI server according to claim 2, characterized in that, The process of the resource priority unit constructing the resource dynamic matching model includes: Obtain the resource priority usage queue of the resource load data corresponding to the AI server by performing a dynamic arrangement of the resource load data in ascending order according to the proportion of the resource load data in the resource load data pie chart corresponding to the resource load view; Divide the resource priority usage queue into three sub-queues on average, and mark the three sub-queues as the optimal usage sub-queue, the available usage sub-queue, and the waiting usage sub-queue respectively according to the order of the resource priority usage queue; Let the sub-queues match themselves first to generate the optimal resource matching group, the available resource matching group, and the waiting resource matching group respectively; then perform a free combination of two or more resource load data according to different sub-queues to generate the free resource matching group, and then arrange them in the order of the optimal resource matching group, the available resource matching group, the waiting resource matching group, the free resource matching group, and the important resource load data set to generate the resource dynamic matching model.
7. The dynamic allocation system for computing power resources of an AI server according to claim 1, characterized in that The process of the decision-making generation unit obtaining the latest position node processing queue in real time includes: Obtain the task data of the position nodes corresponding to the distributed tenant network at the same time point in real time, and the task data includes the task type and the number of tasks; set the resource demand coefficient corresponding to the number of tasks, and then calculate the product of the level weight and the resource demand coefficient corresponding to the task data to obtain the task level corresponding to the task data; then arrange the task data in descending order of the task level to obtain the task data processing queue; Obtain the sum of the task levels corresponding to the task data processing queue, and mark it as the position node task level corresponding to the position node; then arrange the position nodes in descending order of the position node task level to obtain the position node processing queue; Judge the types of the tenants corresponding to each position node in the position node processing queue; If it is a historical tenant location node, obtain the historical tenant level and multiply it by the position node task level to obtain the task scheduling level corresponding to the position node; If it is an existing tenant location node, obtain the location influence coefficient corresponding to the hierarchical tenant location node, and multiply it by the location node task level to obtain the task scheduling level of the corresponding location node; Schedule the corresponding location nodes in the location node processing queue according to the task scheduling level to obtain the latest location node processing queue.
8. The computing power resource dynamic allocation system of an AI server according to claim 7, characterized in that The process of obtaining the optimal decision-making resources based on the resource dynamic matching model includes: Schedule the corresponding matching groups one by one according to the arrangement order of the resource dynamic matching model for the task types of the task data corresponding to the task data processing queue of the location nodes corresponding to the latest location node processing queue, and mark them as the optimal decision-making resources.
9. The dynamic allocation system for computing power resources of an AI server according to claim 8, characterized in that The process of the resource allocation module allocating AI server resources according to the optimal decision-making resources includes: Send the optimal decision-making resources to the central node to allocate AI server resources and then process the task data.
Citation Information
Patent Citations
System, device and process for dynamic tenant structure adjustment in distributed resource management system
CN109565515A
Distributed multi-tenant data security isolation system and method
CN119402233A
Multi-tenant management method and system based on container and Kubernetes
CN119668886A
Distributed management system and method for cloud container cluster
CN119728592A
Geographically distributed data center resource scheduling method
CN120011028A
Cited By
Server leasing data maintenance method and application system
CN120415926A