Distributed AI training method and system for processing heterogeneous training task populations

By using the Monte Carlo method and the multi-network physical information neural network method to solve large-scale multi-group average field game models, the optimal resource allocation strategy is generated, and the problem of high computational complexity of traditional numerical methods is solved, and efficient computing resource allocation and load balancing of distributed AI training is realized.

CN119761462BActive Publication Date: 2025-05-13SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510264628.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-05-13
Estimated Expiration
2045-03-07

AI Technical Summary

Technical Problem

When solving large-scale multi-group average field game models, the existing traditional numerical methods have high computational complexity and poor computational efficiency, making it difficult to meet the real-time and efficient requirements of distributed AI training.

Method used

Through fast and accurate algorithms, including the Monte Carlo method and the multi-network physical information neural network method based on network parallelism, large-scale multi-group average field game models are solved, and the optimal resource allocation strategy is generated to ensure efficient allocation of computing resources and node load balancing.

Benefits of technology

It realizes rapid adjustment of computing resources when processing large-scale tasks, adapts to dynamically changing training environments, and improves the computing efficiency and resource utilization of distributed AI training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119761462B_ABST
    Figure CN119761462B_ABST
Patent Text Reader

Abstract

The present invention proposes a distributed AI training method and system for processing heterogeneous training task populations, which belongs to the field of artificial intelligence deep learning, including: real-time collection of data related to distributed AI training; unified processing of the collected data and construction of a task state matrix and a resource state matrix; construction of a large-scale multi-population mean field game model: the training progress of the task, the resource request strategy of the task, the global task allocation and resource usage, and the impact caused by resource competition and synchronization between tasks are used as state variables, control variables, system states, and non-local interactions to construct a large-scale multi-population mean field game model; the above-mentioned large-scale multi-population mean field game model is solved, and an optimal resource allocation strategy is generated based on the solution result to allocate the computing resources required for the task and ensure node load balancing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence deep learning, and in particular to a distributed AI training method and system for processing heterogeneous training task populations. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] With the rapid development of artificial intelligence technology, distributed AI training has attracted widespread attention due to its ability to split and distribute tasks to multiple computing nodes to achieve efficient processing of ultra-large-scale models and data sets, and has become a core component supporting the development of modern AI technology. At the same time, the large-scale multi-population mean field game (MFG) model has become an important tool for solving complex resource allocation problems because it can effectively simulate resource competition and collaboration in multi-task systems.

[0004] However, existing traditional numerical methods have high computational complexity and poor computational efficiency when solving large-scale MFG models, and it is difficult to meet the real-time and high-efficiency requirements of distributed AI training. Summary of the invention

[0005] In order to overcome the shortcomings of the above-mentioned prior art, the present invention provides a distributed AI training method for processing heterogeneous training task populations. The technical solution of the present invention solves the large-scale multi-population MFG model through a fast and accurate algorithm to ensure that the allocation of computing resources can be quickly adjusted when processing large-scale tasks to adapt to the dynamically changing training environment.

[0006] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions:

[0007] In a first aspect, a distributed AI training method for processing a heterogeneous population of training tasks is disclosed, including:

[0008] Collect data related to distributed AI training in real time;

[0009] The collected data is processed uniformly and a task status matrix and a resource status matrix are constructed. The task status matrix represents the resource requirements of the tasks and their interdependencies, and the resource status matrix represents the distribution of available resources of the computing nodes.

[0010] Construct a large-scale multi-population mean field game model: The training progress of the task, the resource request strategy of the task, the global task allocation and resource usage, and the impact caused by resource competition and synchronization between tasks are used as state variables, control variables, system states, and non-local interactions to construct a large-scale multi-population mean field game model;

[0011] The above-mentioned large-scale multi-population mean field game model is solved, and the optimal resource allocation strategy is generated based on the solution results to allocate the computing resources required for the task and ensure node load balancing.

[0012] As a further technical solution, the data related to the distributed AI training includes task status, resource usage, dependencies between tasks, and global system load;

[0013] The task status includes the current progress of the task, the remaining computation amount and the model parameter scale;

[0014] The resource usage includes GPU, memory and network bandwidth availability of each computing node;

[0015] The inter-task dependencies include gradient synchronization requirements and data sharding allocation conditions;

[0016] The global system load includes overall resource occupancy and task queue status.

[0017] As a further technical solution, the collected data are uniformly processed on the central dispatch server, including data cleaning, standardization and normalization operations to eliminate noise and outliers.

[0018] As a further technical solution, the large-scale multi-population mean field game model mentioned above is solved, including:

[0019] The nonlocal terms in the equations of the model are numerically approximated using Monte Carlo methods;

[0020] The approximated multi-population MFG model is solved by using a multi-network physical information neural network method based on network parallelism, and the global equilibrium solution and its corresponding initial distribution in the game process are obtained.

[0021] Generate a task resource allocation strategy based on the obtained global equilibrium solution.

[0022] As a further technical solution, the Monte Carlo method is used to numerically approximate the non-local terms in the equations of the model, specifically including:

[0023] Generate N independent and identically distributed random numbers through any probability density function;

[0024] These random numbers are then used for sampling to obtain independent and identically distributed samples;

[0025] Make an integral approximation.

[0026] As a further technical solution, the approximated multi-population MFG model is solved by using a multi-network physical information neural network method based on network parallelism, specifically including:

[0027] After approximating the non-local terms based on the Monte Carlo method, the approximated large-scale multi-population mean field game model with non-local interactions is expressed as the first equation;

[0028] For the first equation, a multi-network Monte Carlo physical information neural network algorithm based on network parallelism is implemented in combination with initial-terminal and boundary conditions, including: based on network parallelism, M neural networks are constructed, each neural network has the same input and network structure, and the outputs correspond to the equilibrium solutions of different training task populations.

[0029] In the second aspect, a distributed AI training system for processing a population of heterogeneous training tasks is disclosed, including:

[0030] The data collection module is configured to: collect data related to distributed AI training in real time;

[0031] The matrix construction module is configured to: uniformly process the collected data and construct a task status matrix and a resource status matrix, wherein the task status matrix represents the resource requirements of the tasks and their interdependencies, and the resource status matrix represents the available resource distribution of the computing nodes;

[0032] The model building module is configured to: construct a large-scale multi-population mean field game model: the training progress of the task, the resource request strategy of the task, the global task allocation and resource usage, and the impact caused by resource competition and synchronization between tasks are used as state variables, control variables, system states, and non-local interactions to construct a large-scale multi-population mean field game model;

[0033] The solution module is configured to solve the large-scale multi-population mean field game model and generate an optimal resource allocation strategy based on the solution results to allocate the computing resources required for the task and ensure node load balancing.

[0034] One or more of the above technical solutions have the following beneficial effects:

[0035] In order to apply it to distributed AI training in real time and efficiently and realize efficient allocation of computing resources, the technical solution of the present invention solves the large-scale multi-population mean field game model with non-local interactions by quickly and accurately using an efficient machine learning algorithm to ensure that the allocation of computing resources can be quickly adjusted when processing large-scale tasks to adapt to the dynamically changing training environment.

[0036] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0038] Figure 1 This is a diagram showing the working principle of resource allocation in distributed AI training;

[0039] Figure 2 is a flow chart of a solution system executed in the solver of the present invention;

[0040] Figure 3 is a flow chart of a Monte Carlo algorithm in an embodiment of the present invention;

[0041] Figure 4 A structural diagram of a Monte Carlo multi-network physical information neural network in an embodiment of the present invention;

[0042] Figure 5 The present embodiment of the present invention is a method for processing the average field game model of 20 populations with strong repulsive potential to estimate each and Error accuracy diagram of ;

[0043] Figure 6 The present embodiment of the present invention is a method for processing the average field game model of 50 populations with strong repulsive potential to estimate each and Error accuracy diagram of ;

[0044] Figure 7 The example method of this embodiment in the embodiment of the present invention processes the average field game model of 100 populations with strong repulsive potential to estimate each and Error accuracy diagram of ;

[0045] Specifically, is the optimal value function for the rth individual, guiding the individual on how to make decisions to optimize gains or reduce losses; It is the distribution density of the rth population, reflecting the evolution law of individual states in the overall system. DETAILED DESCRIPTION

[0046] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.

[0047] It should be noted that the terms used herein are for describing specific embodiments only and are not intended to be limiting of exemplary embodiments according to the present invention.

[0048] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.

[0049] The mean field game (MFG) system was originally proposed independently by Lasry and Lions in the engineering community to model and analyze the decision-making process involving a large number of indistinguishable rational agents. The agents in the system are assumed to be identical, any single agent has little influence on the outcome of the game, and each agent aims to minimize a certain cost, and the strategy adopted by the agent is affected by the average value of a function of the states of other agents.

[0050] When the actual application environment is not limited to the case of a single homogeneous population, the large-scale multi-population MFG model that describes multiple heterogeneous populations has attracted widespread attention from researchers. The large-scale multi-population MFG model is a type of mathematical model used to study the decision-making behavior of multiple interacting heterogeneous populations in large-scale systems. Compared with the single-population MFG model, this model greatly reduces the dimension and computational complexity of the model by simplifying the complex interactions between a large number of individuals in each population into the interaction between "individuals and the mean field". Therefore, it can better handle the heterogeneity between populations and the dynamic evolution of individual strategies. At the same time, it can also better adapt to the randomness and uncertainty of the external environment. It has the advantages of being able to handle large-scale dynamic systems, simplify computational complexity, and adapt to the heterogeneity of multiple populations, and has become an important tool for analyzing and optimizing the behavior of complex multi-population systems.

[0051] Generally, the large-scale multi-population MFG model derived from the individual optimization problem and the mean field hypothesis can be described by the coupled equations consisting of the Hamilton-Jacobi-Bellman (HJB) equation and the Fokker-Planck (FP) equation. The HJB equation describes the optimal control problem of individuals, and the FP equation is used to describe the evolution of population density. These partial differential equations can effectively describe the strategy evolution of individuals in the population and their interaction with other individuals, thereby revealing the dynamic laws of multi-population systems. At present, the main methods for solving MFG models include analytical methods, numerical iteration methods, and distributed computing methods. However, analytical methods that are only applicable to solving specific idealized models cannot handle complex nonlinear and random dynamic scenarios. In addition, the dynamic evolution of the population and the randomness of strategy selection, strong nonlinearity and heterogeneity of the population make traditional numerical methods face a series of disadvantages in numerical simulation, such as high computational complexity, low efficiency, poor adaptability, insufficient real-time performance, high memory and storage requirements, etc. There is an urgent need to explore more efficient and flexible solution methods. This embodiment provides an unsupervised neural network algorithm to solve the MFG equation. The output is the global equilibrium solution of various tasks, and the execution feedback is the allocation of computing resources based on the equilibrium solution setting.

[0052] Embodiment 1

[0053] This embodiment discloses a distributed AI training method for processing a heterogeneous training task population, including:

[0054] Step 1: Collect key data in real time through the distributed AI training platform, including task status, resource usage, inter-task dependencies, and global system load.

[0055] The task status includes the current progress of the task, the remaining computing amount, and the scale of model parameters; resource usage includes the GPU, memory, and network bandwidth availability of each computing node; inter-task dependencies include gradient synchronization requirements and data sharding allocation; and the global system load includes overall resource occupancy and task queue status.

[0056] Step 2: The collected data is processed uniformly on the central scheduling server, and noise and outliers are eliminated through data cleaning, standardization and normalization operations. The task status matrix and resource status matrix are constructed to represent the resource requirements of the tasks and their interdependencies and the available resource distribution of the computing nodes respectively.

[0057] The task state matrix is ​​T: Assume that there are N tasks in the system, and each task requires K resources (such as GPU, memory, network bandwidth); each row of the task state matrix T corresponds to a task, and each column corresponds to a specific resource requirement: .

[0058] In terms of the representation of matrix elements, percentage, FLOPs that the task needs to execute, and model size are used to describe the current progress of the task, the remaining computing amount, and the scale of model parameters.

[0059] Resource status matrix R: Assume that there are M nodes in the system, and each node provides J types of resources (such as GPU, memory, bandwidth). Each row of the matrix R corresponds to a node, and each column corresponds to the availability of a specific resource type: .

[0060] Similarly, in the form of matrix elements, GPU occupancy, memory remaining amount and node bandwidth remaining amount are used to describe the GPU availability, memory availability and network bandwidth availability of the computing node respectively.

[0061] Subsequent use: After obtaining the balanced solution, prioritize the tasks to the adaptation priority On the highest node. It is obtained by the following formula: ; Indicates the task At the node The matching priority on Indicates the task resource requirements, Representation Node resource distribution.

[0062] Step 3: The task training progress, task resource request strategy, global task allocation and resource usage, and the impact of resource competition and synchronization between tasks are used as state variables, control variables, system state, and non-local interactions to construct the MFG model.

[0063] It should be noted that the training progress, task resource request strategy, global task allocation and resource usage, and the impact caused by resource competition and synchronization between tasks are abstract variables extracted and processed based on the "task status, resource usage, inter-task dependency, and global system load" obtained in step one.

[0064] Training progress: The training progress is derived from the "total computing requirements" and "remaining computing amount" in the task status : ,in, Indicates the task In time training progress; Indicates the task Total computing requirements; Indicates time Time Task Remaining computational requirements.

[0065] The resource request strategy of the task is constructed by combining the task status and resource usage with the specific requirements of the task for computing resources (such as GPU, memory, and bandwidth). The task status provides the computing demand scale of the task, and the resource usage provides the total amount of resources available in the current system: ;in, Indicates the task In time The resource request vector of Indicates the global load status; Indicates the task Priority of function It reflects the dynamic demand of tasks for resources and is generally an empirical formula.

[0066] Global task allocation and resource usage: obtained through comprehensive analysis of global system load and resource usage. Global system load indicates the overall resource occupancy rate and task queue status of the system, reflecting the global situation of current resource allocation; resource usage indicates the specific resource status of each computing node, and optimizes the task allocation strategy in combination with global load.

[0067] Global task allocation strategy: ,in, Indicates time The global task allocation strategy represents the resources allocated to each task. Description Task In time resource allocation strategy; Indicates the task In time Resource request strategy; Indicates at time The available status of system resources; Indicates the task Loss function for resource allocation.

[0068] The impact of resource competition and synchronization between tasks: derived from the analysis of inter-task dependencies and global system load. Inter-task dependencies reflect the interaction between tasks; global system load reflects the pressure of current resource allocation.

[0069] in, , Indicates the task Impacts due to competition or synchronization with other tasks; Indicates the task and tasks The competitive relationship kernel function; For the task The progress of the task; For the task resource requests; Indicates the task The inter-dependency weight.

[0070] Construct MFG: First, initialize the state and control variables. According to the collected data, define the state variables and control variables as and ,in Represents the global computing load and task Task training progress, tasks The remaining computational workload and tasks GPU requests, tasks Memory requests and tasks bandwidth requirements.

[0071] Then use the task and The dependencies between them define the interaction kernel function of non-local terms: , where the first and second terms represent the positive and negative interaction effects, respectively, is the weight parameter, and An infinitesimal quantity.

[0072] Based on the execution goal of the task, define the Hamiltonian function And the cost function: ,in Represents population Corresponding The optimal value function of .

[0073] in addition, ; is the resource request cost, , and They are respectively the computational cost weight of the task, the storage cost weight of the task, and the communication cost weight of the task.

[0074] Finally, based on the above definition, we give Population MFG model: ;

[0075] .

[0076] in Represents population Corresponding The optimal value function of Represents population In time When in state The distribution density of .

[0077] Step 4: An efficient solution algorithm is used for the above large-scale multi-population mean field game (MFG) model. Based on the obtained equilibrium solution, the optimal allocation strategy for computing resources in the distributed AI training system is given, and the resource allocation is dynamically adjusted again based on the environmental feedback after the strategy is given, the computing resources required for the task are allocated and the node load is balanced. Finally, dynamic resource scheduling and global optimization of computing resources are achieved to improve the training efficiency and resource utilization of the system.

[0078] Generation strategy: Nash equilibrium solution based on the above equation and , substituting into the Hamiltonian function to obtain the task The optimal strategy and tasks In Status Specifically, according to optimal control theory, the optimal strategy Determined by the Hamiltonian function, that is: ; Further according to the optimal strategy Generates a resource allocation strategy for tasks.

[0079] Computing resources: Based on middle Assigning tasks Required GPU; memory resources: Based on middle Assigning tasks Memory; Communication resources: According to middle Assigning tasks bandwidth.

[0080] When solving the above step 4, first, the Monte Carlo method is used to numerically approximate the non-local terms in the MFG equation; secondly, the approximated multi-population MFG model is solved with the help of a multi-network physical information neural network method based on network parallelism to obtain the global equilibrium solution and its corresponding initial distribution in the game process; finally, a task resource allocation strategy is generated based on the obtained global equilibrium solution to ensure the fairness and efficiency maximization of tasks in the distributed system.

[0081] The specific embodiment of the present invention considers the problem of numerical solution of large-scale multi-population MFG in distributed AI training with M heterogeneous training task populations. By using the algorithm of the present invention to approximate the equilibrium solution in the large-scale multi-population MFG model, the optimal resource request strategy for each task is predicted, so that resource allocation is dynamically adjusted based on environmental feedback. The working principle is as follows Figure 1 shown.

[0082] The embodiment provides a machine learning algorithm for solving the MFG model with 20, 50 and 100 populations participating in non-local interactions, see Figure 2 The flow chart of the solution process in the solver shown in the figure, the steps of this method are:

[0083] Step 4-1: Consider the following large-scale multi-population MFG model for the case of strong repulsive potential between populations:

[0084] ;

[0085] ;

[0086] in, , is the population number, is the Hamiltonian, , express The set of probability measures with finite first-order moments on , the strong repulsion between different populations is given by express, for dimensional real number field, is the time variable, is a spatial variable, satisfying .

[0087] Heterogeneity is reflected in the upper right subscripts r and q. r and q represent the rth population and the qth population. If r and q are different, there are heterogeneous populations. The number of populations is M. If it is a single population, then M=1; if there are multiple populations, then M>1.

[0088] The above equation has the following initial-terminal conditions and Dirichlet boundary conditions:

[0089] , , ;

[0090] , .

[0091] Step 4-2: For the above coupled model, perform Figure 3 The numerical approximation process shown is as follows:

[0092] First, through any probability density function ,generate Independent and identically distributed random numbers .

[0093] These random numbers are then used for sampling to obtain independent and identically distributed samples .

[0094] Finally The value of The purpose is to simplify the calculation of high-dimensional integrals. The reason is the idea of ​​Monte Carlo approximate integration, that is, through random sampling, the integral is transformed into the average value of the function in the integral domain.

[0095] Step 4-3: The MFG model approximated by the method in step 4-2 is as follows Figure 2 The solver process shown is used to solve the problem, specifically:

[0096] After approximating the nonlocal terms based on the Monte Carlo method, the large-scale multi-population mean field game model with nonlocal interactions can be written as follows:

[0097] ;

[0098] .

[0099] in is the population number, is the number of approximate points of Monte Carlo approximate integration, . For the above equation, combine the following initial-terminal and boundary conditions:

[0100] , , ;

[0101] , .

[0102] Implement a multi-network Monte Carlo physical information neural network algorithm based on network parallelism. The algorithm structure diagram is as follows Figure 4 Specifically, based on network parallelism, construct Each neural network has the same input and network structure, and the output corresponds to the equilibrium solution of different training task populations. ,Right now , is the optimal value function for the rth individual, guiding the individual on how to make decisions to optimize gains or reduce losses; is the distribution density of the rth population, reflecting the evolution law of individual states in the overall system, where . Indicates The output of a neural network. Therefore, the loss function can be written as follows: .

[0103] Specifically, ;

[0104] ;

[0105] .

[0106] in , and denote the mean square error constructed according to the equation, initial-terminal conditions and boundary conditions, respectively. is the population size, Indicates Equations corresponding to functions The known terminal functions of Indicates Equations corresponding to functions The known initial function of , and is the penalty weight, , , , They represent the number of training points for equation penalty, initial condition penalty, final condition penalty, and boundary penalty, respectively, and , , , All are randomly generated by the Sobol sequence. Then the Adam optimizer is used to iteratively optimize the above loss function until the expected error accuracy is achieved.

[0107] An efficient and intelligent algorithm for solving large-scale multi-population MFG problems with non-local interactions. The algorithm can effectively handle the high-dimensional, strongly coupled and large-scale computing challenges brought by the model, reduce the demand for computing resources and time, avoid the "dimensionality disaster" problem caused by the discretization of traditional numerical methods, and provide a reliable idea for numerical simulation of actual large-scale multi-population random control problems. The present invention can effectively improve the efficiency and rationality of distributed AI training to provide resource allocation strategies for each task, ensure that the system is always in the best operating state, and avoid mismatch or bottleneck of system resources due to lagging calculations.

[0108] like Figure 5-Figure 7 As shown, the numerical accuracy of the sub-method of this embodiment in processing the non-local mean field game model of 20, 50 and 100 populations is demonstrated. Figure 5 , Figure 6 , Figure 7 From left to right, the average field game model with 20, 50 and 100 populations estimates the and For the sake of convenience, the error accuracy is divided into , , Arranged into a bar graph for display. From the numerical results, it can be seen that the method of this embodiment is very effective, and the error accuracy is , the calculation time is several hours. It should be noted here that if the multi-objective solution is directly generated as the network output, it is impossible to calculate this system with 20 populations on such a device without the help of network parallelism, let alone solve this system with 100 (or even more) populations. Table 1 shows the average field game model with 20, 50, and 100 populations with strong repulsive potential processed by the sub-method of this embodiment of the present invention. Related error table; Moreover, it can be seen from Table 1 that although the numerical simulation time of the 100 population case is longer, it still has obvious advantages over the current numerical methods, because the existing numerical methods are completely unable to solve such problems due to the dimensionality disaster caused by spatial discretization and large-scale calculations. The method of this embodiment can effectively avoid the dimensionality disaster caused by discretization, and can solve large-scale coupled models in a simple way.

[0109] Table 1 Correlation error table

[0110]

[0111] As the scale of tasks and the number of computing nodes increase, resource scheduling and allocation become more complex, especially in a dynamically changing training environment. Therefore, the large population MFG model that describes the resource scheduling problem in distributed AI training has also become more high-dimensional and complex, with significant non-locality and strong coupling between tasks. For such models, based on the physical information neural network method, the Monte Carlo method is used to approximate non-local features. While avoiding the "curse of dimensionality", it can effectively deal with the strong coupling, strong nonlinearity and high-dimensional challenges of the model, ultimately achieving efficient computing performance while ensuring that the error accuracy remains within an acceptable range.

[0112] The present invention combines the Monte Carlo method with a physical information neural network algorithm, greatly improving the efficiency of solving high-dimensional integrals in machine learning.

[0113] The present invention is based on automatic differentiation of physical information neural networks, and can directly calculate differential derivatives without discretization, avoiding the "curse of dimensionality" problem faced by traditional numerical methods in solving the model, and can effectively deal with the strong coupling, strong nonlinearity and high-dimensional challenges of the model, and can perform large-scale calculations.

[0114] The present invention establishes a multi-network physical information neural network architecture, which is based on network parallelism, greatly reduces the requirements of large-scale output on the complexity of the network structure and reduces the need for computing resources.

[0115] The present invention realizes for the first time the numerical solution of large-scale MFG problems of 100 (or even more) populations with non-local interactions, which has important guiding significance for solving practical multi-population MFG problems.

[0116] Embodiment 2

[0117] The purpose of this embodiment is to provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the program.

[0118] Embodiment 3

[0119] The purpose of this embodiment is to provide a computer-readable storage medium.

[0120] A computer-readable storage medium stores a computer program, which executes the steps of the above method when executed by a processor.

[0121] Embodiment 4

[0122] The purpose of this embodiment is to provide a distributed AI training system for processing heterogeneous training task populations, including:

[0123] The data collection module is configured to: collect data related to distributed AI training in real time;

[0124] The matrix construction module is configured to: uniformly process the collected data and construct a task status matrix and a resource status matrix, wherein the task status matrix represents the resource requirements of the tasks and their interdependencies, and the resource status matrix represents the available resource distribution of the computing nodes;

[0125] The model building module is configured to: construct a large-scale multi-population mean field game model: the training progress of the task, the resource request strategy of the task, the global task allocation and resource usage, and the impact caused by resource competition and synchronization between tasks are used as state variables, control variables, system states, and non-local interactions to construct a large-scale multi-population mean field game model;

[0126] The solution module is configured to solve the large-scale multi-population mean field game model and generate an optimal resource allocation strategy based on the solution results to allocate the computing resources required for the task and ensure node load balancing.

[0127] Embodiment 5

[0128] The purpose of this embodiment is to provide a computer program product containing instructions, which, when running on a computer, enables the computer to execute the methods and functions involved in any of the above embodiments.

[0129] The steps involved in the apparatus of the above embodiment correspond to the method embodiment 1, and the specific implementation method can refer to the relevant description part of embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.

[0130] Those skilled in the art should understand that the modules or steps of the present invention described above can be implemented by a general-purpose computer device, or alternatively, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.

[0131] Although the above describes the specific implementation mode of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.

Claims

1. A distributed AI training method for handling a heterogeneous population of training tasks, characterized by: include: Collect data related to distributed AI training in real time; The collected data is processed uniformly and a task status matrix and a resource status matrix are constructed. The task status matrix represents the resource requirements of the tasks and their interdependencies, and the resource status matrix represents the distribution of available resources of the computing nodes. Construct a large-scale multi-population mean field game model: The training progress of the task, the resource request strategy of the task, the global task allocation and resource usage, and the impact caused by resource competition and synchronization between tasks are used as state variables, control variables, system states, and non-local interactions to construct a large-scale multi-population mean field game model; The above-mentioned large-scale multi-population mean field game model is solved, and the optimal resource allocation strategy is generated based on the solution results to allocate the computing resources required for the task and ensure node load balancing.

2. The distributed AI training method for processing heterogeneous training task populations according to claim 1, characterized in that: The data related to the distributed AI training includes task status, resource usage, dependencies between tasks, and global system load; The task status includes the current progress of the task, the remaining computation amount and the model parameter scale; The resource usage includes GPU, memory and network bandwidth availability of each computing node; The inter-task dependencies include gradient synchronization requirements and data sharding allocation conditions; The global system load includes overall resource occupancy and task queue status.

3. The distributed AI training method for processing heterogeneous training task populations according to claim 1, characterized in that: The collected data are processed uniformly on the central dispatch server, including data cleaning, standardization and normalization operations to eliminate noise and outliers.

4. The distributed AI training method for processing heterogeneous training task populations according to claim 1, characterized in that: Solve the above large-scale multi-population mean field game model, including: The nonlocal terms in the equations of the model are numerically approximated using Monte Carlo methods; The approximated multi-population MFG model is solved by using a multi-network physical information neural network method based on network parallelism, and the global equilibrium solution and its corresponding initial distribution in the game process are obtained. Generate a task resource allocation strategy based on the obtained global equilibrium solution.

5. The distributed AI training method for processing heterogeneous training task populations according to claim 4, characterized in that: The Monte Carlo method is used to numerically approximate the non-local terms in the equations of the model, including: Generate N independent and identically distributed random numbers through any probability density function; These random numbers are then used for sampling to obtain independent and identically distributed samples; Make an integral approximation.

6. The distributed AI training method for processing heterogeneous training task populations according to claim 4, characterized in that: The approximated multi-population MFG model is solved by using a multi-network physical information neural network method based on network parallelism, including: After approximating the nonlocal terms based on the Monte Carlo method, the large-scale multi-population mean field game model with nonlocal interactions is expressed as the first equation; For the first equation, a multi-network Monte Carlo physical information neural network algorithm based on network parallelism is implemented in combination with initial-terminal and boundary conditions, including: based on network parallelism, M neural networks are constructed, each neural network has the same input and network structure, and the outputs correspond to the equilibrium solutions of different training task populations.

7. A distributed AI training system for processing a heterogeneous population of training tasks, characterized by including: The data collection module is configured to: collect data related to distributed AI training in real time; The matrix construction module is configured to: uniformly process the collected data and construct a task status matrix and a resource status matrix, wherein the task status matrix represents the resource requirements of the tasks and their interdependencies, and the resource status matrix represents the available resource distribution of the computing nodes; The model building module is configured to: construct a large-scale multi-population mean field game model: the training progress of the task, the resource request strategy of the task, the global task allocation and resource usage, and the impact caused by resource competition and synchronization between tasks are used as state variables, control variables, system states, and non-local interactions to construct a large-scale multi-population mean field game model; The solution module is configured to solve the large-scale multi-population mean field game model and generate an optimal resource allocation strategy based on the solution results to allocate the computing resources required for the task and ensure node load balancing.

8. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method described in any one of claims 1 to 6 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 6 are performed.

Citation Information

Patent Citations

  • Data offloading rate determination using mean field games

    US20230090549A1

  • Resource scheduling method and apparatus based on computing cluster twin modeling

    WO2025044022A1