Two-stage strategy gradient optimization automatic parallelization method based on self-adaptive Fimanban
Through the adaptive Tumamba neural network and perturbation dual-stage policy gradient optimization algorithm, the lack of feature representation and strategy optimization of device placement algorithms in large-scale neural network training is solved, and efficient device parallel strategy generation is achieved, which significantly improves training efficiency and strategy search speed.
Patent Information
- Application Number
- CN202510356490.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-04
AI Technical Summary
The existing device placement algorithm has insufficient computing graph feature representation capabilities and defects in optimization stability of reinforcement learning strategies in large-scale neural network training, resulting in the inability to obtain the optimal parallel strategy.
Adaptive Tumamba neural network is used to construct a weight adaptive residual structure, combined with the perturbation two-stage policy gradient optimization algorithm, and through dynamic weight adaptive residual connection and controllable perturbation noise generation, the device placement strategy is optimized to achieve global optimization and local adaptability.
It effectively alleviates the problem of excessive smoothing of traditional graph neural networks, improves the global optimality and local adaptability of device placement strategies, and significantly reduces single-round iteration training time and strategy search time.
Smart Images

Figure CN120258044A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of device placement in the distributed training of large neural networks of AI models, and particularly relates to a two-stage policy gradient optimization automatic parallel method based on adaptive Tumanba. Background Art
[0002] In recent years, with the rapid development of complex tasks such as image classification, speech recognition, and machine translation, the number of parameters of deep neural network models has shown an exponential growth trend. During the model training and inference processes, the dual pressures of computing requirements and storage capacity have increasingly highlighted the performance bottleneck of a single hardware device, forcing the training of large-scale artificial intelligence models to rely on a multi-device collaborative computing architecture. This technical challenge has promoted innovative research on distributed parallelization strategies. Among them, the device placement algorithm, as the core technology for optimizing computing resource allocation, is becoming a key research direction for breaking through the bottleneck of large-scale training.
[0003] Traditional device placement algorithms mainly adopt an evolutionary path that combines the manual design paradigm and automated search techniques in parallel. Traditional manual design methods rely on domain experts to customize parallel schemes for specific model architectures. Although they show high efficiency in early simple architectures such as convolutional neural networks, their technical limitations are significantly exposed when facing complex models such as Transformers with parameter scales exceeding tens of billions and Mixture of Experts (MoE). Domain experts often need to spend hundreds of hours on manual analysis and construct parallel schemes through means such as parsing computational graph dependencies, estimating memory occupancy, and simulating communication overheads. Especially in new computing paradigms such as dynamic sparse computing and mixed-precision training, manual design methods are not only inefficient but also difficult to guarantee the optimality of the strategy.
[0004] To address the above technical challenges, the research focus has gradually shifted to automated policy search techniques. This technical system mainly includes two core methods: data parallelism and model parallelism. Data parallelism realizes a linear expansion of data throughput through a multi-device synchronous parameter update mechanism; model parallelism relies on techniques such as operator partitioning and pipeline parallelism to decompose the model computing task into multiple sub-task modules and deploy them to heterogeneous computing devices through intelligent scheduling algorithms. The collaborative optimization of these two types of techniques drives the evolution of automated parallel frameworks and provides basic support for the training of ultra-large-scale models.
[0005] In the field of automated policy search, the combination of reinforcement learning algorithms and graph encoding techniques constitutes the mainstream technical route. Existing technical solutions, such as ColocRL, hierarchical device placement method (HDP), and other reinforcement learning-based device placement methods, although they can improve policy search performance, still have the inherent defect of too long training cycles. Solutions such as Placeto and GraphSAGE that use Markov decision process (MDP) modeling, although they optimize the decision-making path through the state transition mechanism, are limited by the temporal independence assumption of MDP, resulting in insufficient effective utilization of historical state information. It is worth noting that improved solutions such as Post and Spolight based on the proximal policy optimization (PPO) algorithm, although they improve training stability through importance sampling and policy constraint mechanisms, their continuous action space optimization paradigm has an inherent conflict with the high-dimensional discrete characteristics of the device placement scenario, easily leading to pipeline resource competition problems among multiple GPUs.
[0006] The representation defects of graph encoding techniques further exacerbate the policy deviation problem: existing solutions generally have the common problem of missing physical topology information. Typically, systems such as Placeto and Trinity do not incorporate the physical location of devices (such as the cross-rack topological distance) into the feature space, resulting in ineffective modeling of cross-rack communication delays; the random neighbor sampling strategy adopted by GraphSAGE shows an exponential increase in the probability of losing key features when the model complexity increases; while the position encoding mechanisms designed by solutions such as P-GNN and Aware, although they attempt to enhance the graph structure representation through the aggregation of the features of the farthest nodes, are affected by the locality bias of message passing neural networks (MPNNs), and their global topological feature expression ability still shows a significant degradation in ultra-large-scale models. A deeper technical defect is that existing graph encoding schemes generally ignore the semantic feature differences at the operator level (such as the local computational characteristics of convolutional operations and the global communication requirements of fully connected layers), ultimately resulting in the inability to obtain a relatively optimal parallel policy. Summary of the Invention
[0007] In view of the insufficient ability of the computational graph feature representation in the existing device placement algorithm and the defect of the stability optimization of the reinforcement learning strategy, a two-stage policy gradient optimization automatic parallel method based on adaptive Graph Mamba is proposed. First, in terms of computational graph feature representation, a weight-adaptive residual Graph Mamba neural network architecture is constructed. With the Graph Mamba neural network as the basic feature extraction module, accurate modeling of the long-range dependence relationship between cross-level nodes is realized. An innovative dynamic weight adaptive residual connection mechanism is proposed. By introducing a learnable weight parameter matrix, the feature smoothing problem caused by multi-layer stacking in traditional graph neural networks is alleviated, and the autonomous balance between local fine-grained features and long-range features is achieved. Then, the obtained computational graph features are input into the perturbed two-stage policy gradient optimization algorithm. In its internal update stage, by injecting controllable perturbation noise, a differential device placement strategy is constructed, and the adaptability of the strategy to the dynamic environment is improved based on the proximal policy optimization mechanism (PPO). Subsequently, in the external update stage, global policy optimization based on cumulative reward is implemented. A reward function is constructed with the device execution time constraint as the goal, and long-term value estimation is performed by calculating the difference in policy execution time to ensure the global optimality of the policy. Finally, the optimal device parallel strategy is obtained.
[0008] The present invention provides a two-stage policy gradient optimization automatic parallel method based on adaptive Graph Mamba, and the steps are as follows:
[0009] Step 1: Obtain a publicly available AI model dataset, perform operator fusion on each computational graph G in the dataset, and obtain the feature matrix of the computational graph where n represents the number of nodes of the computational graph G, and f is the feature dimension of the nodes in the computational graph G.
[0010] Step 2: Take X as the initial encoding matrix X(0) and input it into the Graph Mamba neural network. Generate the hidden encoding H(t) through the SSM mechanism in the Graph Mamba neural network, and combine with X(t) to output the global encoding matrix M(t). Dynamically fuse M(t), X(t), and X(0) using the residual structure to obtain the feature encoding X(t + 1) at time t + 1. Iterate the above process T times to obtain the final feature matrix X(T), and generate features through a multi-layer perceptron (MLP)
[0011] Step 3: Input into the perturbed two-stage policy gradient algorithm. For each training cycle, j device placement strategies are generated through perturbation noise in the internal stage and update the parameters involved in the process of generating the perturbation noise. In the external stage, perform policy evaluation, select the strategy with the minimum training time as O k , and perform global parameter update. After K cycles, output the optimal device placement strategy O.
[0012] The beneficial effects of the present invention are as follows:
[0013] (1) The weight-adaptive residual graph Mamba neural network effectively alleviates the over-smoothing problem in traditional graph neural networks and dynamically balances global and local node features through a weight-adaptive residual structure to enrich graph encoding. Specifically, the graph Mamba neural network solves the limited receptive field of traditional graph neural networks by effectively extracting node features and capturing long-term dependencies. The weight-adaptive residual module effectively balances feature information in different dimensions and alleviates over-smoothing through the weight-adaptive residual structure.
[0014] (2) The perturbed two-stage policy gradient optimization algorithm uses an internal and external two-stage update mechanism to optimize the device placement strategy. The internal update stage focuses on short-term node-level optimization and enhances policy expression through perturbed noise to generate different and effective device placement items. The external update stage evaluates the performance of the long-term policy to ensure that the policy converges to the global optimal solution. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 is the overall architecture of the two-stage policy gradient optimization automatic parallel method based on the adaptive graph Mamba;
[0016] Figure 2 is the overall architecture of the weight-adaptive residual graph Mamba neural network;
[0017] Figure 3 is the overall architecture of the perturbed two-stage policy gradient optimization algorithm. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] The following will further illustrate the present invention in conjunction with the drawings and specific implementation steps:
[0019] Figure 1 Shown is the overall process of a two-stage policy gradient optimization automatic parallel method based on the adaptive graph Mamba. An AI model is transformed into a graph G(V, E), where V represents the set of computational graph operators (nodes), and the operator v i ∈V represents a single operator (such as matrix multiplication, convolution, etc.); E represents the set of directed edges between nodes, and e i,j ∈E represents the data communication dependency between operators i and j. Given D = {d1, d2, d3, …, d m} representing the available device resources, where d i ∈D represents a certain computing device (such as CPU or GPU). Given the computing resources D and the computational graph G, find the mapping O (i.e., the device placement strategy), where the device placement item o i ∈O such that each operator v i corresponds to a device d i. The ultimate goal is to find an optimal device placement strategy O such that when the operators in the computation graph G are placed according to the strategy O, the single-round iterative training time F(G, O) of the computation graph G is minimized.
[0020] Specifically: First, take the neural network computation graph G as the input, and use the weight-adaptive residual graph Mamba neural network to obtain the feature matrix Then, traverse each node v i in the node feature representation x i , and obtain the device placement item o through the perturbation-based two-stage policy gradient optimization reinforcement learning algorithm i , and according to o i assign v i to the device d i . During this period, through internal and external parameter updates, finally obtain the device placement strategy O that minimizes the single-round iterative training time F(G, O).
[0021] Step 1: Obtain the publicly available AI model dataset, perform operator fusion on each computation graph G in the dataset, and obtain the feature matrix of the computation graph where n represents the number of nodes in the computation graph G, and f is the feature dimension of the nodes in the computation graph G. Specifically as follows:
[0022] Obtain the publicly available AI model dataset, perform operator fusion on each computation graph G in the AI model dataset through the ColocRL algorithm. The fused computation graph contains the relevant attributes of all operators in the computation graph, including operator type, input, output, and access identifier. Construct the original feature vector of the operator, that is, the node encoding initialization, and obtain the complete feature matrix X of the computation graph. n represents the number of nodes in the computation graph G, and f is the feature dimension of the nodes in the computation graph G.
[0023] Step 2: Take X as the initial encoding matrix X(0) and input it into the Mamba neural network. Generate the hidden encoding H(t) through the SSM mechanism, and combine with X(t) to output the global encoding matrix M(t). Dynamically fuse M(t), X(t), and X(0) using the residual structure, iterate T times to obtain the final feature matrix X(T), and generate The overall process is as Figure 2 shown.
[0024] Step 2.1: At time t, input the encoding matrix X(t) into the selective state management (SSM) mechanism of the Mamba neural network to generate the hidden encoding matrix and combine H(t) with X(t) to output the global encoding matrix
[0025] Step 2.1.1: Introduce five learnable weight matrices and Perform linear transformations on the input encoding matrix X(t) respectively to obtain the projection parameters and the discrete step size
[0026] A(t) = W A(t) X(t), B(t) = W B(t) X(t), C(t) = W c(t) X(t),
[0027] D(t) = W D(t) X(t), Δ(t) = W Δ(t) X(t)
[0028] Step 2.1.2: Apply the Softplus activation function to Δ(t) to obtain Ensure its non-negativity; then, based on A(t), B(t) and Calculate the discretized evolution parameters and
[0029] Step 2.1.3: Based on and X(t), obtain the hidden encoding matrix H(t) at time t. By multiplying C(t) with , obtain the contribution of the hidden encoding H(t) to the global encoding. At the same time, multiply D(t) with X(t) to obtain the contribution of the input feature X(t) to the global encoding. Then, add these two contributions to obtain the global encoding matrix M(t) at the current time:
[0030]
[0031] Step 2.2: Concatenate M(t) and X(t) column-wise, and generate the weight dynamic adaptive factor through the learnable weight and the bias term after passing through the Sigmoid activation function. Use α(t) to perform weighted fusion on M(t) and X(t) to obtain the transformed encoding matrix
[0032] α(t) = Sigmoid[(M(t))‖X(t) × W α (t) + b α (t)]
[0033] P(t) = (1 - α(t)) × M(t) + α(t) × X(t)
[0034] Step 2.3: Combine P(t) with the initial encoding matrix X(0) through residual connection to generate the encoding matrix X(t + 1) at time t + 1:
[0035] X(t + 1) = P(t) + X(0)
[0036] Step 2.4: Repeat the above Steps 2.1, 2.2, and 2.3 until the preset number of iterations T is reached to obtain the final feature matrix Input X(T) into the multi-layer perceptron MLP to obtain the final feature matrix
[0037] Step 3: Input into the perturbed two-stage policy gradient algorithm. For each training epoch, in the internal stage, generate j device placement strategies through perturbed noise and update the internal parameters; in the external stage, perform policy evaluation, select the strategy with the minimum training time as O k , and perform global parameter update. After K epochs, output the optimal device placement strategy O. The overall network framework is as Figure 3 shown, and the overall algorithm process is as shown in Algorithm 1.
[0038]
[0039]
[0040]
[0041] Step 3.1: In each training epoch k, in the internal stage, perform j iterations to generate j device placement strategies Among them, the device placement strategy is obtained by traversing the node features in to obtain the node status and adding controllable perturbed noise to obtain the perturbed node status Finally, obtain the device placement item of the node The device placement strategies of all nodes constitute a device placement strategy In addition, based on perform parameter gradient update in the internal stage
[0042] Step 3.1.1: In the internal update stage of the algorithm, traverse each node feature in the feature representation Process it through a fully connected layer to obtain the node status where m represents the resource quantity of available devices. The current node status Add controllable perturbation noise to obtain the perturbed node state Subsequently Obtain the probability distribution of the device placement generated by node i Device placement item Device placement probability Calculate the clipping loss through these parameters Perform gradient update on the parameters in the internal stage.
[0043] Step 3.1.1.1: Process the features of the node Through the fully connected layer to obtain the node state Take the current moment state And the previous moment latent state Through the GRU network, generate the node state And the latent state Hidden state Take the mean to obtain Obtain the standard deviation through exponential operation And based on Introduce random noise Finally generate the perturbed state And through the hyperparameters And Take the weighted combination of the node state And the perturbed state To obtain the node state Subsequently, take Through the fully connected layer and Softmax activation function calculation, obtain the probability distribution of device placement
[0044] Step 3.1.1.2: According to the obtained probability distribution of device placement Obtain the device placement probability And the device placement item Subsequently, these three parameters are in the internal update stage, and the parameters in Step 3.1.1.1 are gradient updated. Specifically:
[0045] According to the obtained Take the logarithm to obtain the device placement probability Take its average maximum value as the device placement item of the current node And for Take the mean to obtain To reflect the central tendency of the node on all possible device placement items. Subsequently, take Perform exponential calculation to obtain the device placement ratio By calculating And Calculate the advantage value through the sparse Softmax cross-entropy loss function And the learnable variable baseline value Calculate the probability value through the Sigmoid function Through And Multiply to get the surrogate loss Finally And Get the clipped loss And based on Update the gradients of the parameters in step 3.1
[0046] Step 3.1.2: Place the device placement items obtained in the internal update phase In the device placement policy in sequence And based on Calculate the training time required for a single iteration In addition, by traversing the internal phase J times, J device placement policies are obtained, and J execution times F are obtained therefrom, constituting the training time set
[0047] Step 3.2: The external phase of the algorithm determines the optimal device placement policy O in the current training cycle based on the j strategies generated in the internal phase of step 3.1 k . After K training cycles, obtain the device placement policy with convergence, that is, the minimum training time for a single iteration
[0048] Step 3.2.1: In the external update phase of the algorithm, select the minimum value from the training time set As the minimum training time at the current moment The difference between And Between adjacent training cycles is used as the reward value
[0049] Step 3.2.2: Multiply the reward value R k By the discount factor And add it to the historical discounted return L k-1 To get the current moment's discounted return Use L k Perform gradient update to optimize the global parameters, including the parameters of the Tumanba neural network in step 2 and the parameters of the perturbed two-stage policy gradient algorithm in step 3. The optimal device placement policy O is obtained through continuous iterative optimization
[0050] Example:
[0051] To verify the effectiveness of a two-stage policy gradient optimization automatic parallel method based on adaptive Tumanba, three carefully constructed datasets, namely PTB, CIFAR10, and NMT, were selected for testing. Specifically, PTB, CIFAR10, and NMT correspond to the recurrent neural network architecture, convolutional neural network, and the parameters of the adjusted neural machine translation model respectively. Each dataset contains 17 randomly selected computational graph samples, and the average number of operators in the computational graphs of NMT, CIFAR10, and PTB is 190, 300, and 500 respectively. Two key metrics were used for performance evaluation: the single-round iteration training time reflects the execution efficiency of the parallel policy (the shorter the time, the higher the efficiency), and the policy search time measures the policy generation speed (the less time-consuming, the faster the speed). Experimental data reveals that compared with existing device placement methods such as Placeto, GraphSAGE, P-GNN, CP-GNNAK, GNN, GCN, and GAT, the two-stage policy gradient optimization automatic parallel method based on adaptive Tumanba proposed and implemented in this invention demonstrates significant advantages in both policy search speed and parallel policy quality.
[0052] Table 1 intuitively shows the comparison of the single-round iteration training time of various methods under different datasets and GPU configurations. The results show that the automatic parallel method adopted in this invention demonstrates significant advantages compared with other methods. Specifically, compared with the Placeto method, when applying the parallel policy searched by this invention, the single-round iteration time is reduced by 10.24% to 28.66%. Further, Table 2 details the computational time data of each method under different datasets and the number of GPUs. From these data, it can be seen that this invention also performs excellently in terms of policy search speed. Compared with the Placeto method, the average search speed of this invention in finding the optimal parallel policy is increased by 95.89%, and this data further verifies the excellent performance of this invention in terms of efficiency.
[0053] Table 1 Comparison of single-round iteration training time of different methods on different datasets with different numbers of GPUs (unit: sec)
[0054]
[0055] Table 2 Comparison of policy search time of different methods on different datasets with different numbers of GPUs (unit: sec)
[0056]
[0057]
Claims
1. A two-stage policy gradient optimization automatic parallel method based on adaptive Tumanba, characterized in that, Including the following steps: Step 1: Obtain the publicly available AI model dataset, perform operator fusion on each computational graph G in the dataset, and obtain the feature matrix X of the computational graph; Step 2: Use X as the initial encoding matrix and input it into the Tumanba neural network to obtain the feature encoding X(t + 1) at time t + 1, and generate features through a multi-layer perceptron Step 3: Take the input perturbation two-stage policy gradient algorithm. For each training cycle, in the internal stage, j device placement policies are generated through perturbation noise and the parameters involved in the process of generating the perturbation noise are updated; in the external stage, policy evaluation is performed, the policy with the minimum training time is selected for global parameter update, and the optimal device placement policy O is output.
2. The automatic parallel method based on adaptive Tumanba's two-stage policy gradient optimization according to claim 1, wherein The specific process of obtaining the feature matrix X of the computational graph is as follows: Obtain the publicly available AI model dataset, perform operator fusion on each computational graph G in the AI model dataset through the ColocRL algorithm. The fused computational graph contains the relevant attributes of all operators in the computational graph, including operator type, input, output, and access identifier. Construct the original feature vector of the operator, that is, initialize the node encoding according to the relevant attributes of the operator, and obtain the complete feature matrix X of the computational graph. n represents the number of nodes in the computational graph G, and f is the feature dimension of the nodes in the computational graph G.
3. The automatic parallel method based on adaptive TuManBa for two-stage policy gradient optimization according to claim 2, wherein The specific implementation process of Step 2 is as follows: Step 2.1: At time t, input the encoding matrix X(t) into the Selective State Management (SSM) mechanism of the Tumanba neural network to generate a hidden encoding matrix and combine H(t) with X(t) to output a global encoding matrix Step 2.2: Concatenate M(t) and X(t) column-wise, and generate a weight dynamic adaptation factor through learnable weights and bias terms through the Sigmoid activation function Use α(t) to perform weighted fusion on M(t) and X(t) to obtain a transformed encoding matrix α(t) = Sigmoid[(M(t))‖X(t)×W α (t) + b α (t)] P(t) = (1 - α(t)) × M(t) + α(t) × X(t) Step 2.3: Combine P(t) with the initial encoding matrix X(0) through residual connection to generate the encoding matrix X(t + 1) at time t + 1: X(t + 1) = P(t) + X(0) Step 2.4: Repeat Steps 2.1, 2.2, and 2.3 until the preset number of iterations T is reached to obtain the feature matrix Input X(T) into the multi-layer perceptron MLP to obtain the final feature matrix 4. The automatic parallel method based on adaptive Tumanba for two-stage policy gradient optimization according to claim 3, characterized in that, The specific implementation process of Step 2.1 is as follows: Step 2.1.1: Introduce five learnable weight matrices and perform linear transformations on the input encoding matrix X(t) respectively to obtain the projection parameter and the discrete step size A(t) = W A(t) X(t), B(t) = W B(t) X(t), C(t) = W c(t) X(t), D(t) = W D(t) X(t), Δ(t) = W Δ(t) X(t) Step 2.1.2: Apply the Softplus activation function to Δ(t) to obtain and ensure its non-negativity; Subsequently, based on A(t), B(t), and calculate the discretized evolution parameter and Step 2.1.3: Based on and X(t), obtain the implicit encoding matrix H(t) at time t; by multiplying C(t) with to obtain the contribution of the implicit encoding H(t) to the global encoding; meanwhile, multiply D(t) with X(t) to obtain the contribution of the input feature X(t) to the global encoding; subsequently, add these two contributions to obtain the global encoding matrix M(t) at the current time:
5. The automatic parallel method for optimizing two-stage policy gradient based on adaptive Tumanba according to claim 4, characterized in that The specific implementation process of Step 3 is as follows: Step 3.1: In each training cycle k, the internal stage performs j iterations to generate j device placement strategies Among them, the device placement strategy is obtained by traversing the node features in to obtain the node status and adding controllable perturbation noise to obtain the perturbed node status to obtain the device placement item of the node The device placement strategies of all nodes constitute a device placement strategy In addition, based on perform parameter gradient update for the internal stage; Step 3.2: Based on the j strategies generated in the internal stage of the algorithm in Step 3.1, the optimal device placement strategy O in the current training cycle is determined in the external stage of the algorithm k ; After K training cycles, the device placement strategy with convergence, i.e., the minimum single-round iterative training time, is obtained 6. The automatic parallel method based on adaptive Tumanba with two-stage policy gradient optimization according to claim 5, wherein The specific implementation process of Step 3.1 is as follows: Step 3.1.1: In the internal update phase, traverse each node feature in the to obtain the node state through a fully connected layer where m represents the resource quantity of available devices; Add controllable perturbation noise to obtain the perturbed node state Subsequently obtain the device placement probability distribution generated by node i Device placement item Device placement probability Calculate the clipping loss through parameter calculation Update the gradients of the parameters in the internal phase; Step 3.1.2: Place the device items obtained in the internal update phase in the device placement strategy in sequence, and calculate the training time required for a single round of iteration according to By traversing the internal phase J times, J device placement strategies are obtained, and J execution times F are obtained, forming a training time set 7. The automatic parallel method for optimizing the dual-stage policy gradient based on the adaptive Tumanba according to claim 6, wherein The specific implementation of Step 3.1.1 is as follows: Step 3.1.1.1: Take the node features through a fully connected layer to obtain the node state Take and the previous moment's latent state through a GRU network to generate the node state and the latent state Take take the mean to obtain Obtain the standard deviation through exponential operation and based on introduce random noise to generate the perturbation state and through hyperparameters and Take and to obtain the node state through their weighted combination Take through a fully connected layer and the Softmax activation function to obtain the device placement probability distribution Step 3.1.1.2: According to the obtained device placement probability distribution Obtain the device placement probability And the device placement item Subsequently, these three parameters enter the internal update stage, and the parameters in Step 3.1.1.1 are updated by gradient; the specific implementation is as follows: According to what is obtained Take the logarithm to obtain the device placement probability Take the average maximum value of it as the device placement item of the current node And for Take the mean value to obtain Reflect the central tendency of the node on all possible device placement items; Subsequently, Perform an exponential calculation to obtain the device placement ratio By calculating And Calculate the advantage value through the sparse Softmax cross-entropy loss function And the learnable variable baseline value Calculate the probability value through the Sigmoid function Through And Multiply to obtain the surrogate loss Finally And Obtain the clipped loss And according to Update the parameters in step 3.1 by gradient 8. The automatic parallel method for optimizing the two-stage policy gradient based on the adaptive Tumanba according to claim 7, characterized in that, The specific implementation process of Step 3.2 is as follows: Step 3.2.1: In the external update stage of the algorithm, select the minimum value from the training time set as the minimum training time at the current moment Take the difference between and of adjacent training cycles as the reward value Step 3.2.2: Multiply the reward value R k by the discount factor , and add it to the historical discounted return L k-1 to obtain the current moment's discounted return Use L k to perform gradient update to optimize the global parameters, including the parameters of the Tu Mamba neural network in Step 2 and the parameters of the perturbed two-stage policy gradient algorithm in Step 3; continuously iterate and optimize to obtain the optimal device placement strategy O.