Language model intelligent training task scheduling method based on reinforcement learning

By constructing a set of task attributes and training state vectors, and combining an improved PPO model and buffer control mechanism, the problem of poor adaptability of scheduling strategies in language model training is solved, achieving efficient and stable multi-task scheduling and improving the overall performance of language model training.

CN121704976APending Publication Date: 2026-03-20MARM BRANCH OF STATE ENERGY GROUP QINGHAI ELECTRIC POWER CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511844609.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-09
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing language model training scheduling schemes fail to fully consider the rhythmic changes of training tasks at different stages, the dynamics of resource requirements, and the convergence status of parameters. This results in poor adaptability of scheduling strategies to training tasks, making them unable to cope with problems such as uneven load between tasks and gradient oscillations caused by frequent switching.

Method used

We construct a set of task attributes and a training state vector, and generate a scheduling strategy by driving the improved PPO model through the scheduling input sequence. We introduce a task splitting mechanism and a conflict weight calculation method, combined with a buffer control mechanism, to improve the continuity of training and the smooth transition of model parameters. In the policy update stage, we construct a smooth reward vector to dynamically control the magnitude of policy distribution changes and update frequency.

Benefits of technology

It significantly improves the scheduling stability and adaptability of language model training, effectively alleviates gradient fluctuations and resource waste caused by task switching conflicts, and achieves multi-task scheduling with high rhythm, high concurrency and high resource utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121704976A_ABST
    Figure CN121704976A_ABST
Patent Text Reader

Abstract

The invention discloses a language model intelligent training task scheduling method based on reinforcement learning, and the method comprises the following steps: constructing a task pool, and generating a task attribute set; performing rhythm coding and memory priority processing to generate a training state vector, inputting the training state vector to the improved PPO model, and outputting scheduling action distribution; executing task segmentation, and inserting fragment boundary mark information to generate a scheduling action set; calculating a step length difference value, a training duration difference value and a rhythm difference value between tasks, generating a conflict weight, and executing scheduling conflict constraint and buffer control; scheduling behaviors and training feedback are collected, training efficiency, parameter stability, task switching smoothness and task distribution diversity scores are calculated and fused to generate a smooth reward vector, and a task pool is updated and refreshed. According to the method, self-adaption, high efficiency and rhythm coordination of task scheduling are realized, and the self-adaption capability and execution efficiency of training task scheduling can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of language model training optimization technology, and in particular to a method for intelligent training task scheduling of language models based on reinforcement learning. Background Technology

[0002] With the widespread application of large-scale language models in tasks such as question-answering generation, intelligent customer service, and code completion, the multi-task management and resource scheduling issues involved in the training process have gradually become key factors affecting model performance and training efficiency. Most existing language model training scheduling schemes rely on fixed strategies or static rules for task allocation, failing to fully consider the rhythmic changes of training tasks at different stages, the dynamic nature of resource requirements, and the convergence state of parameters. This results in scheduling strategies having poor adaptability to training tasks and being unable to cope with problems such as uneven load between tasks and gradient oscillations caused by frequent switching.

[0003] In multi-task parallel training scenarios, current methods generally neglect the impact of task scheduling order on overall training stability and lack a dynamic evaluation mechanism for potential conflicts between training segment duration and switching rhythm. This easily leads to phenomena such as scheduling frequency imbalance, task selection bias, and resource distribution polarization. Furthermore, existing scheduling systems mostly use a single reward signal for policy optimization when facing complex scheduling feedback, failing to comprehensively consider multi-dimensional indicators such as training efficiency, parameter stability, task switching smoothness, and scheduling diversity. This results in slow convergence speed and weak generalization ability of scheduling policies, making it difficult to meet the long-term scheduling requirements of high-performance language model training.

[0004] Existing reinforcement learning scheduling methods mostly focus on coarse-grained resource allocation, lacking modeling of the fine-grained interaction between training task rhythm characteristics and scheduling behavior, making it difficult to construct a stable, efficient, and dynamically adaptable scheduling strategy model.

[0005] Therefore, how to provide an intelligent training task scheduling method for language models based on reinforcement learning is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0006] One objective of this invention is to propose an intelligent training task scheduling method for language models based on reinforcement learning. This invention combines task state features, training rhythm information, and historical scheduling behavior to construct a task attribute set and a training state vector. It then drives an improved PPO model to generate a scheduling strategy through a scheduling input sequence, addressing problems such as frequent task switching conflicts, disordered scheduling rhythm, uneven resource utilization, and unstable strategy updates in existing training processes. By constructing a rhythm-aware scheduling input sequence, it achieves accurate modeling of training task states and segment boundaries. It introduces a task segmentation mechanism and conflict weight calculation method to manage and suppress conflicts in the duration, step size, and rhythm fluctuations of training tasks in segments. A buffer control mechanism is used to temporarily store and fuse gradient information during task switching, effectively improving the continuity of training and the smooth transition of model parameters. In the strategy update stage, this invention constructs a smooth reward vector based on feedback information such as training efficiency, parameter stability, task switching smoothness, and task distribution diversity. It executes an adaptive entropy-controlled contraction and gradual update mechanism to dynamically control the magnitude of policy distribution changes and update frequency, improving the stability and adaptability of the scheduling strategy.

[0007] The intelligent training task scheduling method for language models based on reinforcement learning according to embodiments of the present invention includes the following steps: Step 1: Construct a task pool based on language model training tasks, collect information of each task in the task pool, and generate a set of task attributes; Step 2: Perform rhythm encoding on the attribute fields in the task attribute set, combine historical scheduling records with memory-first processing, and generate a training state vector; Step 3: Combine the training state vectors in a set order to form a scheduling input sequence, input it into the improved PPO model, and output the scheduling action distribution; Step 4: If the training duration of the selected training task is greater than the preset duration threshold, then the selected training task is split, and segment boundary marker information is inserted into the scheduling input sequence to generate a set of scheduling actions. Step 5: Calculate the conflict weight based on the step size difference, training duration difference, and rhythm difference between the trained task and the previous training task, and perform conflict constraint processing on the scheduling action distribution; if there is a task switching time, the improved PPO model generates a cache strength indication signal and triggers a buffer control operation. Step 6: Collect scheduling behavior records and training feedback information, calculate training efficiency score, parameter stability score, task switching smoothness score and task distribution diversity score, perform multi-reward fusion processing to form a smooth reward vector; Step 7: Execute the policy update operation of the improved PPO model based on the smoothed reward vector; collect the task state information and rhythm parameters of the trained task, refresh the task pool, and construct a new round of scheduling input sequence.

[0008] Preferably, step one specifically comprises: Based on the language model training task, a task pool is constructed according to a preset task structure; input requests for multiple language model training tasks are received, and task identifiers are registered according to the task generation order; Information collection is performed on each task in the task pool; the information of the task includes task type, training duration, loss change, resource consumption and scheduling interval information; the task type records task category information, the training duration records the historical training time of the task, the loss change records the changing trend of the loss function in the continuous training phase, the resource consumption records the computational resource consumption information when the task is executed, and the scheduling interval information records the time interval between two scheduling of the task. Perform a uniform formatting operation on the collected task information, mapping each piece of information to task attribute fields; group and combine the formatted attribute fields according to the task identifier to construct a task attribute set.

[0009] Preferably, the training state vector includes task type encoding, rhythm feature encoding, cost encoding, loss trend encoding, and time interval encoding, specifically: Perform rhythm encoding on each set of attribute fields in the task attribute set, calculate the rhythm fluctuation amplitude and periodic parameters of the task based on the training duration and scheduling interval information, and generate rhythm feature encoding; Based on task type, resource consumption and loss change information, a task feature sequence is constructed, and structured mapping rules are called to generate task type code, cost code and loss trend code; Collect historical scheduling records of tasks, count the selection frequency and relative position of each task in different scheduling cycles, and construct a scheduling preference sequence; perform memory-first processing on the scheduling preference sequence to generate a scheduling preference vector and combine it with the aforementioned encoded information; Generate time interval codes based on the time interval between two task schedulings; The task type encoding, rhythm feature encoding, cost encoding, loss trend encoding, and time interval encoding are concatenated into a training state vector according to a set structure, and the mapping relationship between the training state vector and the task identifier is recorded.

[0010] Preferably, the step of combining the training state vectors in a predetermined order to form a scheduling input sequence specifically involves: Collect all training state vectors in the task pool, set sorting priority parameters according to task identifier order, task historical scheduling frequency and remaining training time, perform priority sorting on the training state vectors, and generate a sorting queue. Based on the order of the sorted queue, read the contents of the training state vectors sequentially and construct a vector permutation list; Perform structural unification processing on the vector permutation list, set unified input dimension standards and position information identification rules, and generate a structured vector matrix; Perform positional encoding and sequence normalization operations on the structured vector matrix to generate a scheduling input sequence with sequential encoding; record the index mapping information between the scheduling input sequence and the original training state vector to construct a scheduling input sequence index table.

[0011] Preferably, the improved PPO model includes an input encoding structure, an action prediction structure, a state evaluation structure, and a policy update structure, specifically: The input encoding structure receives the scheduled input sequence, performs embedding mapping processing, sequential position mapping processing, and normalization processing, and generates an encoded input tensor. The action prediction structure receives the encoded input tensor and outputs a scheduled action distribution based on a multi-layer feedforward neural structure and an action decision distribution function; the scheduled action distribution includes an action probability vector corresponding to the training task number and a boundary identifier component of the task segment boundary; The state evaluation structure receives the encoded input tensor, constructs a state value function path, and outputs the state value estimate corresponding to the training state. The strategy update structure includes an entropy-controlled contraction unit and a slow-cut update unit; the entropy-controlled contraction unit performs dynamic compression processing on the entropy range of the scheduling action distribution to adjust the convergence amplitude of the strategy distribution; the slow-cut update unit includes a shearing amplitude adjustment structure and a parameter delay control structure, performs shearing boundary updates based on the policy change amplitude, and performs step-by-step update processing based on the update delay period; The improved PPO model constructs a smooth reward vector based on training efficiency score, parameter stability score, task switching smoothness score, and task distribution diversity score collected during training. It calculates the advantage function by combining historical trajectory records, performs joint optimization on the parameters of the action prediction structure and the state evaluation structure, and outputs an updated scheduling strategy.

[0012] Preferably, step four specifically includes: A duration comparison operation is performed based on the training duration of the selected training task and a preset duration threshold. When the training duration exceeds the preset duration threshold, the segmentation logic function is called to divide the selected training task into multiple training segments according to the set segment length parameter and segment boundary rules. Each training segment is assigned a unique task segment identifier, and segment boundary marker information is inserted into the scheduling input sequence; the segment boundary marker information includes the training segment start position, end position, task segment number, and segment order label; Establish a mapping relationship between the training fragment identifier and the original task identifier, update the structure field of the corresponding task in the task attribute set, and generate a fragment structure attribute set; The fragment structured attribute set is used as input to construct a fragment-level training state vector and participate in the encoding process of scheduling the input sequence; During the scheduling action output phase, the action probability vector and boundary identifier component are jointly decoded to generate a set of scheduling actions containing training segment numbers and segment boundary indicators.

[0013] Preferably, step five specifically includes: Collect the training segment numbers and training state vectors of the training task and the previous training task; based on the step size information, training duration information and rhythm feature information recorded in the training state vector, calculate the step size difference, training duration difference and rhythm difference between the two tasks. Based on the preset conflict weight calculation function, the weight combination mechanism is invoked to map the step size difference, training duration difference, and rhythm difference into conflict weight values; the conflict weight values ​​are used to determine the switching conflict level between scheduling actions. Based on the switching conflict level, the conflict suppression rules are invoked to perform conflict constraint processing on the action probability vector in the scheduling action distribution; the conflict constraint processing includes probability compression, candidate elimination and order priority adjustment operations, which are executed in sequence according to the conflict level. If a task switching time is determined to exist, the improved PPO model outputs a cache strength indicator signal; the cache strength indicator signal includes the switching level, buffer period, and fusion bias label. The buffer control operation is triggered based on the cache strength indicator signal; the buffer control operation includes: saving the gradient state of the current training task, temporarily storing the gradient state in the cache, setting the parameter delay update period, calling the gradient fusion function to perform the fusion of the old and new gradients; and binding the gradient state after buffer control with the policy update structure.

[0014] Preferably, step six specifically includes: The scheduling behavior records and training feedback information generated during the training process are collected to construct a behavior feedback set; the scheduling behavior records include scheduling action number, task switching time point, training segment identifier and scheduling interval; the training feedback information includes task loss value change, gradient convergence status, resource consumption record and training completion status. Training efficiency scores are calculated based on behavioral feedback sets; the training efficiency scores are quantified as the ratio of the decrease in loss value per unit time to the execution time of the training segment; Calculate the parameter stability score; the parameter stability score is evaluated based on the gradient variance and parameter change rate between consecutive segments; Calculate the task switching smoothness score; the task switching smoothness score is calculated based on the degree of resource consumption jump at the switching point and the stability of model output; Calculate the task distribution diversity score; the task distribution diversity score is based on the distribution balance of task numbers in historical scheduling; A weighted fusion operation is performed on the training efficiency score, parameter stability score, task switching smoothness score, and task distribution diversity score to construct the initial value of the reward vector; The initial value of the reward vector is subjected to temporal smoothing and anomaly compensation processing to generate a smoothed reward vector. The smoothing process includes sliding window weighting, score normalization and reward clipping, and the anomaly compensation process includes zero-reward filling and spike correction. The smoothed reward vector is then associated with the training task identifier corresponding to the scheduling input sequence.

[0015] Preferably, the policy update operation based on the smoothed reward vector to perform the improved PPO model includes adaptive entropy-controlled shrinkage and gradual shearing update; the gradual shearing update includes adaptive shearing amplitude adjustment and grouped delay processing, specifically: A sequence of advantage function estimation is constructed based on smoothed reward vectors and historical scheduling trajectory data; the action prediction structure is invoked to output the distribution of scheduling actions, and the state evaluation structure is invoked to output the state value estimate; the policy loss function and the value loss function are calculated to form a joint optimization objective. The adaptive entropy control shrinkage operation dynamically adjusts the entropy value change range of the scheduling action distribution; the shrinkage coefficient is determined according to the entropy value change rate to compress the dispersion of the policy distribution; the shrinkage policy gradient is generated based on the compression result, and the update amplitude of the policy distribution is adjusted. The adaptive shearing amplitude adjustment operation monitors the change amplitude of the strategy and the fluctuation trend of the advantage function, and calculates the shearing factor; based on the shearing factor, the shearing boundary range is set to constrain the gradient change of the joint optimization objective. By using a grouped update operation, the parameters of the improved PPO model are divided into multiple update groups based on the rhythm information and parameter change frequency of the training task; the parameters of each group are updated step by step according to the preset update cycle; and the corresponding parameters in the action prediction structure and state evaluation structure are adjusted synchronously within each update cycle. The compression policy gradient of the adaptive entropy control shrinkage output and the constraint gradient of the gradual update output are fused to generate a joint optimization gradient tensor. The parameter state of the improved PPO model is updated based on the joint optimization gradient tensor to complete the policy update operation.

[0016] Preferably, the step of collecting the task state information and rhythm parameters of the trained task and refreshing the task pool to construct a new round of scheduling input sequence specifically involves: The system collects the status information and rhythm parameters of the trained task after the scheduling and training operations are completed. The task status information includes the training segment number, training progress, loss change trend, and parameter convergence rate. The rhythm parameters include the training time update value, the scheduling interval update value, and the rhythm change amplitude. Update the task attribute fields based on the task status information and rhythm parameters, and reconstruct the task attribute set; The task identifier is determined based on the task completion status. If the task training is complete, the corresponding identifier is removed from the task pool. For incomplete tasks, the identifier is retained and its status is updated to pending scheduling. Perform rhythm encoding and memory-first processing on the updated task attribute set to construct a new set of training state vectors; The training state vector set is combined and sorted according to the set order and scheduling priority to generate a new round of scheduling input sequence; Record the index mapping relationship of the scheduling input sequence and establish an input connection with the improved PPO model input encoding structure to enter the next scheduling cycle.

[0017] The beneficial effects of this invention are: To address issues such as frequent task switching conflicts, low training resource utilization, lack of coordination in scheduling rhythm, and unstable training policy updates in existing training task scheduling processes, this invention constructs a task pool and collects multi-source task information. A unified format is used to generate a set of task attributes. A training state vector is generated through rhythmic encoding and memory-first processing. A scheduling input sequence is dynamically constructed by combining historical task scheduling behavior and rhythm patterns. This sequence is then input into an improved PPO model with input encoding, action prediction, state evaluation, and policy update structures. The output is a distribution of scheduling actions including task boundary markers, achieving fragmented scheduling control for multiple tasks. During the policy update phase, this invention constructs a smooth reward vector based on multiple training feedback metrics, integrating scores for training efficiency, parameter stability, task switching smoothness, and distribution diversity. An adaptive entropy-controlled contraction and gradual update mechanism dynamically adjust the policy convergence process, significantly improving the model's scheduling stability and convergence efficiency. During task switching, a conflict weight and cache control mechanism are introduced to effectively mitigate gradient fluctuations and resource waste caused by scheduling conflicts, ensuring smooth transitions and consistent parameter fusion across multi-task training processes. Ultimately, this achieves high rhythmicity, high concurrency, and high resource utilization efficiency in the language model training task scheduling process, providing intelligent and stable scheduling strategy support for large-scale language model training in a multi-task scheduling environment. Attached Figure Description

[0018] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of the intelligent training task scheduling method for language models based on reinforcement learning proposed in this invention; Figure 2 This is a schematic diagram of the improved PPO model proposed in this invention. Detailed Implementation

[0019] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0020] refer to Figure 1 and Figure 2 The intelligent training task scheduling method for language models based on reinforcement learning includes the following steps: Step 1: Construct a task pool based on language model training tasks, collect information of each task in the task pool, and generate a set of task attributes; Step 2: Perform rhythm encoding on the attribute fields in the task attribute set, combine historical scheduling records with memory-first processing, and generate a training state vector; Step 3: Combine the training state vectors in a set order to form a scheduling input sequence, input it into the improved PPO model, and output the scheduling action distribution; Step 4: If the training duration of the selected training task is greater than the preset duration threshold, then the selected training task is split, and segment boundary marker information is inserted into the scheduling input sequence to generate a set of scheduling actions. Step 5: Calculate the conflict weight based on the step size difference, training duration difference, and rhythm difference between the trained task and the previous training task, and perform conflict constraint processing on the scheduling action distribution; if there is a task switching time, the improved PPO model generates a cache strength indication signal and triggers a buffer control operation. Step 6: Collect scheduling behavior records and training feedback information, calculate training efficiency score, parameter stability score, task switching smoothness score and task distribution diversity score, perform multi-reward fusion processing to form a smooth reward vector; Step 7: Execute the policy update operation of the improved PPO model based on the smoothed reward vector; collect the task state information and rhythm parameters of the trained task, refresh the task pool, and construct a new round of scheduling input sequence.

[0021] This implementation constructs a language model training task pool and collects task information to generate a task attribute set, enabling structured management of multi-task scheduling objects and providing comprehensive and controllable basic information support for subsequent scheduling strategies. By performing rhythm encoding and memory priority processing on task attribute fields, a training state vector containing task periodicity and historical scheduling preferences is generated, further improving the time sensitivity of state representation and the ability to associate scheduling context. Combining the training state vectors in a set order to form a scheduling input sequence and inputting it into the improved PPO model can fully explore the relative priority relationships between tasks and enhance the structural rationality of action output. When the training duration of the selected training task exceeds a threshold, a segmentation operation is performed and the segment boundaries are marked, effectively improving the processing granularity and scheduling flexibility of long tasks. Furthermore, conflict weights are constructed based on the step size difference, training duration difference, and rhythm difference between the training task and the previous training task, and conflict constraint processing of the scheduling action distribution is performed, while triggering buffer control operations. This can suppress unstable behavior caused by high-frequency switching and improve training continuity and parameter fusion stability. By collecting scheduling behavior records and training feedback information, multi-dimensional training scores are calculated and fused into a smooth reward vector, enabling global quantitative feedback on training effectiveness and enhancing the reliability and discriminative power of the reward signal. Based on this, policy update operations are performed using the smooth reward vector, further improving the directionality and response accuracy of policy learning. Simultaneously, task state information and rhythm parameters of the trained task are collected, refreshing the task pool to construct a new round of scheduling input sequences, forming a closed-loop iterative mechanism to achieve dynamic collaborative optimization of the scheduling policy and task state.

[0022] In this embodiment, step one specifically includes: Based on the language model training task, a task pool is constructed according to a preset task structure. The task structure includes a task identifier field, a task type field, a training period field, a loss state field, a resource parameter field, and a scheduling history field. During the construction of the task pool, task identifiers are registered in the order of task input time, and the corresponding task access source, model category, and task priority label are recorded. It receives input requests from multiple language model training tasks, performs task structure initialization operations on each received task, establishes the binding relationship between task fields and task identifiers, and collects the external parameter configurations during task initialization, including training objective function, optimizer configuration, training batch size and device type identifier, and writes them into the task structure as task auxiliary fields. Information collection is performed on each task in the task pool. The collected information includes task type, training duration, loss variation, resource usage, and scheduling interval. The task type field records the model category, data scale type, and task target identifier of the task. The training duration field records the cumulative training time from the first scheduling to the present. The loss variation field collects the loss value in each training round, constructs a mapping sequence between training rounds and loss values, and calls a multinomial fitting function to fit the loss variation trend, generating trend slope coefficients, fitting residuals, and stability indicators. The resource usage field records the computational resource request information of the task during execution, including GPU memory usage, computation time utilization, and memory load. The scheduling interval field records the scheduling time points of the task in the task pool, calculates the time interval sequence between consecutive scheduling time points, and generates the interval mean and fluctuation amplitude values ​​through a moving average function. All collected raw fields are uniformly formatted to construct a standard task attribute field set. The formatting operation includes field type conversion, unit standardization, missing completion and structure mapping. Task type information is converted into an enumeration encoding format, training duration and scheduling interval information are uniformly converted into minute units, resource usage information is normalized to a percentage value, and loss change information is mapped into a trend description vector according to a preset function structure. All formatted task attribute fields are grouped and combined according to task identifiers to construct a task attribute set. Each item in the task attribute set contains a set of standardized attribute fields and their corresponding task identifiers, and the field order and field type information are recorded through an index table. The task attribute set is provided as input to the subsequent rhythm encoding operation and training state vector construction process.

[0023] In this embodiment, the training state vector includes task type encoding, rhythm feature encoding, cost encoding, loss trend encoding, and time interval encoding, specifically: Perform rhythm encoding on each set of attribute fields in the task attribute set, calculate the rhythm fluctuation amplitude and periodic parameters of the task based on the training duration and scheduling interval information, and generate rhythm feature encoding; Based on task type, resource consumption and loss change information, a task feature sequence is constructed, and structured mapping rules are called to generate task type code, cost code and loss trend code; Collect historical scheduling records of tasks, count the selection frequency and relative position of each task in different scheduling cycles, and construct a scheduling preference sequence; perform memory-first processing on the scheduling preference sequence to generate a scheduling preference vector and combine it with the aforementioned encoded information; Generate time interval codes based on the time interval between two task schedulings; The task type encoding, rhythm feature encoding, cost encoding, loss trend encoding, and time interval encoding are concatenated into a training state vector according to a set structure, and the mapping relationship between the training state vector and the task identifier is recorded.

[0024] In this embodiment, the step of combining the training state vectors in a predetermined order to form a scheduling input sequence specifically involves: Collect all training state vectors from the task pool; the training state vectors include task type encoding, rhythm feature encoding, cost encoding, loss trend encoding, and time interval encoding; set sorting priority parameters according to task identifier order, historical task scheduling frequency, and remaining training time; the task identifier order is obtained by calculating the task generation time and scheduling participation time, the historical task scheduling frequency is calculated by counting the cumulative occurrences of the corresponding task identifier in the scheduling input sequence, and the remaining training time is calculated by the difference between the total training time and the completed time; call the multi-factor priority fusion function to fuse the above three factors and generate a priority score value; form a mapping set by combining all training state vectors and their corresponding priority scores, sort the mapping set according to the score value from high to low, and generate a training state vector sorting queue; According to the order in the sorting queue, the encoded content of each training state vector is read sequentially to construct a vector permutation list; a dimension alignment operation is performed on each state vector in the vector permutation list; the dimension alignment operation includes missing value filling, redundant field removal, and field order unification processing; the dimension mapping specification rules are called to align all state vectors to a unified structure template and output a standardized vector set; The standardized vector set is subjected to structural unification processing; the processing includes setting a unified input dimension standard and position information identification rules; the input dimension standard is set according to the improved PPO model input tensor format, including the total vector length, field encoding order and nesting structure requirements; the position information identification rules set the mapping relationship between the position label of each state vector in the input sequence and the original task label; all standardized state vectors are stacked in the order of arrangement to generate a structured vector matrix; Position encoding and sequence normalization operations are performed on the structured vector matrix. The position encoding adopts a sine and cosine function position embedding mechanism. By setting a frequency group, the sine and cosine components of each position are calculated to generate a position embedding vector. The position embedding vector and the state vector are concatenated and fused to form a state code with sequential characteristics. The sequence normalization operation includes input tensor dimension reshaping, format verification and type declaration to generate a scheduling input sequence that can be input into the improved PPO model. Record the index mapping information between the scheduling input sequence and the original training state vector, and construct a scheduling input sequence index table; the index table includes the position index of the state vector in the input sequence, the original task identifier, the scheduling round to which it belongs, and the segment number, which is used to support subsequent segment boundary insertion, caching operations, and reward feedback association operations.

[0025] This implementation method effectively enhances the temporal modeling capability of training tasks and the context awareness capability of strategy optimization by constructing a scheduling input sequence with a unified structure, clear sequence, and complete information mapping, thereby improving the stability and accuracy of the scheduling strategy.

[0026] In this embodiment, the improved PPO model includes an input encoding structure, an action prediction structure, a state evaluation structure, and a policy update structure, specifically: The input encoding structure receives the scheduled input sequence; it calls the embedding mapping rule to perform independent embedding processing on each field in the scheduled input sequence; the embedding mapping rule sets different dimensions and initialization methods according to the field category, calls the linear transformation embedding function for numerical fields, and calls the word embedding matrix generation function for categorical fields; it performs order position mapping processing on the embedded vector, constructs a position encoding tensor using sine and cosine functions, and concatenates and fuses the embedded vector and the position encoding tensor; the fused vector is input to a normalization function, and amplitude compression is performed using mean-variance normalization or maximum-minimum normalization rules to generate an encoded input tensor; The action prediction structure receives an encoded input tensor and sequentially generates a sequence of hidden state vectors through a multi-layer feedforward neural network structure. The feedforward neural network structure includes multiple fully connected layers, activation functions, and residual connection modules. The activation function uses ReLU or GELU. An action decision distribution function is invoked to perform action probability estimation on the hidden state vectors, outputting a scheduling action distribution. This scheduling action distribution includes an action probability vector and a boundary marker component. The action probability vector represents the probability that each training task or task segment is selected as the next scheduling object, and the boundary marker component records whether the task is segmented and the segment boundary position information. The state evaluation structure receives an encoded input tensor and calls the state value function path. The state value function path shares the first two network parameters with the action prediction structure, and is then separated into independent network paths to construct a state value estimation network. For each input state vector, a corresponding state value estimate is output, and the mapping relationship between the state and the state value is recorded. The state value estimate is used to evaluate the expected performance of the current scheduling state in long-term rewards. The strategy update structure includes an entropy-controlled contraction unit and a gradual update unit. The entropy-controlled contraction unit performs entropy analysis on the action probability distribution and calls a dynamic compression function to fit the entropy range. The dynamic compression function generates a compression coefficient based on the deviation between the average entropy value of the current round and the target entropy threshold, adjusting the convergence speed of the strategy distribution. The gradual update unit includes a shearing amplitude adjustment structure and a parameter delay control structure. The shearing amplitude adjustment structure monitors the policy change amplitude, uses the KL divergence index to calculate the policy offset, determines whether to adjust the shearing boundary based on a set threshold, and outputs a shearing correction coefficient. The parameter delay control structure calls the policy gradient function and the optimizer structure, sets the update period window and step size control factor, performs step-by-step gradient update operations, and outputs the updated network parameters. During training, training efficiency score, parameter stability score, task switching smoothness score, and task distribution diversity score are collected; a score weighted fusion function is called to construct a smooth reward vector; the reward vector is combined with historical scheduling trajectory records to calculate the advantage function; the advantage function is obtained by fitting using the GAE method and is generated based on the weighted difference between the current state value and the future reward value; the advantage function is used as the basis for gradient optimization to perform joint optimization on all trainable parameters in the action prediction structure and the state evaluation structure.

[0027] This implementation improves the convergence efficiency of the scheduling strategy and the accuracy of action output by constructing an improved PPO model with fine embedding, clear structure, and stable updates, thereby enhancing the intelligent scheduling performance and long-term learning ability of the language model training task.

[0028] In this embodiment, step four specifically includes: A duration comparison operation is performed based on the training duration of the selected training task and a preset duration threshold; the task duration determination function is called to compare the historical training duration field of the training task with the threshold parameter field; if the determination result is that the training duration exceeds the preset duration threshold, the segmentation logic function is called to perform segment construction processing; The segmentation logic function is invoked to construct a segmentation index list based on the set segment length parameters, boundary constraint rules, and the original structure of the training task. The segmentation index list includes the start index, end index, segment number, and sequence label of each training segment. Based on the segmentation index list, the original training task is divided into multiple training segments. The segment identifier generation module is invoked to generate an independent task segment identifier for each training segment and record the mapping pairs between task identifiers and segment identifiers. The fragment identifier, start position, end position, and sequence label are combined to form a boundary marker field, which is then inserted into the corresponding position of the scheduling input sequence to construct a sequence input matrix containing fragment structure information. The attribute reorganization function is called to decompose and reorganize the original task fields in the task attribute set. Based on the mapping relationship between the task identifier and the fragment identifier, the structure field is updated and a fragment structured attribute set is generated. The fragment structured attribute set is input into the training state encoding path, and rhythm encoding, type encoding, cost encoding, trend encoding and time interval encoding operations are performed; a fragment-level training state vector set is generated; the fragment-level training state vector set is concatenated and fused with the boundary marker field to form an extended state vector containing structural boundary information, which participates in the scheduling input sequence encoding process; During the scheduling action output phase, a joint decoding function is invoked to jointly decode the action probability vector and boundary identifier components output by the action prediction structure; the segment number and boundary indication information corresponding to each scheduling action are extracted to form a scheduling action set; the scheduling action set includes the scheduling probability value, the task number to which it belongs, the boundary label and the segment order position parameter for each training segment.

[0029] This implementation improves the ability of long tasks to be executed in segments and the ability to control the granularity of scheduling by introducing a task segmentation mechanism based on training duration and a segment boundary labeling method. It also enhances the model's adaptability to tasks with uneven training intensity and its ability to distinguish between different tasks.

[0030] In this embodiment, step five specifically includes: The training segment numbers and training state vectors of the training task and the previous training task are collected; the step size field, training duration field, and rhythm feature field are extracted from the training state vector and recorded as step size parameter, duration parameter, and rhythm parameter, respectively; according to the scheduling order when switching tasks, the parameter differences between the current task and the previous task are matched, and the step size difference, training duration difference, and rhythm difference are calculated; the difference calculation is constructed through an absolute difference function, and the amplitude is uniformly processed by standardization. The step size encoding field records the average parameter update magnitude during the historical training phase of the task. According to the set gradient change evaluation function, the average gradient magnitude change of several consecutive training segments is used as the step size index. The difference function is called to calculate the absolute difference between the step size indices of the two tasks, generating a step size difference. The training duration field records the cumulative training time for each task. The standardization function is called to scale the original training time to a normalized interval, and then the training durations of the two tasks are calculated to generate a training duration difference. The rhythm feature encoding field includes rhythm fluctuation amplitude encoding and periodic parameter encoding. The rhythm fluctuation amplitude encoding is quantified based on the standard deviation of the task scheduling interval change, and the periodic parameter encoding is constructed by extracting the main frequency from the task occurrence frequency sequence based on Fourier transform. The rhythm distance calculation function is called, inputting the rhythm fluctuation encoding and periodic parameter encoding of the two tasks respectively, calculating the Euclidean distance and normalizing it to generate a rhythm difference. The preset conflict weight calculation function is invoked to calculate the conflict weight value based on the difference combination; the conflict weight calculation function is a polynomial weighted function, and the weight factor is obtained by fitting through an offline optimization process. The minimum variance fitting is performed based on the task stability label to ensure the sensitivity and resolution of the conflict weight mapping. Based on the conflict weight value, the conflict level classification rule is invoked, and multi-level conflict level labels are set. The action probability vector in the scheduling action distribution is used as the processing object, and conflict constraint processing operations are performed step by step according to the conflict level. The conflict constraint processing includes: performing probability compression on high conflict level action components, calling the compression function to reduce their distribution density; performing candidate elimination on action components that exceed the conflict threshold, removing their possibility of participating in scheduling selection; and prioritizing the execution order of medium and low conflict level actions to improve the ranking priority of low conflict tasks in the distribution. The switching detection function is invoked to determine whether there is a difference in task number or segment number between the training task and the previous training task; if a task switching is detected, a cache strength indicator signal is output; the cache strength indicator signal includes a switching level label, a set buffer period parameter, and a fusion bias label; The buffer control module is invoked to trigger buffer control operations based on the switching level and buffer period. During the buffer control process, the gradient state snapshot module is invoked to save the model gradient state of the current task. The state writing module is invoked to temporarily store the gradient state in the buffer area. The parameter delay scheduling module is invoked to set the delay update period. During the delay update period, the gradient fusion function is invoked to control the weighting ratio of the new and old gradients based on the fusion bias label and perform parameter fusion operations. The fused gradient state after buffer control is structurally bound to the policy update structure to form a delayed update sub-path in the policy update path; the index correspondence between the fused gradient state and the scheduling action is recorded and added to the historical trajectory record set.

[0031] This implementation method effectively reduces the risk of sudden changes and parameter fluctuations in task scheduling by introducing a conflict weight-driven scheduling action suppression mechanism and a task switching buffer control mechanism, thereby improving the stability of the strategy and the smoothness of task switching during the training process.

[0032] In this embodiment, step six specifically includes: Collect scheduling behavior records and training feedback information generated during training to construct a behavior feedback set; the scheduling behavior records include scheduling action number, task switching time point, training segment identifier and scheduling interval parameter; the training feedback information includes task loss value change trajectory, gradient convergence status within training segment, resource consumption index sequence and task training completion status identifier. The behavioral feedback set is analyzed to construct a unit time loss value reduction rate index; a duration standardization function is called to calculate the actual execution time of each training segment; the ratio of loss value reduction rate to execution time is used as the basic index for training efficiency scoring; a non-linear scaling function is called to map the scoring index and generate a training efficiency score. Calculate the gradient variance between consecutive training segments and extract the parameter change rate sequence; according to the set stability coefficient rule, weight the gradient variance and the change rate to generate a parameter stability score; the weighted combination weight coefficient is obtained by fitting the historical best model trajectory and training with a mean square error minimization strategy. Collect resource usage curves before and after task switching and calculate resource jump amplitude index; collect model output sequence and calculate output continuity fluctuation amplitude; according to the set smoothness mapping function, fuse jump amplitude and output fluctuation amplitude to generate task switching smoothness score; The frequency of task numbers in historical scheduling sequences is counted; the standard deviation and distribution entropy of each task number are calculated; based on the principle of minimizing distribution entropy, the balance of task distribution is scored to generate a task distribution diversity score. The training efficiency score, parameter stability score, task switching smoothness score, and task distribution diversity score are input into the fusion scoring function; the scoring weight adjustment module is called to set an adjustment weight for each score. The weight values ​​are generated by fitting the task importance label and the training stage number; the weighted fusion operation is performed, and the initial value of the reward vector is output. The initial value of the reward vector is initialized by calling the sliding window processing function to perform local smoothing operation; the sliding weight is adjusted according to the window length setting strategy and time decay parameter to generate preliminary smoothing results; the score normalization module is called to standardize the smoothing results using the maximum and minimum scaling method; the reward pruning function is called to truncate reward components that exceed the positive and negative thresholds to improve the stability of reward values. The anomaly compensation module is invoked to perform correction operations on segments in the reward sequence that have zero rewards or spikes in scores; zero-reward segments are filled using a local mean fitting method to generate filler values; spike correction uses a gradient stability constraint model to calculate predicted values ​​to replace outliers. The final smoothed reward vector is indexed and associated with the training task identifier in the scheduling input sequence; the index mapping table and the reward value mapping table are recorded.

[0033] This implementation method enhances the continuity and discriminative power of the reward distribution while ensuring the authenticity of the reward signal by constructing a multi-dimensional task feedback index system and performing fusion scoring and time-series smoothing operations. This improves the convergence efficiency and stability of the improved PPO model in the scheduling strategy learning process.

[0034] In this embodiment, the policy update operation based on the smoothed reward vector to perform the improved PPO model includes adaptive entropy-controlled shrinkage and gradual shearing update; the gradual shearing update includes adaptive shearing amplitude adjustment and grouped delay processing, specifically: Construct an advantage function estimation sequence; the advantage function estimation sequence is calculated based on the smoothed reward vector and historical scheduling trajectory data, and the advantage function is a measure of the long-term benefit advantage brought by the scheduling behavior, specifically obtained by fitting the GAE function; The system calls the action prediction structure to receive the scheduling input sequence and outputs the corresponding scheduling action distribution; it calls the state evaluation structure to output the state value estimate; it constructs the policy loss function based on the action distribution probability and the advantage function estimate; it constructs the value loss function based on the difference between the state estimation deviation and the actual smoothed reward; and it weights and fuses the policy loss function and the value loss function to generate a joint optimization objective. An adaptive entropy-controlled shrinkage operation is performed to dynamically adjust the entropy value change range of the scheduling action distribution; a shrinkage coefficient is set based on the rate of change of entropy value during training; a policy compression function is called to perform a compression operation on the dispersion of the policy distribution, generating an entropy-controlled shrinkage gradient; the shrinkage coefficient is obtained by fitting the derivative of the entropy change curve. Perform adaptive shearing amplitude adjustment operation; collect the change difference value between the strategy output of the previous round and the current round, and calculate the shearing factor in combination with the fluctuation trend of the advantage function; the shearing factor is obtained by fitting the shearing amplitude function to control the amplitude boundary of the strategy update; call the shearing boundary function to constrain the gradient change of the joint optimization objective and output the shearing constraint gradient; Perform grouped extension operation; collect rhythm information and parameter change frequency of training task, divide model parameters into multiple extension groups according to preset rhythm matching rules; set an independent extension period for each group of parameters; within each extension period, call the extension control logic to synchronously update the corresponding parameter subsets in the action prediction structure and state evaluation structure; the extension period is obtained by fitting the parameter fluctuation period function; The compression strategy gradient of the adaptive entropy control shrinkage output and the constraint gradient of the gradual update output are combined, and the joint optimization gradient tensor is generated by calling the joint gradient construction function. The joint optimization gradient tensor is then input into the optimizer to perform gradient descent operation, update the parameter state of the improved PPO model, and complete this round of policy update.

[0035] This implementation dynamically controls the convergence direction of the strategy through an adaptive entropy control contraction mechanism, and controls the update frequency and magnitude through shear boundary adjustment and parameter grouping delay strategies, effectively improving the stability of the model strategy, convergence efficiency and scheduling generalization ability.

[0036] In this embodiment, the step of collecting the task state information and rhythm parameters of the trained task and refreshing the task pool to construct a new round of scheduling input sequence specifically includes: The system collects the status information and rhythm parameters of the trained task after the scheduling and training operations are completed. The task status information includes the training segment number, training progress, loss change trend, and parameter convergence rate. The training progress is calculated based on the ratio of the training segment execution time to the total time. The loss change trend is obtained by fitting the difference curve of the loss function value within adjacent training cycles. The parameter convergence rate is estimated by the rate of change of the continuous gradient norm. The rhythm parameters include the training time update value, the scheduling interval update value, and the rhythm change amplitude. The rhythm change amplitude is obtained by fitting the rate of change of scheduling time points within the sliding window period. The task status information and rhythm parameters are mapped to the original task attribute fields, and the corresponding field values ​​in the task pool are updated; the task attribute refresh function is called to perform a unified update on the attribute set of all tasks, and a new set of task attributes is constructed. Perform a screening and judgment operation on the task completion status field in the task attribute set; if the task training is determined to be completed, delete the task identifier and release its resource allocation flag; if the task is determined to be incomplete, retain the task identifier and reset its scheduling status field to pending scheduling, update the task priority field value and synchronize it to the task status vector label. Perform rhythm encoding processing on the updated task attribute set; call the rhythm embedding module to perform period normalization, volatility compression and rhythm label mapping on the rhythm parameters to generate a rhythm encoding tensor; combine the historical scheduling trajectory and call the memory priority function to perform weighted replay processing on the task execution records to generate a task memory vector; The rhythm encoding tensor is concatenated with the task memory vector to construct a new set of training state vectors; the scheduling order sorting module is called to perform combined sorting on the training state vector set; the sorting priority is determined by the task identification order, rhythm confidence score and memory playback intensity; the rhythm confidence score is obtained by fitting the rhythm fluctuation stability function, and the memory playback intensity is calculated by the historical scheduling coverage function. A new round of scheduling input sequence is generated based on the sorting results; the index position of each state vector in the scheduling input sequence is recorded, and a scheduling input sequence index mapping table is constructed; the scheduling input sequence is connected with the input encoding structure of the improved PPO model, and the next scheduling cycle begins.

[0037] This implementation method dynamically constructs a set of task attributes and a set of training state vectors by jointly collecting and encoding the state feedback and rhythm parameters of the trained task, and optimizes and updates the input order and scheduling priority in each scheduling iteration, thereby effectively improving the adaptive adjustment capability and task allocation rationality of the scheduling model.

[0038] Example 1: To verify the feasibility of this invention in practice, it was applied to the task scheduling system of an enterprise-level language model training platform. This platform handles over 200 language model training tasks of varying structures and scales daily, including dialogue generation, summary extraction, named entity recognition, and code generation. The model architectures involved include Transformer, BERT, GPT, and T5, and the training framework is based on a hybrid deployment of PyTorch and TensorFlow. Previously, the platform's scheduling strategy primarily relied on static priority and time-based polling, which resulted in problems such as concentrated model scheduling, frequent training segment interruptions, and high task switching overhead, impacting overall training throughput and hardware resource utilization efficiency.

[0039] After introducing the reinforcement learning-based intelligent training task scheduling method for language models proposed in this invention, the system dynamically generates policies and adaptively allocates tasks for language model training through an improved PPO model. Task attributes are uniformly structured and input into the scheduling model. The scheduling model continuously optimizes the policy based on indicators such as training efficiency, parameter stability, and scheduling smoothness fed back by the smoothed reward vector, ultimately achieving task granularity adjustment, conflict buffer management, and policy self-updating.

[0040] The specific deployment process of this invention includes: collecting task execution logs and model training feedback from the past two months, constructing task status labels and reward templates; initializing the improved PPO model and performing 10 rounds of pre-training; embedding the model into the platform scheduling module and setting the task collection cycle to every 15 minutes, the minimum unit for training segment segmentation to 300 seconds, the entropy control shrinkage threshold to 0.7, and the adaptive shearing adjustment threshold to 0.25; after formal deployment, the system runs continuously for three months, and the method of this invention is compared with the traditional static scheduling method to generate the following experimental data.

[0041] In actual operation, the scheduling system processes an average of about 240 training tasks per day. For ease of observation and analysis, the following data are taken from typical representative time windows (Month 1, Month 2, and Month 3), and the comparison items include average training segment execution success rate, average loss jitter value after switching, GPU utilization, and average single round scheduling time.

[0042] Table 1 Comparison of Task Scheduling and Operation Indicators As shown in Table 1, the method of this invention achieves significant performance improvement in training task scheduling. Traditional static scheduling methods suffer from high interruption rates in training segments due to their inability to dynamically adapt to training rhythms and model differences between tasks, resulting in an average execution success rate of only 83.5%. In contrast, the method of this invention, by introducing a training duration segmentation mechanism and a task rhythm conflict control mechanism, steadily improves the execution success rate to 94.2%. Furthermore, in task switching scenarios, traditional methods lack buffering mechanisms, leading to large short-term fluctuations in loss values. This method, through training segment buffering and gradient fusion strategies, effectively reduces the post-switching loss jitter from 0.035 to 0.011, improving training stability. Regarding resource utilization, this invention significantly improves GPU utilization efficiency from 71.3% to 87.8% through scheduling action probability optimization and policy gradient updates, while simultaneously reducing scheduling time from 142.6ms to 89.7ms, resulting in an overall training throughput improvement of approximately 24%.

[0043] To verify the performance of the model scheduling in terms of task distribution diversity and policy adaptation, the distribution of task types processed by the system, the average number of training rounds per task, and the proportion of repetitive task scheduling were statistically analyzed over three training cycles.

[0044] Table 2 Comparison of Scheduling Behavior Diversity and Adaptability Table 2 reflects the continuous improvement of the scheduling system in terms of task diversity and policy update capability as the training rounds increase. In the early part of month 1, the scheduling strategy was still in the initial adaptation phase, with a high concentration of task types and significant repetitive training behavior, with a repetitive task scheduling ratio as high as 18.4%. However, guided by the improved PPO model, the system continuously strengthened its focus on low-frequency tasks by dynamically adjusting the scheduling distribution and optimizing the reward mechanism. By month 3, the number of different task types had expanded to 11, while significantly reducing the repetitive scheduling ratio to 7.3%. The average number of training rounds per task increased from 4.6 to 5.8, indicating that the task segmentation mechanism and task delay mechanism effectively improved segment coverage and training depth. Furthermore, the policy update frequency increased from 1.2 times per thousand steps to 3.6 times, indicating that the scheduling model gradually formed an adaptive scheduling strategy during long-term training, not only responding to the immediate training state but also predicting the phased scheduling offset trend.

[0045] This embodiment fully verifies the adaptability and efficiency of the present invention in language model training scenarios. Through multi-dimensional reward design, entropy control compression mechanism, and task fragment buffering mechanism, it solves the problems of unbalanced task scheduling, unstable fragment execution, and waste of training resources in traditional scheduling methods. It has significant advantages in improving training success rate, scheduling stability, and resource utilization, and has good engineering application value and promotion prospects.

[0046] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for intelligent training task scheduling of language models based on reinforcement learning, characterized in that, Includes the following steps: Step 1: Construct a task pool based on language model training tasks, collect information of each task in the task pool, and generate a set of task attributes; Step 2: Perform rhythm encoding on the attribute fields in the task attribute set, combine historical scheduling records with memory-first processing, and generate a training state vector; Step 3: Combine the training state vectors in a set order to form a scheduling input sequence, input it into the improved PPO model, and output the scheduling action distribution; Step 4: If the training duration of the selected training task is greater than the preset duration threshold, then the selected training task is split, and segment boundary marker information is inserted into the scheduling input sequence to generate a set of scheduling actions. Step 5: Calculate the conflict weight based on the step size difference, training duration difference, and rhythm difference between the trained task and the previous training task, and perform conflict constraint processing on the scheduling action distribution; if there is a task switching time, the improved PPO model generates a cache strength indication signal and triggers a buffer control operation. Step 6: Collect scheduling behavior records and training feedback information, calculate training efficiency score, parameter stability score, task switching smoothness score and task distribution diversity score, perform multi-reward fusion processing to form a smooth reward vector; Step 7: Execute the policy update operation of the improved PPO model based on the smoothed reward vector; collect the task state information and rhythm parameters of the trained task, refresh the task pool, and construct a new round of scheduling input sequence.

2. The intelligent training task scheduling method for language models based on reinforcement learning according to claim 1, characterized in that, Step one specifically involves: Based on the language model training task, a task pool is constructed according to a preset task structure; input requests for multiple language model training tasks are received, and task identifiers are registered according to the task generation order; Information collection is performed on each task in the task pool; the information of the task includes task type, training duration, loss change, resource consumption and scheduling interval information; the task type records task category information, the training duration records the historical training time of the task, the loss change records the changing trend of the loss function in the continuous training phase, the resource consumption records the computational resource consumption information when the task is executed, and the scheduling interval information records the time interval between two scheduling of the task. Perform a uniform formatting operation on the collected task information, mapping each piece of information to task attribute fields; group and combine the formatted attribute fields according to the task identifier to construct a task attribute set.

3. The intelligent training task scheduling method for language models based on reinforcement learning according to claim 2, characterized in that, The training state vector includes task type encoding, rhythm feature encoding, cost encoding, loss trend encoding, and time interval encoding, specifically: Perform rhythm encoding on each set of attribute fields in the task attribute set, calculate the rhythm fluctuation amplitude and periodic parameters of the task based on the training duration and scheduling interval information, and generate rhythm feature encoding; Based on task type, resource consumption and loss change information, a task feature sequence is constructed, and structured mapping rules are called to generate task type code, cost code and loss trend code; Collect historical scheduling records of tasks, and statistically analyze the selection frequency and relative position of each task in different scheduling cycles to construct a scheduling preference sequence; The scheduling preference sequence is processed by memory priority to generate a scheduling preference vector, which is then combined with the aforementioned encoded information. Generate time interval codes based on the time interval between two task schedulings; The task type encoding, rhythm feature encoding, cost encoding, loss trend encoding, and time interval encoding are concatenated into a training state vector according to a set structure, and the mapping relationship between the training state vector and the task identifier is recorded.

4. The intelligent training task scheduling method for language models based on reinforcement learning according to claim 3, characterized in that, The step of combining the training state vectors in a predetermined order to form a scheduling input sequence is specifically as follows: Collect all training state vectors in the task pool, set sorting priority parameters according to task identifier order, task historical scheduling frequency and remaining training time, perform priority sorting on the training state vectors, and generate a sorting queue. Based on the order of the sorted queue, read the contents of the training state vectors sequentially and construct a vector permutation list; Perform structural unification processing on the vector permutation list, set unified input dimension standards and position information identification rules, and generate a structured vector matrix; Perform positional encoding and sequence normalization operations on the structured vector matrix to generate a scheduling input sequence with sequential encoding; record the index mapping information between the scheduling input sequence and the original training state vector to construct a scheduling input sequence index table.

5. The intelligent training task scheduling method for language models based on reinforcement learning according to claim 4, characterized in that, The improved PPO model includes an input encoding structure, an action prediction structure, a state evaluation structure, and a policy update structure, specifically: The input encoding structure receives the scheduled input sequence, performs embedding mapping processing, sequential position mapping processing, and normalization processing, and generates an encoded input tensor. The action prediction structure receives the encoded input tensor and outputs a scheduled action distribution based on a multi-layer feedforward neural structure and an action decision distribution function; the scheduled action distribution includes an action probability vector corresponding to the training task number and a boundary identifier component of the task segment boundary; The state evaluation structure receives the encoded input tensor, constructs a state value function path, and outputs the state value estimate corresponding to the training state. The strategy update structure includes an entropy-controlled contraction unit and a slow-cut update unit; the entropy-controlled contraction unit performs dynamic compression processing on the entropy range of the scheduling action distribution to adjust the convergence amplitude of the strategy distribution; the slow-cut update unit includes a shearing amplitude adjustment structure and a parameter delay control structure, performs shearing boundary updates based on the policy change amplitude, and performs step-by-step update processing based on the update delay period; The improved PPO model constructs a smooth reward vector based on training efficiency score, parameter stability score, task switching smoothness score, and task distribution diversity score collected during training. It calculates the advantage function by combining historical trajectory records, performs joint optimization on the parameters of the action prediction structure and the state evaluation structure, and outputs an updated scheduling strategy.

6. The intelligent training task scheduling method for language models based on reinforcement learning according to claim 5, characterized in that, Step four specifically involves: A duration comparison operation is performed based on the training duration of the selected training task and a preset duration threshold. When the training duration exceeds the preset duration threshold, the segmentation logic function is called to divide the selected training task into multiple training segments according to the set segment length parameter and segment boundary rules. Each training segment is assigned a unique task segment identifier, and segment boundary marker information is inserted into the scheduling input sequence; the segment boundary marker information includes the training segment start position, end position, task segment number, and segment order label; Establish a mapping relationship between the training fragment identifier and the original task identifier, update the structure field of the corresponding task in the task attribute set, and generate a fragment structure attribute set; The fragment structured attribute set is used as input to construct a fragment-level training state vector and participate in the encoding process of scheduling the input sequence; During the scheduling action output phase, the action probability vector and boundary identifier component are jointly decoded to generate a set of scheduling actions containing training segment numbers and segment boundary indicators.

7. The intelligent training task scheduling method for language models based on reinforcement learning according to claim 6, characterized in that, Step five specifically involves: Collect the training segment numbers and training state vectors of the training task and the previous training task; based on the step size information, training duration information and rhythm feature information recorded in the training state vector, calculate the step size difference, training duration difference and rhythm difference between the two tasks. Based on the preset conflict weight calculation function, the weight combination mechanism is called to map the step size difference, training duration difference and rhythm difference into conflict weight values. The level of switching conflict between scheduling actions is determined based on the conflict weight value; Based on the switching conflict level, the conflict suppression rules are invoked to perform conflict constraint processing on the action probability vector in the scheduling action distribution; the conflict constraint processing includes probability compression, candidate elimination and order priority adjustment operations, which are executed in sequence according to the conflict level. If a task switching time is determined to exist, the improved PPO model outputs a cache strength indication signal; Cache strength indication signals include switching level, buffer duration, and fusion bias label; Buffer control operations are triggered based on the cache strength indicator signal; The buffer control operation includes: saving the gradient state of the current training task, temporarily storing the gradient state in the buffer area, setting the parameter delay update period, calling the gradient fusion function to perform the fusion of the old and new gradients; and binding the buffered gradient state with the policy update structure.

8. The intelligent training task scheduling method for language models based on reinforcement learning according to claim 7, characterized in that, Step six specifically involves: The scheduling behavior records and training feedback information generated during the training process are collected to construct a behavior feedback set; the scheduling behavior records include scheduling action number, task switching time point, training segment identifier and scheduling interval; the training feedback information includes task loss value change, gradient convergence status, resource consumption record and training completion status. Training efficiency scores are calculated based on behavioral feedback sets. The training efficiency score is quantified based on the ratio of the decrease in loss value per unit time to the execution time of the training segment; Calculate the stable score of the parameters; The parameter stability score is evaluated based on the gradient variance and parameter change rate between consecutive segments; Calculate the task switching smoothness score; the task switching smoothness score is calculated based on the degree of resource consumption jump at the switching point and the stability of model output; Calculate the task distribution diversity score; The task distribution diversity score is based on the balanced distribution of task numbers in historical scheduling. A weighted fusion operation is performed on the training efficiency score, parameter stability score, task switching smoothness score, and task distribution diversity score to construct the initial value of the reward vector; The initial value of the reward vector is subjected to temporal smoothing and anomaly compensation processing to generate a smoothed reward vector. The smoothing process includes sliding window weighting, score normalization and reward clipping, and the anomaly compensation process includes zero-reward filling and spike correction. The smoothed reward vector is then associated with the training task identifier corresponding to the scheduling input sequence.

9. The intelligent training task scheduling method for language models based on reinforcement learning according to claim 8, characterized in that, The policy update operation based on the smoothed reward vector to perform the improved PPO model includes adaptive entropy-controlled contraction and gradual shearing update; the gradual shearing update includes adaptive shearing amplitude adjustment and grouped delay processing, specifically: A sequence of advantage function estimation is constructed based on smoothed reward vectors and historical scheduling trajectory data; the action prediction structure is invoked to output the distribution of scheduling actions, and the state evaluation structure is invoked to output the state value estimate; the policy loss function and the value loss function are calculated to form a joint optimization objective. The adaptive entropy control shrinkage operation dynamically adjusts the entropy value change range of the scheduling action distribution; the shrinkage coefficient is determined according to the entropy value change rate to compress the dispersion of the policy distribution; the shrinkage policy gradient is generated based on the compression result, and the update amplitude of the policy distribution is adjusted. The adaptive shearing amplitude adjustment operation monitors the change amplitude of the strategy and the fluctuation trend of the advantage function, and calculates the shearing factor; based on the shearing factor, the shearing boundary range is set to constrain the gradient change of the joint optimization objective. By using a grouped update operation, the parameters of the improved PPO model are divided into multiple update groups based on the rhythm information and parameter change frequency of the training task; the parameters of each group are updated step by step according to the preset update cycle; and the corresponding parameters in the action prediction structure and state evaluation structure are adjusted synchronously within each update cycle. The compression policy gradient of the adaptive entropy control shrinkage output and the constraint gradient of the gradual update output are fused to generate a joint optimization gradient tensor. The parameter state of the improved PPO model is updated based on the joint optimization gradient tensor to complete the policy update operation.

10. The intelligent training task scheduling method for language models based on reinforcement learning according to claim 9, characterized in that, The process of collecting the task state information and rhythm parameters of the trained task, refreshing the task pool, and constructing a new round of scheduling input sequence is as follows: The system collects the status information and rhythm parameters of the trained task after the scheduling and training operations are completed. The task status information includes the training segment number, training progress, loss change trend, and parameter convergence rate. The rhythm parameters include the training time update value, the scheduling interval update value, and the rhythm change amplitude. Update the task attribute fields based on the task status information and rhythm parameters, and reconstruct the task attribute set; The task identifier is determined based on the task completion status. If the task training is complete, the corresponding identifier is removed from the task pool. For incomplete tasks, the identifier is retained and its status is updated to pending scheduling. Perform rhythm encoding and memory-first processing on the updated task attribute set to construct a new set of training state vectors; The training state vector set is combined and sorted according to the set order and scheduling priority to generate a new round of scheduling input sequence; Record the index mapping relationship of the scheduling input sequence and establish an input connection with the improved PPO model input encoding structure to enter the next scheduling cycle.