Multi-model collaborative operation method based on large model efficient training

Through resource monitoring, heterogeneous resource spatio-temporal graph construction, dynamic resource redistribution and cross-model knowledge distillation, the problems of resource waste and repeated calculations in multi-model training are solved, and efficient and stable multi-model collaborative training is achieved.

CN120295784APending Publication Date: 2025-07-11SHENZHEN BOAN CLOUD TECHNOLOGY CO LTD
View PDF 0 Cites 20 Cited by

Patent Information

Application Number
CN202510370153.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In multi-model parallel training scenarios, traditional resource allocation strategies lead to resource waste and repeated calculations, reducing training efficiency and cost, and lacking an effective cross-model knowledge sharing mechanism.

Method used

A multi-model collaborative operation method based on large models is adopted to optimize resource allocation and knowledge sharing through resource monitoring, heterogeneous resource spatio-temporal graph construction, dynamic resource redistribution, cross-model knowledge distillation and adaptive collaborative training.

Benefits of technology

It significantly improves resource utilization efficiency, reduces repeated calculations, improves model training efficiency and performance, and ensures training stability and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295784A_ABST
    Figure CN120295784A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-model collaborative operation method based on large-model efficient training, and relates to the technical field of artificial intelligence, and the method comprises the steps: firstly, monitoring model training resource indexes, such as activation tensor memory occupation and gradient calculation intensity, and carrying out the optimization through a three-stage optimization mechanism; constructing a heterogeneous resource space-time diagram to predict resource requirements, and combining multi-modal feature dynamic fusion and prediction error online compensation to improve prediction accuracy; then dynamically allocating resources, determining a migration priority by evaluating a training income gradient, an energy consumption efficiency factor and the like, and implementing video memory block management and pipeline scheduling; carrying out cross-model hierarchical knowledge distillation and self-adaptive cooperative training, and adjusting a cooperative mode according to model similarity; the problems of heterogeneous model resource competition and isolated knowledge precipitation are effectively solved, the resource utilization rate is improved, repeated calculation is reduced, the overall performance and robustness of multi-model collaborative operation are enhanced, the training cost is reduced, and the training efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and specifically provides a multi-model collaborative operation method based on efficient training of large models. Background Art

[0002] In the scenario of multi-model parallel training, traditional fixed resource allocation strategies have exposed serious drawbacks. During the training process of different models, the demands for GPU memory and computing units in the forward propagation, backward propagation, and parameter update phases vary greatly. During peak periods, the resource utilization rate can exceed 90%, while during low periods, it is below 0%. Such drastic fluctuations result in an overall resource waste rate as high as 42%, with a large amount of computing resources wasted during idle periods, seriously reducing the resource utilization efficiency, increasing the training cost, and restricting the overall performance improvement of multi-model training.

[0003] When multiple models are independently trained in the same task domain, it is found that there is an overlapping area of more than 78% in the intermediate layer feature distributions. However, there are significant deficiencies in cross-model knowledge sharing in the existing technology, lacking effective mechanisms to utilize these overlapping features. This leads to a large amount of duplicate calculations during the model training process, with the amount of duplicate calculations increasing by 3.2 times, not only wasting computing resources but also prolonging the training time, hindering the improvement of model training efficiency and performance optimization.

[0004] The problems of heterogeneous model resource competition and isolated knowledge precipitation seriously affect the effect and efficiency of multi-model collaborative training. The waste of resources increases the training cost, and the duplicate calculations prolong the training cycle, making it difficult for model training to fully utilize the potential of hardware resources and difficult to complete complex tasks quickly and efficiently. Therefore, there is an urgent need for an innovative method that can effectively solve these problems, achieve optimal resource allocation and knowledge sharing during the training process of multi-models, and improve the overall performance of multi-model collaborative operation.

[0005] To address the above deficiencies, a technical solution is provided herein. Summary of the Invention

[0006] The purpose of the present invention is to solve the resource waste caused by heterogeneous model resource competition in traditional multi-model parallel training, as well as the problem of a 3.2-fold increase in duplicate calculations caused by isolated knowledge precipitation during the independent training of multi-models in the same task domain, and to propose a multi-model collaborative operation method based on efficient training of large models.

[0007] The purpose of the present invention can be achieved through the following technical solutions:

[0008] A multi-model collaborative operation method based on efficient training of large models, comprising the following steps:

[0009] S1: Resource monitoring, monitor the memory occupancy of the activation tensors during the forward propagation of model training, the intensity of gradient calculation during the backward propagation, and the memory requirements of the optimizer state during the parameter update phase. And adopt a three-level optimization mechanism to optimize the monitoring of the gradient calculation intensity;

[0010] S2: Resource demand prediction, construct a heterogeneous resource spatio-temporal graph, analyze the resource demand through the spatio-temporal graph attention network and the dynamic fusion of multi-modal features, and establish a closed-loop feedback mechanism to compensate for the prediction error;

[0011] S3: Dynamic resource reallocation, deploy a video memory demand prediction model, analyze the factors affecting the migration priority of parameter evaluation, and perform video memory block management and pipeline scheduling to improve the resource allocation efficiency;

[0012] S4: Cross-model hierarchical knowledge distillation, use BERT-Large as the teacher model, and promote the student model to learn effective features by establishing a knowledge feature alignment layer, a multi-granularity knowledge distillation strategy, and a progressive knowledge fusion;

[0013] S5: Adaptive collaborative training, construct a task similarity evaluation matrix, dynamically adjust the collaboration mode according to the similarity, and elastically schedule the training to improve the model training effect.

[0014] Furthermore, the specific process of S1 is as follows:

[0015] Monitor the resource metrics during the model training process, including:

[0016] The memory occupancy of the activation tensors in the forward propagation stage. During the forward propagation of the model, use a memory monitoring tool or the memory analysis interface provided by the programming language to obtain the information on the memory occupied by the activation tensors;

[0017] The intensity of gradient calculation in the backward propagation stage. Use the tools provided by CUDA programming to record the execution time of the Kernel; start timing before the backward propagation starts, and stop timing when the CUDA Kernel involving gradient calculation in the backward propagation is completed. Thus, the elapsed time of this Kernel is obtained, and the intensity of gradient calculation in the backward propagation stage is calculated by measuring the elapsed time of the CUDA Kernel; and optimize it using a three-level optimization mechanism;

[0018] The memory requirements of the optimizer state in the parameter update phase. When using the Adam optimizer to update the parameters during the model training, the Adam optimizer will maintain the momentum variable, and read the memory size occupied by the data structure by accessing the momentum variable attribute in the optimizer object.

[0019] Furthermore, the process of the three-level optimization mechanism in S1 is as follows:

[0020] Hardware-level Execution Trace Tracking: Through the PC Sampling hardware module of the NVIDIA Ampere architecture, directly collect the number of execution cycles of computing instructions inside the stream processor, and record the actual operation duration of each gradient calculation Kernel with a precision of 10 nanoseconds; for GPUs with Volta architecture and above, activate the CycleCounter register of each SM cluster to capture the full process clock cycles from instruction emission to result write-back in real time, eliminating the errors introduced by CUDA runtime scheduling; establish a Kernel fingerprint library, perform instruction pattern recognition on the backpropagation operators in the cuDNN library, automatically label the gradient calculation core Kernel and isolate auxiliary memory operations;

[0021] Deploy a gap monitor at the CUDA stream level. When an execution interval exceeding 5 microseconds is detected between two adjacent Kernels, automatically trigger the following compensation strategies:

[0022] For the waiting gaps caused by resource contention, by analyzing the shared memory bank conflict pattern, calculate the difference between the theoretical maximum throughput and the actual throughput, and deduct the invalid waiting duration proportionally;

[0023] For the forced gaps caused by data dependencies, establish a computational graph dependency relationship model. Based on the output tensor size and video memory bandwidth of the predecessor Kernel, deduce the minimum necessary transmission time and eliminate the excessive waiting from the total elapsed time;

[0024] Implement runtime gap classification compensation. When the gap type is identified as video memory controller contention, start the hardware performance counter to track the arbitration cycle count of the memory controller and quantify the latency caused by the conflict;

[0025] Adaptive Sparse Monitoring Strategy: In the initial stage, adopt the full-scale monitoring mode to perform hardware-level cycle counting on each gradient calculation Kernel;

[0026] When the calculation intensity volatility in 10 consecutive iteration cycles is lower than 8%, switch to the critical path sampling mode, that is, only continuously track the Kernels that consume the top 20% of the computing resources;

[0027] Establish a Kernel elapsed time prediction model. Based on the LSTM network to analyze the historical execution time series, when the prediction error is less than 5%, use the predicted value to replace the actual measured value;

[0028] Inject a monitoring agent at the PyTorch framework level to automatically identify the critical path of gradient calculation, that is, the set of Kernels that contribute more than 85% of the total duration, maintain a sampling interval of 200 microseconds, and switch to millisecond-level sparse sampling for non-critical paths.

[0029] Furthermore, the specific operation steps of S2 are as follows:

[0030] Construction of heterogeneous resource spatio-temporal graph: Divide the 40GB video memory of NVIDIA A100 GPU into 256 dynamic units. Each unit is abstracted as a graph node and assigned features such as real-time video memory occupancy, historical occupancy fluctuation, and data transfer frequency between adjacent units. At the same time, CUDA cores are added as independent nodes and assigned features such as SM cluster utilization rate and TensorCore activation status, and two-way edge connections are established; the feature data comes from step S1. Through multi-dimensional characterization of the video memory and computing core status, it provides an information basis for subsequent analysis of resource requirements;

[0031] Spatio-temporal graph attention network: Modeling in the time dimension, using dilated causal convolution to process historical data, with the dilation factor increasing according to the exponential sequence (1, 2, 4), and the convolutional kernel covering 15 iteration cycles to capture long-range time dependencies;

[0032] Modeling in the space dimension, deploying a multi-head graph attention mechanism, setting three dedicated attention heads:

[0033] Head 1 focuses on the Bank conflict pattern between video memory units and calculates the attention weight

[0034] α1 = softmax(LeakyReLU(W1·[h i ||h j ));

[0035] Head 2 analyzes the load balance between computing cores and video memory units, and the weight

[0036] α2 = softmax(LeakyReLU(W2·[h i ⊙h j ));

[0037] Head 3 models the cross-model data transfer requirements, and the weight

[0038]

[0039] Among them, α1, α2, and α3 respectively represent the attention weights calculated by the three dedicated attention heads; softmax is an activation function used to convert the input numerical values into a probability distribution, making all output values between 0 and 1 and the sum equal to 1; LeakyReLU is a variant of the rectified linear unit activation function, which solves the problem that neurons are not activated when the input is negative. The formula is where γ is a positive number; W1, W2, and W3 are all learnable parameter matrices; h i , h j are the feature vectors of graph nodes i and j; [·||·] is the vector concatenation operation; [·⊙·] is the element-wise multiplication operation of vectors; It is a feature fusion operation defined according to specific modeling requirements, used to combine the feature vectors of node i and node j to capture information related to cross-model data transmission requirements;

[0040] Multi-modal feature dynamic fusion: Introduce an external feature encoder to process two types of information: Model structure features: The number of Transformer layers and the number of attention heads are mapped to 128-dimensional embedding vectors;

[0041] Task priority signal: Normalize it to a dynamic weight coefficient in the range of [0.1, 1.0]; Design a gated fusion unit to achieve feature fusion: Calculate the gate value: g = σ(W g1 ·[h graph ||h ext +b g ); Feature fusion: h final = g ⊙ h graph +(1 - g) ⊙ h ext ; where g is the gate value, a scalar between 0 and 1; σ is the sigmoid function, which maps the input value to the range of 0 to 1; W g1 is a learnable weight matrix; h graph is the feature vector output by the spatio-temporal graph attention network, containing feature information related to the time and space dimensions extracted from the video memory unit and the computing core node; h ext is the feature vector obtained after the external feature encoder processes the model structure features and the task priority signal; b g is the bias vector; h final is the finally fused feature vector; ⊙ is the Hadamard product, that is, the element-wise multiplication operation of vectors; h ext is the feature vector obtained after the external feature encoder processes the model structure features and the task priority signal;

[0042] Online compensation for prediction error: Establish a closed-loop feedback mechanism. When the deviation between the measured video memory occupancy and the predicted value exceeds 5%, dynamically reconstruct the spatio-temporal graph structure and increase the connection weight between the abnormal unit and the associated core by 300%;

[0043] Inject Gaussian noise perturbation to enhance the robustness of the model; Perform lightweight distillation every 30 minutes: Compress the prediction model to 40% of the original size using temperature scaling and synchronize it to the edge scheduling node.

[0044] Furthermore, the specific operation steps of S3 are as follows:

[0045] Deploy a video memory demand prediction model. By analyzing the features in the model training stage, the historical resource occupancy pattern, and the task queue metadata, generate a heat map of the video memory demand within the next 1 second 300 milliseconds in advance;

[0046] When it is predicted that the model is about to enter the parameter update phase, actively trigger the pre-labeling process of low-priority video memory blocks, and use the NVLink DMA engine to pre-load migration data packets to reduce the migration preparation time; the migration priority is comprehensively evaluated through a dynamic evaluation mechanism;

[0047] The video memory block management adopts an elastic adaptive strategy. The basic division is 256 units of 156 MB, and it supports dynamic merging into large blocks of 312 MB or 624 MB;

[0048] When it is detected that the model requests more than 256 MB of video memory, the merging protocol automatically integrates adjacent free units and optimizes the video memory space through a periodic defragmentation engine;

[0049] The defragmentation uses the ant colony algorithm to calculate the minimum movement path, which is executed once every 5 minutes to reduce the video memory fragmentation rate; combined with the LSTM network to predict the video memory demand distribution in the next 10 iteration cycles and reserve continuous large blocks of space in advance;

[0050] The pipeline scheduling integrates real-time resource monitoring data to construct a multi-dimensional resource status matrix, covering the utilization rate of SM clusters, the memory bank-level latency, and task dependencies;

[0051] The dynamic priority scheduling strategy injects low-priority preprocessing tasks when the idle rate of the GPU computing unit exceeds 25%, and overlaps the data transfer and computing tasks when the video memory bandwidth utilization rate is lower than 40%;

[0052] The critical path tasks adopt round-robin scheduling to ensure the minimum latency of core operations such as gradient calculation;

[0053] The fault tolerance and recovery mechanism uses the FPGA coprocessor to real-time verify the calculation results. When the error rate exceeds 0.1%, it automatically rolls back to the stable state and reconstructs the pipeline to optimize and compress the idle rate.

[0054] Furthermore, the specific operation steps for the comprehensive evaluation of the migration priority in S3 through the dynamic evaluation mechanism are as follows:

[0055] Analyze the influencing parameters, including:

[0056] Training benefit gradient TD: Calculate the exponential moving average of the loss values of the past 10 batches. The formula is: where i is the batch index, taking values from 1 to 10, representing the past 10 batches; Δloss i is the change in the loss value of the i-th batch;

[0057] Energy consumption efficiency factor EZ: Obtain the GPU power consumption GF and computing throughput FLOP in real time through the NVIDIA SMI interface S, calculate the energy consumption efficiency, which is used to quantify the energy efficiency ratio of task execution;

[0058] Dependency strength YQ: Analyze the parameter sharing rate and data exchange frequency between models, and define:

[0059]

[0060] Design a comprehensive evaluation function for design priorities, integrating the video memory requirement R 显存 , task priority Q 优先级 and influencing parameters: Pri = τ1×R 显存 +τ2×A 优先级 +τ3×TD + τ4×EZ + τ5×YQ. The initial weights are set as τ1 = 0.3, τ2 = 0.25, τ3 = 0.2, τ4 = 0.15, τ5 = 0.1, and online learning adjustment is supported; when the prediction error exceeds the threshold, trigger the backpropagation algorithm to dynamically optimize the weight allocation;

[0061] When the prediction model A enters the parameter update stage, decide whether to trigger video memory migration based on the priority score: when Pri 模型B > 1.5×Pri 模型A , immediately start the migration;

[0062] Otherwise, delay the migration until model A completes the key calculation stage; preload the migration data packet through the NVLink DMA engine, and dynamically adjust the transmission bandwidth in combination with the priority score to ensure that the data packets of high-priority tasks are transmitted first.

[0063] Furthermore, the specific operation steps of S4 are as follows:

[0064] Establish a knowledge feature alignment layer: Select BERT-Large with the largest number of parameters in the model set as the teacher model; perform feature mapping on the 6th, 12th, and 18th layer Transformer modules of other student models;

[0065] Use the Wasserstein distance to measure the similarity of the output distributions of each layer, and set the threshold to 0.35. When the similarity is lower than the set threshold of 0.35, perform subsequent knowledge distillation operations;

[0066] Multi-granularity knowledge distillation strategy: During the forward propagation of each layer, calculate the KL divergence loss between the student model and the corresponding layer of the teacher model, with the temperature parameter τ = 3; construct a joint probability distribution at the model output layer, and use the Jensen-Shannon distance to constrain the logits difference; introduce an adaptive weight adjuster to dynamically allocate loss weights according to the alignment degree of each layer;

[0067] Progressive knowledge fusion: In the first stage, the underlying features are forced to align, and the learning rate is set to 1.2 times the benchmark value; in the second stage, the middle-level constraints are relaxed, a random dropout mechanism is introduced, and the dropout rate is 15%; in the third stage, soft target supervision is only implemented at the output layer, and the differences in high-level features are retained.

[0068] Further, the specific operation steps of S5 are as follows:

[0069] Construct a task similarity evaluation matrix: Extract the feature representations of each model on the validation set; calculate the cosine similarity matrix S ∈ R N×N ; R is the set of real numbers, and N is the number of models; when S > 0.7, it is marked as a similar task group;

[0070] Dynamically adjust the collaboration mode: For the high similarity group, that is, when S > 0.7, start the parameter sharing mechanism and freeze 60% of the common parameters at the bottom layer;

[0071] For the medium similarity group, that is, when 0.4 < S ≤ 0.7, enable gradient cross-accumulation and synchronize the gradients every 3 batches;

[0072] For the low similarity group, that is, when S ≤ 0.4, implement adversarial training and generate perturbation samples through the gradient reversal layer;

[0073] Elastic training scheduling: When it is detected that the loss reduction rate of a certain model is lower than the average value by 30%, automatically trigger the collaboration mode switch;

[0074] Deploy a priority queue in the NVIDIA DGX cluster to allocate an additional 16% of computing resources for the model; achieve cross-node gradient aggregation through the RDMA network, and control the latency within 2 ms.

[0075] Compared with the prior art, the beneficial effects of the present invention are:

[0076] (1) In the present invention, through the resource monitoring and dynamic resource reallocation mechanism, the resource utilization efficiency is significantly improved; in the resource monitoring link, the resource indicators at each stage of model training are accurately monitored, and the three-level optimization mechanism is used to optimize the gradient calculation intensity monitoring, providing a basis for accurate resource allocation; during dynamic resource reallocation, a video memory demand prediction model is deployed to plan the video memory usage in advance, and video memory migration is carried out according to the comprehensively evaluated migration priority. Combining the video memory block management and pipeline scheduling strategies, resource waste is reduced; for example, by dynamically merging video memory units and defragmenting, the success rate of large video memory allocation is improved; tasks are scheduled according to the real-time resource status, reducing the idle time of computing units and video memory bandwidth, balancing the utilization rates of GPU video memory and computing units, effectively avoiding the problems of resource shortage during peak periods and resource idleness during low periods under the traditional fixed resource allocation strategy, and significantly reducing the overall resource waste rate;

[0077] (2) In the present invention, remarkable achievements have been made in cross-model knowledge sharing and training optimization. The cross-model hierarchical knowledge distillation mechanism uses the BERT-Large with the largest number of parameters as the teacher model. Through the knowledge feature alignment layer, multi-granularity knowledge distillation strategy, and progressive knowledge fusion, it guides the student model to learn the effective features of the teacher model, reducing the redundant calculation in the overlapping area of the intermediate layer feature distributions. In the adaptive co-training, a task similarity evaluation matrix is constructed, and the collaboration mode is dynamically adjusted according to the similarity between models. For example, high-similarity groups share parameters, medium-similarity groups cross-accumulate gradients, and low-similarity groups perform adversarial training. At the same time, elastic training scheduling allocates more resources to key models, promoting collaborative learning between models, improving training efficiency and model performance, and avoiding the problem of redundant calculation caused by isolated knowledge precipitation.

[0078] (3) The present invention comprehensively improves the overall performance and robustness of multi-model collaborative operation. The resource demand prediction module accurately predicts resource demands through a spatio-temporal graph attention network and dynamic multi-modal feature fusion, and enhances the adaptability of the model to environmental changes through a prediction error online compensation mechanism. During the training process, combined with dynamic resource reallocation and adaptive co-training, it not only optimizes resource allocation but also improves the training effect of the model. The fault tolerance and recovery mechanism uses an FPGA co-processor to verify the calculation results in real time. When the error rate exceeds the threshold, it automatically rolls back and reconstructs the pipeline, ensuring the stability and reliability of the training, enabling the entire multi-model collaborative operation system to operate efficiently and stably in a complex environment, and providing a reliable technical guarantee for multi-model collaborative training based on large models. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] For the convenience of those skilled in the art to understand, the present invention will be further described below with reference to the accompanying drawings;

[0080] Figure 1 is the overall block diagram of the system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0081] The technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts fall within the scope of protection of the present invention.

[0082] It should be understood that the terms "including" and "comprising" used in the specification and claims of this disclosure indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their collections.

[0083] It should also be understood that the terms used in this disclosure specification are for the purpose of describing specific embodiments only and are not intended to limit the disclosure. As used in this disclosure specification and the claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms. It should further be understood that the term "and / or" used in this disclosure specification and the claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations.

[0084] As Figure 1 shown, a multi-model collaborative operation method based on efficient training of a large model includes the following steps:

[0085] Step 1, resource monitoring, monitoring resource metrics such as the memory occupancy of the activation tensor during forward propagation in model training, the gradient calculation intensity during backpropagation, and the memory requirements of the optimizer state during the parameter update phase, and using a three-level optimization mechanism to optimize the monitoring of the gradient calculation intensity;

[0086] Monitor the resource metrics during the model training process, including:

[0087] The memory occupancy of the activation tensor (accuracy 0.1 GB) in the forward propagation stage. During the forward propagation of the model, use a memory monitoring tool or the memory analysis interface provided by the programming language to obtain information on the memory occupied by the activation tensor; the gradient calculation intensity in the backpropagation stage, using the tool provided by CUDA programming to record the execution time of the Kernel. Before the backpropagation starts, start timing. When the CUDA Kernel involving gradient calculation in the backpropagation is completed, stop timing. Thus, the elapsed time of this Kernel is obtained. By measuring the elapsed time of the CUDA Kernel, the gradient calculation intensity in the backpropagation stage is calculated; use the three-level optimization mechanism for optimization, and the process is as follows:

[0088] Hardware-level Execution Trace Tracing: Through the PCSampling hardware module of the NVIDIA Ampere architecture, directly collect the number of computing instruction execution cycles inside the streaming multiprocessors (SMs), and record the actual operation duration of each gradient calculation Kernel with a precision of 10 nanoseconds. For GPUs of Volta architecture and above, activate the CycleCounter register of each SM cluster to capture the full process clock cycles from instruction emission to result write-back in real time, eliminating the errors introduced by CUDA runtime scheduling. Establish a Kernel fingerprint library to perform instruction pattern recognition on the backward propagation operators in the cuDNN library, automatically label the gradient calculation core Kernel and isolate auxiliary memory operations; deploy a gap monitor at the CUDA stream level. When an execution interval exceeding 5 microseconds is detected between two adjacent Kernels, automatically trigger the following compensation strategies: For the waiting gap caused by resource contention, by analyzing the shared memory bank conflict pattern, calculate the difference between the theoretical maximum throughput and the actual throughput, and deduct the invalid waiting duration proportionally; for the forced gap caused by data dependence, establish a computational graph dependence relationship model, and based on the output tensor size of the predecessor Kernel and the video memory bandwidth, deduce the minimum necessary transfer time and eliminate the excessive waiting from the total duration; implement runtime gap classification compensation. When the gap type is identified as video memory controller contention, start the hardware performance counter to track the arbitration cycle count of the memory controller, and accurately quantify the delay caused by the conflict;

[0089] Adaptive Sparse Monitoring Strategy: In the initial stage, adopt the full-scale monitoring mode to perform hardware-level cycle counting on each gradient calculation Kernel; when the calculation intensity volatility in 10 consecutive iteration cycles is lower than 8%, switch to the critical path sampling mode, and only continuously track the Kernels that consume the top 20% of the computing resources; establish a Kernel duration prediction model, analyze the historical execution time series based on the LSTM network. When the prediction error is less than 5%, use the predicted value to replace the actual measurement value; inject a monitoring agent at the PyTorch framework level to automatically identify the critical path of gradient calculation (the set of Kernels that contribute more than 85% of the total duration), maintain a sampling interval of 200 microseconds for it, and switch to millisecond-level sparse sampling for non-critical paths.

[0090] Memory Requirements of the Optimizer State during the Parameter Update Phase. When using the Adam optimizer for parameter update during model training, the Adam optimizer will maintain momentum variables. By accessing the momentum variable attributes in the optimizer object, read the memory size occupied by its data structure.

[0091] Step 2: Resource Requirement Prediction. Construct a heterogeneous resource spatio-temporal graph, analyze the resource requirements through spatio-temporal graph attention network and multi-modal feature dynamic fusion, and establish a closed-loop feedback mechanism to compensate for prediction errors;

[0092] Construction of the spatio-temporal graph of heterogeneous resources: The 40GB video memory of the NVIDIA A100 GPU is divided into 256 dynamic units (156MB / unit). Each unit is abstracted as a graph node and assigned features such as the real-time video memory occupancy (accuracy ±0.03GB), historical occupancy volatility (standard deviation over 5 iteration cycles), and the data transfer frequency between adjacent units (obtained through the NVLink hardware counter). At the same time, the CUDA cores are added as independent nodes and assigned features such as the SM cluster utilization rate (accuracy ±1.2%) and the activation status of TensorCores, and two-way edge connections are established. The feature data comes from the resource monitoring step, and through the multi-dimensional characterization of the video memory and computing core status, it provides a rich information basis for subsequent analysis of resource requirements;

[0093] Spatio-temporal graph attention network: For time dimension modeling, dilated causal convolution is used to process historical data, and the dilation factor increases exponentially in the sequence (1, 2, 4). The convolutional kernel covers 15 iteration cycles to capture long-range time dependencies; for space dimension modeling, a multi-head graph attention mechanism is deployed, and three dedicated attention heads are set:

[0094] Head 1 focuses on the Bank conflict pattern between video memory units, and calculates the attention weight α1 = softmax(LeakyReLU(W1·[h i ||h j )); Head 2 analyzes the load balance between the computing core and the video memory unit, and the weight α2 = softmax(LeakyReLU(W2·[h i ⊙h j )); Head 3 models the cross-model data transfer requirements, and the weight where α1, α2, and α3 respectively represent the attention weights calculated by the three dedicated attention heads; softmax is an activation function used to convert the input values into a probability distribution, making all output values between 0 and 1 and the sum equal to 1; α1, α2, and α3 are variants of the rectified linear unit activation function, which solves the problem that neurons are not activated when the input of the ReLU function is negative, and the formula is where γ is a positive number; W1, W2, and W3 are all learnable parameter matrices; h i , h j are the feature vectors of graph nodes i and j; [·||·] is the vector concatenation operation; [·⊙·] is the element-wise multiplication operation of vectors; is the feature fusion operation defined according to specific modeling requirements, which is used to combine the feature vectors of nodes i and j to capture information related to cross-model data transfer requirements;

[0095] Multi-modal Feature Dynamic Fusion: Introduce an external feature encoder to process two types of key information: Model structure features: The number of Transformer layers and the number of attention heads are mapped into a 128-dimensional embedding vector; Task priority signal: A dynamic weight coefficient normalized to the interval [0.1, 1.0]; Design a gated fusion unit to achieve feature fusion: Calculate the gate value: g = σ(W g1 ·[h graph ||h ext +b g ); Feature fusion: h final = g⊙h graph +(1 - g)⊙h ext ; where g is the gate value, a scalar between 0 and 1; σ is the sigmoid function, which maps the input value to the interval from 0 to 1; W g1 is a learnable weight matrix; h graph is the feature vector output by the spatio-temporal graph attention network, which contains the feature information related to the time and space dimensions extracted from the video memory unit and the computing core node; h ext is the feature vector obtained after the external feature encoder processes the model structure features and the task priority signal; b g is the bias vector; h final is the finally fused feature vector; ⊙ is the Hadamard product, that is, the element-wise multiplication operation of vectors; h ext is the feature vector obtained after the external feature encoder processes the model structure features and the task priority signal.

[0096] Online Compensation for Prediction Error: Establish a closed-loop feedback mechanism. When the deviation between the measured video memory occupancy and the predicted value exceeds 5%, dynamically reconstruct the spatio-temporal graph structure, and increase the connection weight between the abnormal unit and the associated core by 300%; Inject Gaussian noise perturbation (σ = 0.15) to enhance the robustness of the model; Perform lightweight distillation every 30 minutes: Compress the prediction model to 40% of the original size using temperature scaling (T = 2) and synchronize it to the edge scheduling node.

[0097] Step 3: Dynamic Resource Reallocation, Deploy the video memory demand prediction model, analyze the parameters affecting the evaluation of migration priority, perform video memory block management and pipeline scheduling to improve the resource allocation efficiency;

[0098] Deploy the video memory demand prediction model. By analyzing the features in the model training stage, the historical resource occupancy pattern, and the task queue metadata, generate a heat map of the video memory demand within the next 1 second 300 milliseconds in advance. When it is predicted that a certain model is about to enter the parameter update stage, actively trigger the pre-marking process of the low-priority video memory block, and use the NVLink DMA engine to pre-load the migration data packet to reduce the migration preparation time; The migration priority is comprehensively evaluated through a dynamic evaluation mechanism. The specific process is as follows: Analyze the influencing parameters, including:

[0099] Training Reward Gradient TD: Calculate the Exponential Moving Average (EMA) of the loss values of the past 10 batches. The formula is: It reflects the recent training effect of the model. The larger the gradient, the higher the task reward. Here, i is the batch index, ranging from 1 to 10, representing the past 10 batches; Δloss i is the change in the loss value of the i-th batch; Energy Consumption Efficiency Factor EZ: Obtain the GPU power consumption GF and computing throughput FLOP in real time through the NVIDIA SMI interface S and calculate the energy consumption efficiency: which is used to quantify the energy efficiency ratio of task execution and avoid tasks with high energy consumption and low rewards from preempting resources; Dependency Strength YQ: Analyze the parameter sharing rate and data exchange frequency between models and define: The higher the value, the stronger the collaborative demand between models, and resource continuity needs to be ensured first;

[0100] Design a comprehensive evaluation function for design priority, integrating the video memory requirement R 显存 and task priority Q 优先级 and influencing parameters: Pri = τ1×R 显存 +τ2×A 优先级 +τ3×TD+τ4×EZ+τ5×YQ. The initial weights are set as τ1 = 0.3, τ2 = 0.25, τ3 = 0.2, τ4 = 0.15, τ5 = 0.1, and online learning adjustment is supported. When the prediction error exceeds the threshold, trigger the backpropagation algorithm to dynamically optimize the weight allocation; when the prediction model A is about to enter the parameter update stage, decide whether to trigger video memory migration based on the priority score: when Pri 模型B > 1.5×Pri 模型A , immediately start the migration; otherwise, delay the migration until model A completes the critical calculation stage; preload the migration data packet through the NVLink DMA engine and dynamically adjust the transmission bandwidth in combination with the priority score to ensure that the data packets of high-priority tasks are transmitted first.

[0101] The video memory block management adopts an elastic adaptive strategy. The basic division is 256 units of 156MB, and it supports dynamic merging into large blocks of 312MB or 624MB; when it is detected that the model requests more than 256MB of video memory, the merging protocol automatically integrates adjacent free units and optimizes the video memory space through a periodic fragmentation reorganization engine; the fragmentation reorganization uses the ant colony algorithm to calculate the minimum movement path, which is executed once every 5 minutes to reduce the video memory fragmentation rate; combine the LSTM network to predict the video memory demand distribution in the next 10 iteration cycles, reserve continuous large block spaces in advance, avoid allocation failures caused by sudden demands, and improve the success rate of large block video memory allocation;

[0102] The pipeline scheduling integrates real-time resource monitoring data to construct a multi-dimensional resource status matrix, covering the utilization rate of SM clusters, the memory bank-level latency, and task dependencies; the dynamic priority scheduling strategy injects low-priority preprocessing tasks when the idle rate of GPU computing units exceeds 25%, and overlaps data transfer and computing tasks when the memory bandwidth utilization rate is lower than 40%; the critical path tasks adopt round-robin scheduling to ensure the minimization of the latency of core operations such as gradient calculation; the fault-tolerant recovery mechanism uses an FPGA coprocessor to verify the calculation results in real time. When the error rate exceeds 0.1%, it automatically rolls back to a stable state and reconstructs the pipeline to optimize and compress the idle rate.

[0103] Step 4: Cross-model hierarchical knowledge distillation. Using BERT-Large as the teacher model, by establishing a knowledge feature alignment layer, a multi-granularity knowledge distillation strategy, and progressive knowledge fusion, it promotes the student model to learn effective features.

[0104] Establish a knowledge feature alignment layer: Select BERT-Large with the largest number of parameters in the model set as the teacher model; perform feature mapping on the 6th, 12th, and 18th layer Transformer modules of other student models; use the Wasserstein distance to measure the similarity of the output distributions of each layer (the threshold is set to 0.35). When the similarity is lower than the set threshold of 0.35, perform subsequent knowledge distillation operations to promote the student model to learn the effective features of the teacher model.

[0105] Multi-granularity knowledge distillation strategy: When propagating forward through each layer, calculate the KL divergence loss between the corresponding layers of the student model and the teacher model (the temperature parameter τ = 3); construct a joint probability distribution at the model output layer and use the Jensen-Shannon distance to constrain the logits difference; introduce an adaptive weight adjuster to dynamically allocate loss weights according to the alignment degree of each layer (range 0.1 - 0.9).

[0106] Progressive knowledge fusion: In the first stage, force the alignment of the underlying features (layers 1 - 6), and set the learning rate to 1.2 times the benchmark value; in the second stage, relax the middle layer constraints (layers 7 - 12) and introduce a random dropout mechanism (dropout rate 15%); in the third stage, only perform soft target supervision at the output layer to retain the differences in high-level features.

[0107] Step 5: Adaptive collaborative training. Construct a task similarity evaluation matrix, dynamically adjust the collaboration mode according to the similarity, and perform elastic scheduling training to improve the model training effect.

[0108] Construct a task similarity evaluation matrix: Extract the feature representations of each model in the validation set (the dimension is reduced to 512); calculate the cosine similarity matrix S ∈ R N×N; R is the set of real numbers, and N is the number of models; when S > 0.7, it is marked as a similar task group; dynamically adjust the collaboration mode: the high-similarity group (S > 0.7) starts the parameter sharing mechanism and freezes 60% of the common parameters at the bottom layer; the medium-similarity group (0.4 < S ≤ 0.7) enables gradient cross-accumulation and synchronizes gradients every 3 batches; the low-similarity group (S ≤ 0.4) implements adversarial training and generates perturbation samples through the gradient reversal layer;

[0109] Elastic training scheduling: when it is detected that the loss reduction rate of a certain model is lower than the mean by 30%, automatically trigger the collaboration mode switch; deploy a priority queue in the NVIDIA DGX cluster and allocate an additional 16% of the computing resources to critical models; achieve cross-node gradient aggregation through the RDMA network, and control the latency within 2 ms.

[0110] The preferred embodiments of the present invention disclosed above are only used to help explain the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the present invention to only the specific embodiments. Obviously, according to the content of this specification, many modifications and changes can be made. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the present invention, so that those skilled in the art in the relevant technical field can well understand and utilize the present invention. The present invention is only limited by the claims and their full scope and equivalents.

Claims

1. A multi-model collaborative operation method for efficient training based on large models, characterized in that, It includes the following steps: S1: Resource monitoring, monitoring resource metrics such as the memory occupancy of activation tensors during forward propagation in model training, the intensity of gradient calculation during backpropagation, and the memory requirements of the optimizer state during the parameter update phase, and adopting a three-level optimization mechanism to optimize the monitoring of gradient calculation intensity; S2: Resource demand prediction, constructing a heterogeneous resource spatio-temporal graph, analyzing resource demands through spatio-temporal graph attention networks and multi-modal feature dynamic fusion, and establishing a closed-loop feedback mechanism to compensate for prediction errors; S3: Dynamic resource reallocation, deploying a video memory demand prediction model, analyzing the factors affecting parameter evaluation and migration priorities, and performing video memory block management and pipeline scheduling to improve resource allocation efficiency; S4: Cross-model hierarchical knowledge distillation, using BERT-Large as the teacher model, and promoting the student model to learn effective features by establishing a knowledge feature alignment layer, a multi-granularity knowledge distillation strategy, and progressive knowledge fusion; S5: Adaptive collaborative training, constructing a task similarity evaluation matrix, dynamically adjusting the collaboration mode according to the similarity, and elastically scheduling the training to improve the model training effect.

2. The multi-model collaborative operation method based on efficient training of a large model according to claim 1, wherein The specific process of S1 is as follows: Monitor the resource metrics during the model training process, including: The memory occupancy of activation tensors in the forward propagation stage. During the forward propagation of the model, use a memory monitoring tool or the memory analysis interface provided by the programming language to obtain the information on the memory occupied by the activation tensors; The intensity of gradient calculation in the backpropagation stage. Use the tools provided by CUDA programming to record the execution time of the Kernel; start timing before the backpropagation starts, and stop timing when the CUDA Kernel involving gradient calculation in the backpropagation is completed, so as to obtain the elapsed time of this Kernel. By measuring the elapsed time of the CUDA Kernel, calculate the intensity of gradient calculation in the backpropagation stage; and optimize it using a three-level optimization mechanism; The memory requirements of the optimizer state in the parameter update phase. When using the Adam optimizer to update parameters during model training, the Adam optimizer will maintain momentum variables. By accessing the momentum variable attributes in the optimizer object, read the memory size occupied by the data structure.

3. A multi-model collaborative operation method based on efficient training of large models according to claim 2, characterized in that The process of the three-level optimization mechanism in S1 is as follows: Hardware-level execution trace tracking: Through the PC Sampling hardware module of the NVIDIA Ampere architecture, directly collect the number of execution cycles of the computing instructions inside the stream processor, and record the actual operation duration of each gradient calculation Kernel with a precision of 10 nanoseconds; for the Volta architecture and GPU, activate the CycleCounter register of each SM cluster to capture the full process clock cycles from instruction emission to result write-back in real time, eliminating the error introduced by CUDA runtime scheduling; establish a Kernel fingerprint library, perform instruction pattern recognition on the backpropagation operators in the cuDNN library, automatically label the gradient calculation core Kernel and isolate auxiliary memory operations; Deploy a gap monitor at the CUDA stream level. When it detects an execution interval of more than 5 microseconds between two adjacent Kernels, automatically trigger the following compensation strategy: For the waiting gaps caused by resource contention, by analyzing the shared memory bank conflict pattern, calculate the difference between the theoretical maximum throughput and the actual throughput, and deduct the invalid waiting duration proportionally; For the forced gaps caused by data dependencies, establish a computational graph dependency relationship model, and based on the output tensor size of the predecessor Kernel and the video memory bandwidth, deduce the minimum transmission time, and eliminate the excessive waiting from the total elapsed time; Implement runtime gap classification and compensation. When the gap type is identified as video memory controller contention, start the hardware performance counter to track the arbitration cycle count of the memory controller and quantify the latency caused by the conflict; Adaptive sparse monitoring strategy: In the initial stage, adopt a full-scale monitoring mode to perform hardware-level cycle counting on each gradient calculation Kernel; When the computational intensity volatility in 10 consecutive iteration cycles is lower than 8%, switch to the critical path sampling mode, that is, only continuously track the Kernels that consume the top 20% of the computing resources; Establish a Kernel elapsed time prediction model, analyze the historical execution time series based on the LSTM network. When the prediction error is less than 5%, use the predicted value to replace the actual measured value; Inject a monitoring agent at the PyTorch framework level to automatically identify the critical path of gradient calculation, that is, the set of Kernels that contribute more than 85% of the total duration, maintain a sampling interval of 200 microseconds, and switch to millisecond-level sparse sampling for non-critical paths.

4. A multi-model collaborative operation method based on efficient training of large models according to claim 1, characterized in that, The specific operation steps of S2 are as follows: Heterogeneous resource spatio-temporal graph construction: Divide the 40GB video memory of NVIDIA A100 GPU into 256 dynamic units, abstract each unit as a graph node and assign it real-time video memory occupancy, historical occupancy fluctuations, and adjacent unit data transfer frequency characteristics. At the same time, add CUDA cores as independent nodes and assign SM cluster utilization rate and TensorCore activation status characteristics, and establish bidirectional edge connections; The feature data comes from step S1. Through multi-dimensional characterization of the video memory and computing core status, it provides an information basis for subsequent analysis of resource requirements; Spatio-temporal graph attention network: Modeling in the time dimension, use dilated causal convolution to process historical data, and the dilation factor increases according to the exponential sequence (1, 2, 4), and the convolutional kernel covers 15 iteration cycles to capture long-range time dependencies; Modeling in the space dimension, deploy a multi-head graph attention mechanism, and set three dedicated attention heads: Head 1 focuses on the Bank conflict pattern between video memory units, and calculates the attention weight α1 = softmax(LeakyReLU(W1·[h i ||h j )); The head 2 analyzes and calculates the load balance between the core and the video memory unit, and the weight α2 = softmax(LeakyReLU(W2·[h i ⊙h j )); Head 3 models the cross-model data transfer requirements, weights where α1, α2, and α3 respectively represent the attention weights calculated by the three dedicated attention heads; Softmax is an activation function used to convert the input numerical values into a probability distribution, making all output values between 0 and 1 and the sum equal to 1. LeakyReLU is a variant of the rectified linear unit activation function, which solves the problem that neurons do not activate when the input is negative. The formula is where γ is a positive number; W1, W2, and W3 are all learnable parameter matrices; h i , h j are the feature vectors of graph nodes i and j; [·||·] is the vector concatenation operation; [·⊙·] is the element-wise multiplication operation of vectors; is a feature fusion operation defined according to specific modeling requirements, which is used to combine the feature vectors of nodes i and j to capture information related to cross-model data transmission requirements; Multi-modal feature dynamic fusion: Introduce an external feature encoder to process two types of information: Model structure features: The number of Transformer layers and the number of attention heads are mapped to 128-dimensional embedding vectors; Task priority signal: a dynamic weight coefficient normalized to the interval [0.1, 1.0]; design a gated fusion unit to achieve feature fusion: calculate the gating value: g = σ(W g1 ·[h graph ||h ext +b g ); Feature Fusion: h final = g ⊙ h graph + (1 - g) ⊙ h ext ; where g is the gating value, a scalar between 0 and 1; σ is the sigmoid function that maps the input value to the interval from 0 to 1; W g1 is the learnable weight matrix; h graph is the feature vector output by the spatio-temporal graph attention network, containing the feature information related to the time and space dimensions extracted from the video memory unit and the computing core node; h ext is the feature vector obtained after the external feature encoder processes the model structure features and the task priority signals; b g is the bias vector; h final is the finally fused feature vector; ⊙ is the Hadamard product, that is, the operation of multiplying the corresponding elements of the vectors; h ext is the feature vector obtained after the external feature encoder processes the model structure features and the task priority signals; Online compensation for prediction error: Establish a closed-loop feedback mechanism. When the deviation between the measured video memory occupancy and the predicted value exceeds 5%, dynamically reconstruct the spatio-temporal graph structure and increase the connection weight between the abnormal unit and the associated core by 300%; Inject Gaussian noise perturbation to enhance the robustness of the model; Perform lightweight distillation every 30 minutes: Compress the prediction model to 40% of the original size using temperature scaling and synchronize it to the edge scheduling node.

5. A multi-model collaborative operation method based on efficient training of large models according to claim 1, characterized in that, The specific operation steps of S3 are as follows: Deploy a video memory demand prediction model. By analyzing the characteristics in the model training stage, the historical resource occupancy pattern, and the task queue metadata, generate a heat map of the video memory demand within the next 1 second 300 milliseconds in advance; When it is predicted that the model is about to enter the parameter update stage, actively trigger the pre-labeling process of the low-priority video memory blocks, and use the NVLink DMA engine to preload the migration data packets to reduce the migration preparation time; The migration priority is comprehensively evaluated through a dynamic evaluation mechanism; The video memory block management adopts an elastic adaptive strategy. The basic division is 256 units of 156MB, and it supports dynamic merging into large blocks of 312MB or 624MB; When it is detected that the model requests more than 256MB of video memory, the merging protocol automatically integrates adjacent idle units, and optimizes the video memory space through a periodic fragmentation reorganization engine; The fragmentation reorganization uses the ant colony algorithm to calculate the minimum movement path, which is executed once every 5 minutes to reduce the video memory fragmentation rate; combine with the LSTM network to predict the video memory demand distribution in the next 10 iteration cycles and reserve continuous large blocks of space in advance; The pipeline scheduling integrates real-time resource monitoring data to construct a multi-dimensional resource status matrix, covering the SM cluster utilization rate, the video memory Bank-level latency, and the task dependencies; The dynamic priority scheduling strategy injects low-priority preprocessing tasks when the idle rate of the GPU computing unit exceeds 25%, and overlaps the data transmission and computing tasks when the video memory bandwidth utilization rate is lower than 40%; The critical path tasks adopt round-robin scheduling to ensure the minimization of the latency of the core operations of gradient calculation; The fault tolerance and recovery mechanism uses the FPGA coprocessor to verify the calculation results in real time. When the error rate exceeds 0.1%, it automatically rolls back to the stable state and reconstructs the pipeline to optimize and compress the idle rate.

6. The multi-model collaborative operation method for efficient training based on a large model according to claim 5, wherein The specific operation steps of the comprehensive evaluation of the migration priority in S3 through the dynamic evaluation mechanism are as follows: Analyze the influencing parameters, including: Training Reward Gradient TD: Calculate the exponentially weighted moving average of the loss values for the past 10 batches. The formula is: where i is the batch index, ranging from 1 to 10, representing the past 10 batches; Δloss i is the change in the loss value for the i-th batch; Energy consumption efficiency factor EZ: Real-time obtain the GPU power consumption GF and computing throughput FLOP through the NVIDIA SM interface S , calculate the energy consumption efficiency: used to quantify the energy efficiency ratio of task execution; Dependency strength YQ: Analyze the parameter sharing rate and data exchange frequency between models, and define: Design a comprehensive evaluation function for priorities, integrating the video memory requirement R 显存 , the task priority Q 优先级 and the influencing parameters: Pri = τ1 × R 显存 + τ2 × A 优先级 + τ3 × TD + τ4 × EZ + τ5 × YQ. The initial weights are set as τ1 = 0.3, τ2 = 0.25, τ3 = 0.2, τ4 = 0.15, τ5 = 0.1, and online learning adjustment is supported; when the prediction error exceeds the threshold, the backpropagation algorithm is triggered to dynamically optimize the weight allocation; When the prediction model A enters the parameter update stage, it decides whether to trigger video memory migration based on the priority score: when Pri 模型B > 1.5 × Pri 模型A , immediately start the migration; Otherwise, delay the migration until model A completes the critical calculation stage; preload the migration data packets through the NVLink DMA engine, and dynamically adjust the transmission bandwidth in combination with the priority score to ensure the priority transmission of the data packets of high-priority tasks.

7. A multi-model collaborative operation method based on efficient training of large models according to claim 1, characterized in that The specific operation steps of S4 are as follows: Establish a knowledge feature alignment layer: Select BERT-Large with the largest number of parameters in the model set as the teacher model; perform feature mapping on the 6th, 12th, and 18th layer Transformer modules of other student models; Use the Wasserstein distance to measure the similarity of the output distributions of each layer, and set the threshold to 0.

35. When the similarity is lower than the set threshold of 0.35, perform the subsequent knowledge distillation operation; Multi-granularity knowledge distillation strategy: When each layer propagates forward, calculate the KL divergence loss between the corresponding layers of the student model and the teacher model, and the temperature parameter τ = 3; construct a joint probability distribution at the model output layer, and use the Jensen-Shannon distance to constrain the logits difference; introduce an adaptive weight adjuster to dynamically allocate the loss weights according to the alignment degree of each layer; Progressive knowledge fusion: In the first stage, force the alignment of low-level features, and set the learning rate to 1.2 times the baseline value; in the second stage, relax the middle-level constraints, introduce a dropout mechanism with a dropout rate of 15%; in the third stage, implement soft target supervision only at the output layer and retain the differences in high-level features.

8. A multi-model collaborative operation method based on efficient training of large models according to claim 1, characterized in that, The specific operation steps of S5 are as follows: Construct a task similarity evaluation matrix: Extract the feature representations of each model on the validation set; Calculate the cosine similarity matrix S ∈ R N×N ; R is the set of real numbers, and N is the number of models; When S > 0.7, it is marked as a similar task group; Dynamically adjust the collaboration mode: For high similarity groups, that is, when S > 0.7, start the parameter sharing mechanism and freeze 60% of the common parameters at the bottom layer; For medium similarity groups, that is, when 0.4 < S ≤ 0.7, enable gradient cross-accumulation and synchronize gradients every 3 batches; For low similarity groups, that is, when S ≤ 0.4, implement adversarial training and generate perturbed samples through a gradient reversal layer; Elastic training scheduling: When it is detected that the loss decline rate of a certain model is lower than the average by 30%, automatically trigger the switching of the collaboration mode; Deploy a priority queue in the NVIDIA DGX cluster to allocate an additional 16% of computing resources to the model; achieve cross-node gradient aggregation through the RDMA network, and control the latency within 2 ms.

Citation Information

Cited By

  • Large model reasoning efficiency dynamic optimization and hardware sensing compression method

    CN120494006A

  • Dynamic optimization of large model inference performance and hardware-aware compression methods

    CN120494006B

  • Large model batch reasoning and data flow optimization system oriented to MOE architecture

    CN120849141A

  • Mass inference and data flow optimization system for MOE architecture-oriented large model

    CN120849141B

  • Method and system for realizing resource use and concurrent reasoning of large-model all-in-one machine

    CN120930806A