Efficient adaptation method of credential and credential environment large language model based on resource dynamic allocation

By evaluating hardware resources and task complexity in real time and dynamically adjusting the ranks of each layer of the large language model, the problems of resource waste and insufficient performance in the information innovation environment are solved, and the dynamic allocation of hardware resources and the improvement of model performance are achieved.

CN120295793AInactive Publication Date: 2025-07-11SHENZHEN DINGSHENG FANGYUAN TECHNOLOGY DEVELOPMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510433676.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-11
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the information innovation environment, domestic hardware computing power is limited, and existing technology is difficult to realize the dynamic allocation of resources of large language models, resulting in waste of resources or insufficient performance, and a single rank adjustment takes a long time to cope with the dynamic changes in hardware resources.

Method used

By evaluating the remaining resources and task complexity of the hardware in real time, dynamically adjusting the rank of each layer of the large language model, combining task urgency and hardware resources, dynamic allocation of resources is achieved, avoiding resource waste and hardware overload, and improving model performance.

Benefits of technology

It realizes dynamic allocation of hardware resources in the information innovation environment, improves the training efficiency and performance of large language models, avoids resource tightness and task delays, and adapts to the dynamic changes of hardware resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295793A_ABST
    Figure CN120295793A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of resource allocation, in particular to an efficient adaptation method of a credential environment large language model based on dynamic resource allocation, which comprises the following steps: acquiring video memory margin and effective computing power of each hardware in real time, and the maximum data transmission rate which can be provided when each hardware performs data transmission with other hardware; the gradient tensor of each layer in the large language model, the total time consumption of each layer of task and the estimated duration from the current moment to the completion; evaluating the residual resources of each piece of hardware in real time, and updating the preset reference rank of each layer according to the loss value of each layer and the change condition of the gradient tensor and the proportion of the real-time video memory margin of each piece of hardware in the total video memory amount to obtain the rank of each layer in the large language model; adjusting the rank of each layer through the estimated duration and the total consumed time in real time; and distributing hardware to each layer of tasks. According to the method, the residual resource evaluation and the rank dynamic adjustment are performed on the intelligent chip hardware, so that the dynamic allocation of the resources is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of resource allocation, and particularly to an efficient adaptation method for large language models in the information technology innovation environment based on dynamic resource allocation. Background Art

[0002] Information technology innovation aims to achieve the autonomy and controllability of information technology through independent research and development and technological innovation. However, the computing power of domestic hardware in the information technology innovation environment is relatively limited, and fine-tuning is required when adapting to large language models.

[0003] The Low-Rank Adaptation (LoRA) method significantly reduces the fine-tuning overhead by freezing the original model parameters and injecting trainable low-rank matrices. The size of the rank directly affects the performance of the large language model. An overly small rank will result in the loss of key features, while an overly large rank will prevent effective reduction of the overhead.

[0004] However, in the process of adjusting the rank value in the existing technology, it relies on global gradient statistics and singular value decomposition. Each rank adjustment takes a long time, cannot cope with the dynamic changes of the hardware resources of intelligent chips, and is difficult to achieve dynamic resource allocation. Summary of the Invention

[0005] In view of the above, it is necessary to provide an efficient adaptation method for large language models in the information technology innovation environment based on dynamic resource allocation. Compared with the traditional adaptation method for large language models in the information technology innovation environment based on dynamic resource allocation, it realizes the dynamic resource allocation of intelligent chip hardware in the information technology innovation environment.

[0006] The efficient adaptation method for large language models in the information technology innovation environment based on dynamic resource allocation of this application adopts the following technical solutions:

[0007] An embodiment of this application provides an efficient adaptation method for large language models in the information technology innovation environment based on dynamic resource allocation, and this method includes the following steps:

[0008] Obtain in real time the video memory remaining amount, effective computing power of each hardware in the cluster for training the large language model, the maximum data transmission rate that can be provided when each hardware transfers data with the remaining hardware, the gradient tensors of each layer in the large language model, and the total time consumption of each layer task and the estimated duration from the current moment to completion;

[0009] During the training process of the large language model, evaluate the remaining resources of each hardware in real time by comparing the real-time effective computing power of each hardware with the theoretical peak computing power, and comparing the real-time maximum data transmission rate of each hardware with the theoretical maximum data transmission rate.

[0010] Update the preset benchmark rank of each layer in the large language model based on the changes in the loss values and gradient tensors of each layer, the proportion of the real-time video memory remaining in each hardware in the total video memory, and the real-time evaluation results of the remaining resources of each hardware, to obtain the rank of each layer in the large language model;

[0011] Obtain the time decay degree of each layer in the large language model through the real-time estimated duration and the total duration, and adjust the rank of each layer;

[0012] Allocate hardware for each layer task based on the real-time evaluation results and the adjustment results of the ranks of each layer.

[0013] In one embodiment, the method for real-time evaluating the remaining resources of each hardware is as follows:

[0014] Record the ratio of the real-time effective computing power of each hardware to the theoretical peak computing power as the first ratio;

[0015] Record the ratio of the real-time maximum data transfer rate of each hardware to the theoretical maximum data transfer rate as the second ratio;

[0016] Take the weighted sum value of the first ratio and the second ratio as the resource utilization degree of each hardware, where the weight values of the first ratio and the second ratio are both preset values;

[0017] The remaining resources of each hardware are evaluated through the resource utilization degree.

[0018] In one embodiment, obtaining the rank of each layer in the large language model includes:

[0019] Obtain the comprehensive characteristic coefficient of all hardware through the proportion and the real-time evaluation results;

[0020] Record the ratio of the L2 norm between the real-time gradient tensor of each layer and the gradient tensor at the first training as the gradient ratio;

[0021] Record the ratio of the loss value of each layer at the first training to the preset maximum allowable loss as the loss ratio;

[0022] Calculate the product of the loss ratio and the gradient ratio, and calculate the cumulative value of the product and 1;

[0023] The rank of each layer is obtained by combining the preset benchmark rank of each layer, the cumulative value, and the comprehensive influence coefficient.

[0024] In one embodiment, the expression of the comprehensive characteristic coefficient is:

[0025] Z represents the comprehensive characteristic coefficient of all hardware; J represents the number of hardware in the cluster for training the large language model; TMemj Represents the real-time video memory remaining of the j-th hardware; TMem j Represents the total video memory of the j-th hardware; α j Represents the resource utilization of the j-th hardware; β j Represents the compression coefficient of the j-th hardware, which is inversely proportional to the real-time video memory remaining of the j-th hardware.

[0026] In one embodiment, the rank of each layer is the product of the preset reference rank of each layer, the accumulated value, and the comprehensive influence coefficient.

[0027] In one embodiment, the process of obtaining the time attenuation degree is as follows:

[0028] Denote the ratio of the real-time estimated duration of each layer task to the total duration as the duration ratio; calculate the difference between 1 and the duration ratio; calculate the product value of the attenuation factor of each layer and the difference; where the attenuation factor is inversely proportional to the real-time estimated duration of each layer;

[0029] The time attenuation degree is negatively correlated with the product value.

[0030] In one embodiment, the calculation method of the time attenuation degree is: use the opposite number of the product value as the exponent of the exponential function with the natural constant as the base, and the time attenuation degree is the calculation result of the exponential function.

[0031] In one embodiment, the method of adjusting the rank of each layer is: use the product of the rank of each layer and the time attenuation degree as the adjusted rank of each layer.

[0032] In one embodiment, the method of allocating hardware to each layer task is:

[0033] Arrange all the hardware in the cluster for training the large language model in ascending order according to the resource utilization, arrange all the layer tasks in descending order according to the rank, and allocate each layer task to the hardware at the same position for execution.

[0034] In one embodiment, when the number of tasks is greater than the number of hardware, the tasks are cyclically allocated to the hardware according to the arrangement order until all tasks are allocated to one hardware.

[0035] This application has at least the following beneficial effects:

[0036] By evaluating the remaining resources of the hardware and analyzing the complexity of tasks, this application can dynamically adjust the rank according to the remaining resources, avoid wasting resources or overloading the hardware in case of resource tension, and improve the ability of the large language model to capture complex patterns and the overall performance of the model when resources are sufficient and the task difficulty is high; further adjust the rank according to the urgency of the task. When the task is more urgent, the value of the rank is reduced more to avoid resources being preferentially allocated to tasks with high difficulty but non-urgency, resulting in delays in urgent tasks; furthermore, dynamically allocate appropriate hardware to each task according to the remaining resources of the hardware and the rank of each task; compared with the prior art, it no longer relies on global gradient statistics and singular value decomposition to adjust the rank, but adjusts the rank according to the actual state of the hardware, the complexity and urgency of the tasks in the Xinchuang environment, solves the problem of long time consumption for single rank adjustment, can cope with the dynamic changes of hardware resources, and realizes the dynamic allocation of resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.

[0038] Figure 1 It is a flowchart of the steps of an efficient adaptation method for a large language model in a Xinchuang environment based on dynamic resource allocation provided by this application;

[0039] Figure 2 It is a schematic diagram of the evaluation process of the remaining resources. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] In the description of the embodiments of the present application, words such as "exemplary", "or", "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "exemplary", "or", "for example" aims to present relevant concepts in a specific way.

[0041] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used in this application are only for the purpose of describing specific embodiments and are not intended to limit this application. It should be understood that unless otherwise stated in this application, " / " means "or".

[0042] In addition, it should be noted that the terms "first" and "second" in this application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0043] The following specifically describes the specific solution of the efficient adaptation method of the large language model in the Xinchuang environment based on resource dynamic allocation in conjunction with the accompanying drawings.

[0044] An efficient adaptation method of the large language model in the Xinchuang environment based on resource dynamic allocation provided by an embodiment of this application. Specifically, the following efficient adaptation method of the large language model in the Xinchuang environment based on resource dynamic allocation is provided. Please refer to Figure 1 , and the method includes the following steps:

[0045] Step 1: Real-time obtain the remaining video memory, effective computing power of each hardware in the cluster for training the large language model, the maximum data transfer rate that can be provided when each hardware transfers data with the remaining hardware, the gradient tensors of each layer in the large language model, and the total time consumption of each layer task and the estimated duration from the current moment to completion.

[0046] For each intelligent chip hardware in the cluster for training the large language model, obtain the video memory allocation status through the hardware interface call, and real-time collect the remaining video memory FMem of the hardware, with the unit of GB; real-time monitor the activity of the computing unit through the computing power unit of the hardware to obtain the effective computing power TFLOPS of the hardware; real-time obtain the maximum data transfer rate that can be provided when the hardware transfers data with the remaining hardware through the network protocol, with the unit of GB / s.

[0047] For the fine-tuned large language model, during the backpropagation process, real-time capture the gradient tensors of each layer in the large language model; through the scheduler interface, extract the estimated completion time of each layer task, and combine the length of the task queue and the single-step calculation time consumption to obtain the total time consumption of each layer task and the estimated duration from the current moment to completion of each layer task. Among them, the total time consumption of each layer task is: when training once, the total duration estimated to complete each layer task.

[0048] Step 2: Comprehensively analyze the remaining resources of each hardware, the complexity and urgency of each layer task, and adjust the rank of each layer in the large language model.

[0049] The Xinchuang environment aims to achieve the autonomy and controllability of information technology, but the computing power of domestic hardware is relatively limited. The large language model is huge in scale and has a large number of parameters. If large-scale updates are made to all parameters during the adaptation process, it will generate a huge computational burden, which is difficult to achieve in the Xinchuang environment with limited computing power, and it will also consume a large amount of time and resources, reducing the training efficiency of the large language model.

[0050] The main idea of the Low-Rank Adaptation (LoRA) method is to reduce the number of parameters through low-rank decomposition. Specifically, for the weight matrix W0 of a large language model, instead of fine-tuning all its parameters during training, the LoRA method introduces two low-rank matrices during fine-tuning, such that the update of the weight matrix is ΔW = AB, where the dimension of A is D×r, the dimension of B is r×d, d is the number of rows of the weight matrix W0, and r is a rank much smaller than d. The fine-tuned weight matrix is W = W0 + ΔW = W0 + AB. During the training of the large language model, the weight matrix W0 is fixed, and only A and B are updated, which can significantly reduce the number of parameters while maintaining the performance of the large language model.

[0051] For different layers in the large language model, the traditional LoRA method fixes the rank of all layers before training, resulting in the inability to adapt to resource fluctuations and task changes, and easily causing resource waste or insufficient performance. If the rank is set too high, it will cause resource waste in the case of tight resources; if the rank is set too low, the large language model cannot fully learn when facing complex tasks, resulting in insufficient performance.

[0052] Existing LoRA methods rely on global gradient statistics and singular value decomposition, with a large amount of computation. When the video memory occupancy suddenly increases, they cannot reduce the rank in time, resulting in the problem that the large language model may run out of memory. Moreover, existing LoRA methods rely on a single metric to adjust the rank, without considering the special characteristics of the hardware in the information and communication technology (ICT) environment, and cannot dynamically allocate resources according to the actual situation of the hardware.

[0053] Step 2.1, during the training of the large language model, by comparing the real-time effective computing power and the theoretical peak computing power of each hardware, and comparing the real-time maximum data transfer rate and the theoretical maximum data transfer rate of each hardware, the remaining resources of each hardware are evaluated in real time.

[0054] To achieve the purpose of dynamic resource allocation, during the training of the large language model, by comparing the real-time effective computing power and the theoretical peak computing power of each hardware, and comparing the real-time maximum data transfer rate and the theoretical maximum data transfer rate of each hardware, the resource utilization of each hardware is calculated to evaluate the remaining resources of each hardware in real time. The expression is:

[0055] In the formula; α j represents the resource utilization of the j-th hardware; TFLOPS j , L_TFLOPS j represent the real-time effective computing power and the theoretical peak computing power of the j-th hardware respectively; MemBW j , L_MemBW jrespectively represent the real-time maximum data transfer rate and the theoretical maximum data transfer rate of the j-th hardware; a j and b j both represent the weight value of the j-th hardware. The value ranges of a j and b j are both (0, 1), and the sum of a j and b j is 1. In this embodiment, the values of a j and b j are 0.7 and 0.3 respectively. On this basis, the values of a j and b j need to be determined according to the characteristics of the hardware itself. For high-computing-power hardware, increase the value of a j , and for hardware with strong data transfer capabilities, increase the value of b j . Denote as the first ratio, and denote as the second ratio. Among them, the calculation of the theoretical peak computing power is a well-known technology, which will not be elaborated in this application. The theoretical maximum data transfer rate is a known factory parameter.

[0056] It should be noted that: when the proportion of the real-time effective computing power and the real-time maximum data transfer rate of the j-th hardware in the theoretical value is smaller, it means that the remaining computing resources of the j-th hardware are more, and the resource utilization degree is smaller; when the resource utilization degree is smaller, it is more suitable to allocate tasks with higher ranks in the large language model to the j-th hardware. The schematic diagram of the evaluation process of the remaining resources is as shown in Figure 2 .

[0057] Step 2.2: Update the preset benchmark rank of each layer through the change situations of the loss values and gradient tensors of each layer in the large language model, as well as the proportion of the real-time video memory remaining of each hardware in the total video memory, and combine the real-time evaluation results of the remaining resources of each hardware to obtain the rank of each layer in the large language model.

[0058] To give full play to the advantages of different hardware and avoid resource waste and performance bottlenecks, resource dynamic allocation is required, that is, according to the task requirements and hardware resource status, the tasks are reasonably allocated to the most suitable hardware.

[0059] Furthermore, update the preset benchmark rank of each layer through the change situations of the loss values and gradient tensors of each layer in the large language model, as well as the proportion of the real-time video memory remaining of each hardware in the total video memory, and combine the real-time evaluation results of the remaining resources of each hardware to obtain the rank of each layer in the large language model. The specific process is as follows:

[0060] Obtain the comprehensive characteristic coefficients of all hardware through the proportion of the real-time video memory remaining of each hardware in the total video memory and combine the real-time evaluation results of the remaining resources of each hardware. The expression is:

[0061] Z represents the comprehensive characteristic coefficient of all hardware; J represents the number of hardware in the cluster for training the large language model; FMem j represents the real-time video memory remaining of the j-th hardware; TMem j represents the total video memory of the j-th hardware; α j represents the resource utilization of the j-th hardware; β j represents the compression coefficient of the j-th hardware, which is used to control the relationship between the video memory remaining and the rank, and is inversely proportional to the real-time video memory remaining of the j-th hardware. The value range of the compression coefficient is (0, 1). When the video memory remaining is smaller, the value of the compression coefficient is larger, so when the video memory remaining decreases, the rank decreases more;

[0062] The ratio of the L2 norm between the real-time gradient tensor of each layer and the gradient tensor at the first training is denoted as the gradient ratio; the ratio of the loss value of each layer at the first training to the preset maximum allowable loss is denoted as the loss ratio; calculate the product of the loss ratio and the gradient ratio, and calculate the cumulative value of the product and 1; adding 1 aims to ensure that when the gradient ratio is close to 0, the rank will not be too small, so as to ensure effective parameter update during the training of the large language model;

[0063] The rank of each layer is the product of the preset reference rank of each layer, the cumulative value, and the comprehensive influence coefficient.

[0064] Among them, the maximum allowable loss refers to the upper limit of the loss allowed during training, and training will be terminated if it is exceeded. The gradient tensor at the first training reflects the parameter update direction and amplitude of the large language model in the initial state. By calculating the ratio of the L2 norm between the real-time gradient tensor of each layer and the gradient tensor at the first training, it is convenient to obtain the rank of each layer based on the preset reference rank of each layer.

[0065] In this embodiment, the maximum allowable loss of each layer needs to be adjusted according to the specific task, model architecture, and dataset.

[0066] In this embodiment, the method for determining the value of the compression coefficient is: taking the reciprocal of the sum of the real-time video memory remaining of each hardware and a preset value greater than 1 as the compression coefficient of each hardware. Among them, adding a preset value greater than 1 aims to avoid the denominator being 0 while ensuring that the calculation result of the reciprocal is less than 1. The preset value greater than 1 is 1.01, and the implementer can set the specific value of the preset value greater than 1 by himself, and this application does not make special restrictions.

[0067] In this embodiment, the benchmark rank of each layer is preset according to the functional characteristics and complexity of each layer of the large language model. The higher the complexity of the function of a layer, the higher the benchmark rank; otherwise, the benchmark rank is lower. The preset benchmark rank of the attention layer is 16, the preset benchmark rank of the Feed Forward Network (FFN) layer is 8, and the preset benchmark rank of the output layer is 12. Implementers can set the values of the benchmark ranks of each layer by themselves, and this application does not impose special restrictions.

[0068] It should be noted that: the calculation result of the rank r i is rounded down, and the constraint condition of the rank is r i ∈[0.5Br i , 2Br i , to avoid hardware overflow, that is, when the calculation result of the rank r i is less than 0.5Br i , r i is assigned 0.5Br i , and when the calculation result of the rank r i is greater than 2Br i , r i is assigned 2Br i ;

[0069] is related to the hardware resources. The remaining video memory directly determines the upper limit of the allocable rank. The exponential compression ratio of the remaining video memory releases resources more quickly in a linear response. When the video memory occupancy suddenly increases, it dominates the rank reduction, effectively reducing the overflow risk; at the same time, it considers the remaining resources of the hardware and increases the rank when the hardware resources are sufficient; for is summed up, then the remaining video memory of all hardware is considered. The larger the sum result, the more available resources, and a higher rank can be used to improve the model's capabilities;

[0070] The accumulated value is related to the difficulty of the task of the i-th layer in the large language model. When the task difficulty increases, for example, when switching from a text task to an image processing task, it is quantified through the change amount of the gradient to increase the rank of the i-th layer model to enhance the expression ability of the large language model; the loss ratio determines the sensitivity of the i-th layer task. The larger its value, the more complex the task, and the higher the rank of the i-th layer should be;

[0071] The larger the rank r i , it indicates that the hardware resources are sufficient, and the i-th layer can allocate more computing resources and video memory resources for parameter training; on the contrary, when the hardware resources are tight, the rank r i will decrease to reduce resource occupancy and avoid resource waste or hardware overload; the larger the rank r i , it indicates that the complexity of the i-th layer task is greater. By increasing the rank of the i-th layer, the effective dimension of the trainable parameters can be increased, thereby enhancing the i-th layer's ability to capture and express complex semantics and features, and improving the overall performance of the model.

[0072] Step 2.3: Obtain the time decay degree of each layer in the large language model based on the real-time estimated duration and the total elapsed time, and adjust the rank of each layer.

[0073] However, the above method for adaptively adjusting the rank is determined based on the hardware resources and the task difficulty of each layer in the large language model. However, different tasks require different amounts of time. When updating the rank, not only the hardware resources and task complexity need to be considered, but also the urgency of the task. In actual application scenarios, the processing times required for different tasks vary greatly. Some tasks may not have high time requirements and can relatively leisurely perform parameter updates and model training; while other tasks have strong timeliness, such as tasks in a real-time question-and-answer system, which need to give answers quickly.

[0074] In the case of limited resources, the allocation of hardware resources needs to comprehensively consider the urgency and complexity of the task. If only the hardware resources and task complexity are concerned, it may lead to emergency tasks not being given priority. For example, an emergency task may be delayed because the resources are allocated to a non-emergency but complex task, which will affect the adaptation degree of the large language model in the domestic information technology innovation environment.

[0075] Based on the above analysis, obtain the time decay degree of each layer through the real-time estimated duration and the total elapsed time of each layer task, and the expression is:

[0076] In the formula, T i represents the time decay degree of the i-th layer; e represents the natural constant; λ i represents the decay factor of the i-th layer, and the value range is (0, 1). The decay factor of the i-th layer is inversely proportional to the real-time estimated duration of the i-th layer. When the real-time estimated duration of the task is shorter, the task is more urgent, the value of the decay factor is larger, and the compression rate of the rank increases exponentially to adapt to the urgency of the task; c i represents the ratio of the real-time estimated duration to the total elapsed time of the i-th layer. When the ratio is larger, the task is less urgent and the compression rate of the rank is lower.

[0077] In this embodiment, the method for determining the value of the decay factor is: take the reciprocal of the sum of the real-time estimated duration of each hardware and a positive number greater than 1 as the decay factor of each hardware. The purpose of adding a positive number greater than 1 is: while avoiding the denominator being 0, ensuring that the calculation result of the reciprocal is less than 1. The value of the positive number greater than 1 is 1.01, and the implementer can set the specific value of the positive number greater than 1 by himself / herself, and this application does not make special restrictions.

[0078] Furthermore, adjust the rank of each layer through the time decay degree of each layer, and the expression is:

[0079] r' i = T i × r i ; where r' i represents the adjusted rank of the i-th layer in the large language model; T i represents the time decay degree of the i-th layer; r i represents the rank of the i-th layer in the large language model. The constraint condition for the rank is r' i ∈ [0.5Br i , 2Br i .

[0080] It should be noted that: the shorter the estimated task duration, the more urgent the task. The exponential function is used to rapidly decrease the rank to ensure the smooth completion of the task. Then, the value of the time decay degree is smaller. On the contrary, while ensuring the timeliness of the task, the large language model maintains good performance;

[0081] By introducing the time decay degree, when the estimated task duration is shorter, the compression rate of the rank increases exponentially, enabling the large language model to quickly adjust the ranks of each layer under limited resources to meet the emergency requirements of different tasks.

[0082] Step 3: Allocate hardware for each layer's task based on the real-time evaluation results of the remaining resources of each hardware and the adjustment results of the ranks of each layer.

[0083] To achieve dynamic resource allocation, allocate according to the resource utilization of the hardware and the adjusted rank of each layer in the large language model. Arrange the hardware in ascending order of resource utilization, arrange each layer's task in descending order of rank, and allocate each layer's task to the hardware at the same position for execution. Specifically: for the arranged tasks and hardware, allocate the first task to the first hardware, the second task to the second hardware, and the N-th task to the N-th hardware. When the number of tasks is greater than the number of hardware, for example, there are N + M tasks and N hardware, and when N ≥ M, allocate the (N + 1)-th task to the first hardware, the (N + 2)-th task to the second hardware, and the (N + M)-th task to the M-th hardware; if there are N + M tasks and N hardware, and N < M, allocate the (2N + 1)-th task to the first hardware, the (2N + 2)-th task to the second hardware until all tasks are allocated to the hardware for execution.

[0084] In summary, by evaluating the remaining resources of the hardware and analyzing the complexity of tasks, the present application can dynamically adjust the rank according to the remaining resources, avoiding wasting resources or overloading the hardware in case of resource shortage, and improving the ability of the large language model to capture complex patterns and the overall performance of the model when resources are sufficient and the task difficulty is high; further adjust the rank according to the urgency of the task, and when the task is more urgent, reduce the value of the rank to avoid resources being preferentially allocated to tasks with high difficulty but non-urgent, resulting in delays in urgent tasks; furthermore, dynamically allocate appropriate hardware to each task according to the remaining resources of the hardware and the rank of each task; compared with the prior art, instead of relying on global gradient statistics and singular value decomposition to adjust the rank, the rank is adjusted through the actual state of the hardware, the complexity and urgency of the tasks in the information technology application innovation environment, solving the problem of long time-consuming for single rank adjustment, being able to cope with the dynamic changes of hardware resources and realizing the dynamic allocation of resources.

[0085] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the block may occur in an order different from that marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in an order different from that disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. Each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0086] For those skilled in the art, it is obvious that the present application is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the basic characteristics of the present application. Therefore, from any point of view, the above embodiments of the present application should be regarded as exemplary and non-limiting.

Claims

1. An efficient adaptation method for large language models in the Xinchuang environment based on dynamic resource allocation, characterized in that, The method includes the following steps: Obtain in real time the remaining video memory of each hardware in the cluster for training the large language model, the effective computing power, the maximum data transfer rate that can be provided when each hardware transfers data with the remaining hardware, the gradient tensors of each layer in the large language model, as well as the total time consumption of each layer task and the estimated duration from the current moment to completion; During the training process of the large language model, evaluate the remaining resources of each hardware in real time by comparing the real-time effective computing power of each hardware with the theoretical peak computing power, and comparing the real-time maximum data transfer rate of each hardware with the theoretical maximum data transfer rate; Update the preset benchmark rank of each layer through the change of the loss value and gradient tensor of each layer in the large language model, the proportion of the real-time remaining video memory of each hardware in the total video memory, and combine with the real-time evaluation result of the remaining resources of each hardware to obtain the rank of each layer in the large language model; Obtain the time decay degree of each layer in the large language model through the real-time estimated duration and the total time consumption, and adjust the rank of each layer; Allocate hardware for each layer task through the real-time evaluation result and the adjustment result of the rank of each layer.

2. The efficient adaptation method of the large language model in the Xinchuang environment based on dynamic resource allocation according to claim 1, wherein, The method for evaluating the remaining resources of each hardware in real time is: Denote the ratio of the real-time effective computing power of each hardware to the theoretical peak computing power as the first ratio; Denote the ratio of the real-time maximum data transfer rate of each hardware to the theoretical maximum data transfer rate as the second ratio; Take the weighted sum value of the first ratio and the second ratio as the resource utilization degree of each hardware, where the weights of the first ratio and the second ratio are both preset values; The remaining resources of each hardware are evaluated through the resource utilization degree.

3. The efficient adaptation method of the large language model in the domestic information technology innovation environment based on dynamic resource allocation as claimed in claim 1, wherein, The obtaining of the rank of each layer in the large language model includes: Obtain the comprehensive characteristic coefficient of all hardware through the proportion and the real-time evaluation result; Denote the ratio of the L2 norm between the real-time gradient tensor of each layer and the gradient tensor at the first training as the gradient ratio; Denote the ratio of the loss value of each layer at the first training to the preset maximum allowable loss as the loss ratio; Calculate the product of the loss ratio and the gradient ratio, and calculate the cumulative value of the product and 1; The rank of each layer is obtained by combining the preset benchmark rank of each layer, the cumulative value and the comprehensive influence coefficient.

4. The efficient adaptation method of the large language model in the domestic information technology innovation environment based on dynamic resource allocation according to claim 3, characterized in that, The expression of the comprehensive characteristic coefficient is: Z represents the comprehensive characteristic coefficient of all hardware; J represents the number of hardware in the cluster for training the large language model; FMem j represents the real-time video memory margin of the j-th hardware; TMem j represents the total video memory of the j-th hardware; α j represents the resource utilization of the j-th hardware; β j represents the compression coefficient of the j-th hardware, which is inversely proportional to the real-time video memory margin of the j-th hardware.

5. The efficient adaptation method of the large language model in the Xinchuang environment based on resource dynamic allocation according to claim 3, characterized in that The rank of each layer is the product of the preset benchmark rank of each layer, the cumulative value and the comprehensive influence coefficient.

6. The efficient adaptation method of the large language model in the Xinchuang environment based on dynamic resource allocation according to claim 1, wherein The process of obtaining the time decay degree is: Denote the ratio of the real-time estimated duration of each layer task to the total time consumption as the duration ratio; calculate the difference between 1 and the duration ratio; Calculate the product value of the decay factor of each layer and the difference; Where the decay factor is inversely proportional to the real-time estimated duration of each layer; The time decay degree is negatively correlated with the product value.

7. The efficient adaptation method of the large language model in the domestic information technology innovation environment based on dynamic resource allocation as claimed in claim 6, wherein The calculation method of the time decay degree is: take the opposite number of the product value as the exponent of the exponential function with the natural constant as the base, and the time decay degree is the calculation result of the exponential function.

8. The efficient adaptation method of the large language model in the Xinchuang environment based on resource dynamic allocation according to claim 1, characterized in that, The method for adjusting the rank of each layer is: take the product of the rank of each layer and the time decay degree as the adjusted rank of each layer.

9. The efficient adaptation method of the large language model in the Xinchuang environment based on dynamic resource allocation as claimed in claim 1, wherein The method for allocating hardware for each layer task is: Arrange all the hardware in the cluster for training large language models in ascending order of resource utilization, arrange all layer tasks in descending order of rank, and assign each layer task to the hardware at the same position for execution.

10. The efficient adaptation method of the large language model in the domestic information technology innovation environment based on dynamic resource allocation as claimed in claim 9, wherein When the number of tasks is greater than the number of hardware, the tasks are cyclically assigned to the hardware in the arranged order until all tasks are assigned to a piece of hardware.