Large language model training and tuning method in credential environment
By combining the computing power and data characteristics of domestic servers in the information creation environment, dynamically adjusting the model weights and optimizing the distributed training of large language models, the problem of low training efficiency of large language models in the information creation environment is solved, and higher training accuracy and speed are achieved.
Patent Information
- Application Number
- CN202510499676.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-18
AI Technical Summary
In the information-creation environment, the training and tuning of large language models need to be carried out on domestic computing power. The existing methods ignore the integration and optimization of model parameters during distributed training, resulting in low training efficiency and poor results.
The distributed training strategy is adopted, combined with the computing power and data characteristics of domestic servers, and the weight of model parameters is dynamically adjusted, and the model fitting deviation is obtained by calculating word vector differences and data feature values, and the aggregation process of model parameters is optimized.
The convergence speed and training accuracy of large language models are improved, the slow convergence problem caused by fixed weights is avoided, and the model's ability to reflect global data features is enhanced.
Smart Images

Figure CN120338020A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of large language model training optimization, and specifically relates to a method for training and optimizing large language models in a domestic information technology innovation environment. Background Art
[0002] The domestic information technology innovation environment is an ecosystem built to achieve the independent control of information technology and ensure information security. Its core is to adopt domestic basic software and hardware products such as hardware, operating systems, and databases. With the increasing external competition, large language models need to be completely independently controllable, secure, and efficient. Therefore, based on the requirements for the security and autonomy of the domestic information technology innovation environment, the training and optimization of large language models must be carried out on domestic computing power and domestic operating systems. Since the training of large language models has extremely high requirements for computing resources, and the domestic computing power in the domestic information technology innovation environment is weak, a distributed training strategy needs to be adopted to train large language models.
[0003] The literature "A Survey of Transmission Optimization Techniques for Federated Large Language Model Training" summarizes the training methods and transmission optimization methods of federated large language models, but points out that the existing methods mainly adjust the transmission optimization in the distributed training process of large language models to reduce communication costs, but ignore the integration and optimization of distributed models in the distributed training process.
[0004] During the training process of federated large language models, the FedAVg algorithm is usually used to update model parameters. The FedAVg algorithm depends on the number of distributed terminals, uses the same weight for the training parameters of each distributed terminal, and updates the global model parameters through weighted averaging, resulting in a low training efficiency of the overall federated learning and a poor model update effect. Summary of the Invention
[0005] In view of the above, it is necessary to provide a method for training and optimizing large language models in a domestic information technology innovation environment, which improves the convergence speed and training accuracy of large language models compared with traditional large language model training and optimization methods.
[0006] A method for training and optimizing large language models in a domestic information technology innovation environment of this application adopts the following technical solution:
[0007] An embodiment of this application provides a method for training and optimizing large language models in a domestic information technology innovation environment. The method includes the following steps:
[0008] Create a distributed training task for a large language model, and randomly initialize the model parameters in the large language model; there is a main domestic server among the domestic servers participating in the training task;
[0009] Randomly select data from the dataset according to the domestic computing power of each domestic server. Train each domestic server based on the selected data. After a single round of training, obtain the gradient information of the model parameters in the large language model of each domestic server.
[0010] Construct a token dictionary based on the occurrence frequency of tokens in the dataset. Based on the token dictionary and the data selected by each domestic server, construct a list of token dictionaries for each domestic server. Calculate the difference in the word vectors of the same tokens in the token dictionary lists between any one domestic server and the other domestic servers to obtain the model fitting deviation of the said one domestic server.
[0011] Obtain the data eigenvalue of each domestic server through the distribution of the data selected by each domestic server in the dataset. Through the model fitting deviation of each domestic server, combined with the domestic computing power and data eigenvalue of each domestic server, obtain the characteristic coefficient of each domestic server. Through the distribution of the characteristic coefficients of all domestic servers, obtain the model weight of each domestic server.
[0012] Update the model parameters in the large language model of the main domestic server through the model weights of all domestic servers, combined with the gradient information in the large language models of all domestic servers.
[0013] In one embodiment, the method of randomly selecting data from the dataset according to the domestic computing power of each domestic server is: the proportion of the data volume selected by each domestic server in the dataset is equal to the proportion of the domestic computing power of each domestic server in the total domestic computing power of all domestic servers.
[0014] In one embodiment, constructing the token dictionary based on the occurrence frequency of tokens in the dataset includes: arranging the tokens in the dataset in descending order of occurrence frequency, and the tokens in the token dictionary are the top pre-set number of tokens after the arrangement.
[0015] In one embodiment, constructing the list of token dictionaries for each domestic server includes: each domestic server constructs word vectors for each token in the selected data based on the large language model, and the corresponding word vectors of all tokens in the token dictionary for each domestic server are respectively composed into the list of token dictionaries for each domestic server.
[0016] In one embodiment, the process of obtaining the model fitting deviation is:
[0017] Denote the domestic servers other than the main domestic server as each slave domestic server. Calculate the mean of all the said differences between any one slave domestic server and the other slave domestic servers, and calculate the average of the said mean between the said one slave domestic server and all the other slave domestic servers.
[0018] Denote the mean value of all the differences between any one of the domestic slave servers and the main domestic server as the difference mean value;
[0019] The model fitting deviation of any one of the domestic slave servers is the weighted sum of the average value and the difference mean value, where the weight value of the difference mean value is greater than the weight value of the average value;
[0020] The model fitting deviation of the main domestic server is the mean value of all the differences between the main domestic server and all the domestic slave servers.
[0021] In one embodiment, the premise that the weight value of the difference mean value is greater than the weight value of the average value is: the domestic computing power of the main domestic server is greater than the domestic computing power of any one of the domestic slave servers.
[0022] In one embodiment, the determination process of obtaining the data characteristic values of each domestic server is as follows:
[0023] When each domestic server selects data from the data set, divide the data in the data set into a preset number of parts, number them in text order, count the numbers of all the data parts selected by any one domestic server, and take the mean value of the deviation values between all any two numbers as the data characteristic value of any one domestic server.
[0024] In one embodiment, the method for obtaining the characteristic coefficient is:
[0025] Denote the proportion of the domestic computing power of each domestic server in the total domestic computing power of all domestic servers as the domestic computing power proportion; calculate the product of the domestic computing power proportion and the data characteristic value;
[0026] The characteristic coefficient is directly proportional to the product and inversely proportional to the model fitting deviation of each domestic server.
[0027] In one embodiment, the method for obtaining the model weight is:
[0028] Calculate the sum value of the characteristic coefficients of all domestic servers, and take the ratio of the characteristic coefficient of each domestic server to the sum value as the model weight of each domestic server.
[0029] In one embodiment, the method for updating the model parameters in the large language model of the main domestic server is:
[0030] Take the weighted sum value of the gradient information in the large language models of all domestic servers as the global gradient information, where the weight value of the gradient information of each domestic server is the model weight of each domestic server;
[0031] The expression for updating the model parameters of the large language model on the main domestic server is as follows:
[0032] W t+1 = W t + ηΔT; where W t+1 and W t respectively represent the weights of any neuron node in the large language model during the (t + 1)-th round and the t-th round of training; η represents the preset learning rate; ΔT represents the global gradient value of any neuron node in the global gradient information.
[0033] This application has at least the following beneficial effects:
[0034] This application mainly adjusts and optimizes the aggregation of model parameters during the training process of the distributed large language model. When allocating training data to each domestic server, it combines the domestic computing power of each domestic server, enabling more reasonable resource allocation; based on the differences in the training situations between each domestic server and the other domestic servers, as well as the context coherence degree of the training data of each domestic server, it obtains the characteristic coefficients of each domestic server and dynamically adjusts the weights of each domestic server during model parameter aggregation, improving the adaptation degree of each domestic server's contribution to the overall large language model. Even if the training effect of some domestic servers is not good or the data distribution is relatively concentrated, it will not have too much negative impact on the global model, enabling the large language model to more accurately reflect the characteristics of the global data during aggregation, which helps to improve the training accuracy and learning ability of the large language model; compared with the traditional FedAVg algorithm, it avoids the problem of slow convergence speed caused by fixed weights and greatly improves the convergence speed of the large language model. Description of the Drawings
[0035] In order to more clearly illustrate the technical solutions and advantages in the embodiments of this application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0036] Figure 1 is the framework structure diagram of the federated large language model;
[0037] Figure 2 is the step flow chart of a large language model training and tuning method provided by this application in a Xinchuang environment;
[0038] Figure 3 is the schematic diagram of the process for obtaining model weights. Detailed Embodiments
[0039] In the description of the embodiments of the present application, words such as "exemplary", "or", "for example", etc. are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary", "or", "for example", etc. is intended to present relevant concepts in a specific manner.
[0040] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used in this application are only for the purpose of describing specific embodiments and are not intended to limit this application. It should be understood that unless otherwise stated in this application, " / " means "or".
[0041] In addition, it should be noted that the terms "first" and "second" in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0042] The following specifically describes the specific solution of a large language model training and tuning method provided by this application in combination with the accompanying drawings.
[0043] A large language model training and tuning method provided by an embodiment of this application. To ensure the security and autonomy of the large language model, domestic computing power is selected for large language model training in this embodiment. The available domestic chip series include the Haiguang series of the Chinese Academy of Sciences, the Kunpeng + Ascend series of chips of Huawei, and the Feiteng series of chips of China Electronics.
[0044] Specifically, the structural diagram of the federated large language model framework is as Figure 1 shown, Figure 1 where Z represents the main domestic server, C represents the slave domestic server, L1 represents the global model download link, and L2 represents the distributed model upload link. Due to the computing power limitation of domestic chips, a distributed model training strategy needs to be adopted, that is, a joint training method based on the combination of the cloud main domestic server and the distributed slave domestic servers is used for joint training. During the distributed training process, the main domestic server can issue training tasks to the slave domestic servers, and the slave domestic servers can perform model training locally, upload the model parameters after each round of iterative training to the main domestic server, perform parameter integration on the main domestic server, and then issue the adjusted model again, gradually completing the training of the large language model.
[0045] During the training process of the federal large language model, the FedAVg algorithm is generally used to update the model parameters. However, in the domestic distributed computing power environment of the information technology application innovation (ITAI) scenario, due to the different processing capabilities of domestic servers, there are differences in the model parameters of the large language model. When the FedAVg algorithm updates the model parameters, it may use the same weight for the model parameters of each domestic server, resulting in the final large language model being difficult to learn global effective features and causing deviations in the final large language model. Therefore, training optimization is required during the training process of the large language model.
[0046] Please refer to Figure 2 , which shows the step flowchart of a method for training and optimizing a large language model in the ITAI scenario provided by an embodiment of the present application. The method includes the following steps:
[0047] Step 1, create a distributed training task for the large language model and randomly initialize the model parameters in the large language model.
[0048] Take the domestic server with the strongest single computing power as the main domestic server, and select 4 available domestic servers as each slave domestic server. In this embodiment, the number of floating-point operations per second in single precision is used as a measure of computing power.
[0049] It should be noted that: 4 is only an embodiment of the present application, and the implementer can set it according to the actual situation. The present application does not make special restrictions.
[0050] Deploy the original large language model on the domestic server and randomly initialize the model parameters in the large language model. The model parameters include the weights and bias terms of each neuron node.
[0051] On the main domestic server, encrypt the model parameters in the initialized large language model through the national cryptographic algorithm and transmit them to each slave domestic server.
[0052] In this embodiment, the selected national cryptographic algorithm is SM1. As other implementation methods, on the basis of being able to encrypt the model parameters, the implementer can use other national cryptographic algorithms, such as SM2, SM3, SM4, etc. The present application does not make special restrictions.
[0053] When the model parameters are transmitted to the slave domestic server, decrypt them through the key and load the initialized model parameters of the main domestic server into the large language models of each slave domestic server. It should be noted that: the large language models of the slave domestic servers are the same as those of the main domestic server.
[0054] Step 2: Randomly select data from the dataset according to the domestic computing power of each domestic server, and train each domestic server according to the selected data. After a single round of training, obtain the gradient information of the model parameters in the large language model of each domestic server.
[0055] During the training of the large language model, it depends on the dataset. Therefore, store the common dataset on the main domestic server and the slave domestic servers.
[0056] In this embodiment, the main domestic server and the slave domestic servers store the common dataset. During each round of iteration, the data trained by different domestic servers is inconsistent, and the data volume is allocated according to the domestic computing power of the domestic servers. For example, if the domestic computing power of one domestic server accounts for 20% of the total domestic computing power of all domestic servers, when the dataset is split into 100 parts, during a single round of iterative training, the one domestic server randomly selects 20 parts from the dataset, and the other domestic servers select from the remaining 80 parts.
[0057] After synchronously initializing the large language models of the main domestic server and the slave domestic servers and dividing the dataset, perform distributed training of the large language model on each domestic server. The domestic server constructs the word vectors of each token in the dataset through the large language model training.
[0058] During the distributed training process, the word granularities extracted by the large language models of different domestic servers are inconsistent, which will seriously affect the subsequent training accuracy of the large language model and the accuracy of answer generation. Among them, the word granularity is an index in the large language model to measure the semantic independence of a single token, that is, in the large language models of different domestic servers, the word vectors constructed based on the same token should be consistent. However, during the distributed training process, there may be a large deviation in the word vectors constructed based on the same token by different domestic servers.
[0059] Based on the above analysis, arrange the tokens in the dataset in descending order of occurrence frequency, and form a token dictionary with the top preset number of tokens after the arrangement. When the large language models of each domestic server finish training based on the selected data, form the token dictionary lists of each domestic server with the corresponding word vectors of all tokens in the token dictionary. It should be noted that: since the data volume allocated to each domestic server is relatively large, and the token dictionary only counts the multiple tokens with the highest occurrence frequencies in the dataset, therefore, the word vectors of each token in the token dictionary can be obtained on each domestic server.
[0060] In this embodiment, the value of the preset number is 2000, and the value of the preset number is preset by humans. The implementer can set it according to the actual situation, and this application does not make special restrictions.
[0061] After the single-round training of the large language model on domestic servers in various countries is completed, the gradient information of the model parameters is obtained through backpropagation. Among them, the gradient information refers to the deviation of the parameters of each neuron node during the backpropagation process, and the parameters include weights and bias terms.
[0062] Step 3: Construct a word segmentation dictionary based on the occurrence frequency of word segments in the dataset, and construct a list of word segmentation dictionaries for domestic servers in various countries according to the word segmentation dictionary and the data selected by domestic servers in various countries; by calculating the differences in word vectors of the same word segments in the word segmentation dictionary lists between any domestic server and the other domestic servers, obtain the model fitting deviation of the said any domestic server.
[0063] In the traditional federated learning process, the FedAvg algorithm is used for gradient aggregation. That is, when the main domestic server receives the gradient information of the local models uploaded by each slave domestic server, weighted averaging is performed based on the number of domestic servers. If the number of domestic servers is M, the weight of the gradient information corresponding to each slave domestic server is 1 / M, and the global gradient information is obtained through weighted averaging. It should be noted that: since the structures of the large language models of all domestic servers are the same, when performing weighted averaging, weighted averaging is performed on the parameters of each neuron node, and finally the global gradient information is obtained, which is also the deviation of the parameters of each neuron node.
[0064] In the Xinchuang environment, the training capabilities and training data of different domestic servers are different, resulting in different deviations of the large language models of different domestic servers. Therefore, when using the FedAvg algorithm for model aggregation, it may lead to an increase in deviation, resulting in error accumulation, making the effect of the finally aggregated large language model poor. Therefore, it is necessary to perform training optimization on the model aggregation in the process of distributed federated large language model training.
[0065] After the single-round training of the large language model on domestic servers in various countries is completed, the gradient information and the word segmentation dictionary list corresponding to each domestic server are uploaded to the main domestic server through the national cryptography algorithm, and specific model aggregation training optimization is carried out within the main domestic server. The specific process is as follows:
[0066] If the difference in word vectors of the same word segments between any domestic server and the other domestic servers is greater, it indicates that the fitting capabilities of the large language models between the said any domestic server and the other domestic servers differ more. Therefore, it is necessary to reduce the contribution degree of the said any domestic server during the aggregation of the large language model.
[0067] Based on the above analysis, calculate the difference value of the word vectors of each same word in the word segmentation dictionary list between any domestic server and the remaining domestic servers, calculate the average value of all the above difference values between any slave domestic server and the remaining slave domestic servers, and calculate the average value of the above average values between any slave domestic server and all the remaining slave domestic servers; Denote the average value of all the above differences between any slave domestic server and the master domestic server as the difference average value; Take the weighted sum of the above average value and the difference average value as the model fitting deviation of any slave domestic server, where the weight value of the difference average value is greater than the weight value of the average value; Take the average value of all the above differences between the master domestic server and all the slave domestic servers as the model fitting deviation of the master domestic server.
[0068] In this embodiment, the difference value between word vectors is: the reciprocal of the sum of the cosine similarity between word vectors and a preset positive number. Here, the preset positive number is used to avoid the denominator being 0, and the value of the preset positive number is preset manually, and the implementer can set it by himself. In this embodiment, the value of the preset positive number is 0.01. As other implementation manners, the implementer can use existing technologies such as Euclidean distance to measure the difference between word vectors, and this application does not make special restrictions.
[0069] It should be noted that: Since the domestic computing power of the master domestic server is the strongest during the actual training process, the data fitting ability of the master domestic server is the strongest. Therefore, when calculating the model fitting deviation of the slave domestic server, a higher weight is given to the word vector difference between the slave domestic server and the master domestic server.
[0070] In this embodiment, the weight values of the difference average value and the average value are 0.6 and 0.4 respectively. On the basis of satisfying that the weight value of the difference average value is greater than the weight value of the average value, the implementer can set the weight values of the difference average value and the average value by himself.
[0071] Step 4, obtain the data characteristic values of each domestic server through the distribution of the data selected by each domestic server in the data set; Obtain the characteristic coefficients of each domestic server by combining the model fitting deviation of each domestic server with the domestic computing power and data characteristic values of each domestic server; Obtain the model weights of each domestic server through the distribution of the characteristic coefficients of all domestic servers.
[0072] The training dataset of large language models is usually text data with certain context relationships. During the actual training process, the distribution of the data selected by each domestic server in the dataset plays a crucial role. Since the dataset is split into multiple parts and randomly distributed to different domestic servers, for each domestic server, if the continuity of the data distributed to the domestic server is stronger, it means that the domestic server is more dependent on some information in the dataset during training, and the deviation of the domestic server from the dataset training will be greater. Therefore, the contribution degree of the domestic server in the aggregation process of the large language model will be smaller.
[0073] Based on the above analysis, when the dataset is divided into a preset number of parts, number them in text order, and count the numbers of all the data parts selected by any one domestic server. Take the average value of the deviation values between all any two numbers as the data characteristic value of the any one domestic server.
[0074] In this embodiment, the value of the preset number is 100. The value of the preset number is preset manually, and the implementer can set it by himself / herself. This application does not make special restrictions.
[0075] In this embodiment, the deviation value between the numbers is the absolute value of the difference. As other implementation methods, on the basis of being able to measure the difference between the numbers, the implementer can adopt other calculation methods, such as the square of the difference, the ratio, etc. This application does not make special restrictions.
[0076] Further, by combining the model fitting deviation of each domestic server, the domestic computing power and the data characteristic value of each domestic server, obtain the characteristic coefficient of each domestic server, specifically:
[0077] Denote the proportion of the domestic computing power of each domestic server in the total domestic computing power of all domestic servers as the domestic computing power proportion; calculate the product of the domestic computing power proportion and the data characteristic value; the characteristic coefficient of each domestic server is directly proportional to the product and inversely proportional to the model fitting deviation of each domestic server.
[0078] In this embodiment, the expression of the characteristic coefficient of each domestic server is: F k represents the characteristic coefficient of the kth domestic server; Q k represents the domestic computing power proportion of the kth domestic server; D k represents the data characteristic value of the kth domestic server; C k represents the model fitting deviation of the kth domestic server; ∈ represents a preset value greater than 0, used to avoid the denominator being 0. The value of ∈ is preset manually, and the implementer can set it by himself / herself. In this embodiment, the value of ∈ is 0.01.
[0079] In another embodiment, the expression for the characteristic coefficient of each domestic server is as follows: F k represents the characteristic coefficient of the k-th domestic server; Q k represents the proportion of domestic computing power of the k-th domestic server; Q k represents the data characteristic value of the k-th domestic server; C k represents the model fitting deviation of the k-th domestic server; exp() represents the exponential function with the natural constant as the base, used to map C k to a positive number to avoid a denominator of 0.
[0080] It should be noted that: the stronger the domestic computing power, the more dispersed the selected data, and the smaller the word vector difference between the domestic server and other domestic servers, the stronger the fitting ability of the domestic server to the data, and the greater the weight proportion in the aggregation of the large language model.
[0081] Calculate the sum value of the characteristic coefficients of all domestic servers, and use the ratio of the characteristic coefficient of each domestic server to the sum value as the model weight of each domestic server. The schematic diagram of the acquisition process of the model weight is as Figure 3 shown.
[0082] Step 5, update the model parameters in the large language model of the main domestic server by combining the model weights of all domestic servers with the gradient information in the large language models of all domestic servers.
[0083] Use the weighted sum value of the gradient information in the large language models of all domestic servers as the global gradient information, where the weight value of the gradient information of each domestic server is the model weight of each domestic server.
[0084] The expression for updating the model parameters of the large language model of the main domestic server is:
[0085] W t+1 = W t + ηΔT; in the formula, W t+1 , W t respectively represent the weights of any neuron node in the large language model during the (t + 1)-th and t-th rounds of training; η represents the preset learning rate; ΔT represents the global gradient value of any neuron node in the global gradient information.
[0086] In this embodiment, the value of the preset learning rate is 0.01. The value of the preset learning rate is preset manually, and the implementer can set it by himself / herself. This application does not make special restrictions.
[0087] Update the large language model of the main domestic server using the aggregated global gradient information. Here, the optimization algorithm adopted in this embodiment is the gradient descent method. Implementers can use other optimization algorithms, such as the Adam optimization algorithm, AdaGrad optimization algorithm, etc. This application does not make special restrictions.
[0088] After updating the large language model of the main domestic server, distribute the model parameters of the large language model to each domestic server for a new round of training. After continuous optimization and iteration, obtain the final large language model.
[0089] In summary, this application mainly adjusts and optimizes the aggregation of model parameters during the training process of the distributed large language model. When allocating training data to each domestic server, it combines the domestic computing power of each domestic server, enabling more reasonable resource allocation; based on the differences in the training situations between each domestic server and the other domestic servers, as well as the context coherence degree of the training data of each domestic server, obtain the characteristic coefficients of each domestic server, and dynamically adjust the weights of each domestic server during model parameter aggregation, improving the adaptability of each domestic server's contribution to the overall large language model. Even if the training effect of some domestic servers is not good or the data distribution is relatively concentrated, it will not have too much negative impact on the global model, enabling the large language model to more accurately reflect the characteristics of the global data during aggregation, which helps to improve the training accuracy and learning ability of the large language model; compared with the traditional FedAVg algorithm, it avoids the problem of slow convergence speed caused by fixed weights and greatly improves the convergence speed of the large language model.
[0090] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion thereof that contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functionality involved. In the description corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than disclosed in the description. Sometimes, there is no specific order between different operations or steps. For example, two consecutive operations or steps may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functionality involved. Each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0091] For those skilled in the art, it is obvious that this application is not limited to the details of the above exemplary embodiments, and without departing from the basic characteristics of this application, this application can be implemented in other specific forms. Therefore, from any point of view, the above embodiments of this application should be regarded as exemplary and non-restrictive.
Claims
1. A method for training and optimizing large language models in the Xinchuang environment, characterized in that, The method includes the following steps: Create a large language model distributed training task and randomly initialize the model parameters in the large language model; there is a main domestic server among the domestic servers participating in the training task; Randomly select data from the dataset according to the domestic computing power of each domestic server, and train each domestic server according to the data selected by each domestic server. After a single round of training, obtain the gradient information of the model parameters in the large language model of each domestic server; Construct a token dictionary according to the occurrence frequency of tokens in the dataset, and construct a token dictionary list for each domestic server according to the token dictionary and the data selected by each domestic server; obtain the model fitting deviation of any domestic server by calculating the difference in the word vectors of the same tokens in the token dictionary lists between any domestic server and the other domestic servers; Obtain the data characteristic values of each domestic server through the distribution of the data selected by each domestic server in the dataset; obtain the characteristic coefficients of each domestic server by combining the model fitting deviation of each domestic server with the domestic computing power and data characteristic values of each domestic server; obtain the model weights of each domestic server through the distribution of the characteristic coefficients of all domestic servers; Update the model parameters in the large language model of the main domestic server by combining the model weights of all domestic servers with the gradient information in the large language model of all domestic servers.
2. The large language model training and tuning method in the Xinchuang environment according to claim 1, wherein, The method of randomly selecting data from the dataset according to the domestic computing power of each domestic server is: the proportion of the data volume selected by each domestic server in the dataset is equal to the proportion of the domestic computing power of each domestic server in the total domestic computing power of all domestic servers.
3. The large language model training and tuning method in the Xinchuang environment according to claim 1, characterized in that, The constructing of the token dictionary according to the occurrence frequency of tokens in the dataset includes: arranging the tokens in the dataset in descending order of occurrence frequency, and the tokens in the token dictionary are the top preset number of tokens after the arrangement.
4. The large language model training and tuning method in the domestic information technology innovation environment according to claim 1, characterized in that, The constructing of the token dictionary list for each domestic server includes: each domestic server constructs word vectors for each token in the selected data based on the large language model, and the corresponding word vectors of all tokens in the token dictionary for each domestic server respectively form the token dictionary list of each domestic server.
5. A large language model training and tuning method in the Xinchuang environment according to claim 1, characterized in that, The process of obtaining the model fitting deviation is as follows: Denote the domestic servers other than the main domestic server as each slave domestic server, calculate the mean of all the above differences between any slave domestic server and the other slave domestic servers, and calculate the average of the above means between any slave domestic server and all the other slave domestic servers; Denote the mean of all the above differences between any slave domestic server and the main domestic server as the difference mean; The model fitting deviation of any slave domestic server is the weighted sum of the above average and the difference mean, where the weight value of the difference mean is greater than the weight value of the average; The model fitting deviation of the main domestic server is the mean of all the above differences between the main domestic server and all the slave domestic servers.
6. The large language model training and tuning method in the domestic information technology innovation environment according to claim 5, characterized in that, The premise that the weight value of the difference mean is greater than the weight value of the average is: the domestic computing power of the main domestic server is greater than the domestic computing power of any one slave domestic server.
7. A method for training and optimizing large language models in a domestic information technology innovation environment according to claim 1, characterized in that, The determination process of obtaining the data characteristic values of domestic servers in various countries is as follows: When domestic servers in various countries select data from the dataset, the data in the dataset is divided into a preset number of portions and numbered in text order. The numbers of all data portions selected by any one domestic server are counted, and the average value of the deviation values between any two numbers is used as the data characteristic value of the any one domestic server.
8. A method for training and optimizing large language models in a domestic information technology innovation environment according to claim 1, characterized in that, The method for obtaining the characteristic coefficient is as follows: The proportion of the domestic computing power of domestic servers in various countries in the total domestic computing power of all domestic servers is denoted as the domestic computing power proportion; calculate the product of the domestic computing power proportion and the data characteristic value; The characteristic coefficient is directly proportional to the product and inversely proportional to the model fitting deviation of domestic servers in various countries.
9. A method for training and optimizing large language models in a domestic information technology innovation environment according to claim 1, characterized in that, The method for obtaining the model weight is as follows: Calculate the sum value of the characteristic coefficients of all domestic servers, and use the ratio of the characteristic coefficient of each domestic server to the sum value as the model weight of each domestic server.
10. A method for training and optimizing large language models in a domestic information technology innovation environment according to claim 1, characterized in that, The method for updating the model parameters in the large language model of the main domestic server is as follows: Use the weighted sum value of the gradient information in the large language models of all domestic servers as the global gradient information, where the weight value of the gradient information of each domestic server is the model weight of each domestic server; The expression for updating the model parameters of the large language model of the main domestic server is: W t+1 = W t + ηΔT; where, W t+1 , W t respectively represent the weights of any neuron node in the large language model during the (t + 1)-th round and the t-th round of training; η represents a preset learning rate; and ΔT represents the global gradient value of any neuron node in the global gradient information.