Large language model training method and device, storage medium and computer equipment

Through federated learning and quantum encryption technology, large language models are trained on local terminals, and the model parameters are distilled by ring knowledge, which solves the problem of low model accuracy in private data scenarios, and improves model accuracy and generalization capabilities while protecting data security.

CN120353885APending Publication Date: 2025-07-22TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410078184.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-18
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In scenarios with rich privacy data, the large language model trained based on local data is not very accurate, and there are data breaches and security risks.

Method used

Using federated learning technology and quantum encryption method, quantum encrypted training sample data is obtained from local terminals, and model parameters are passed between multiple local terminals through ring knowledge distillation, and training is combined with knowledge distillation method to improve model accuracy.

Benefits of technology

On the premise of protecting data privacy, expand the number of training corpus, improve the accuracy and generalization capabilities of large language models, and avoid the security risks brought by data sharing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353885A_ABST
    Figure CN120353885A_ABST
Patent Text Reader

Abstract

The invention provides a big language model training method and device, a storage medium and computer equipment. The method is applied to a first local terminal in a federated learning system, the federated learning system comprises a federated central server and a plurality of local terminals, and the method comprises the steps of obtaining first training sample data subjected to quantum encryption; receiving a first large language model parameter sent by a second local terminal; determining a teacher model according to the first large language model parameter and a local model; predicting the first training sample data based on a teacher model to obtain prediction probability data output by the teacher model; training a local model according to the first training sample data and the prediction probability data to obtain a student model; sending a second large language model parameter of the student model to a third local terminal; and sending the second large language model parameter of the student model to the federal central server. According to the method, the accuracy of the big language model obtained through training can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and particularly to a method, apparatus, storage medium, and computer device for training a large language model. Background Art

[0002] A large language model (LLM) refers to a deep learning model trained using a large amount of text data, which can understand the meaning of language text or generate natural language text. The large language model can handle various natural language tasks, such as text classification, question answering, and dialogue, etc., and is an important approach to artificial intelligence.

[0003] However, in scenarios where some text data contains a large amount of private information, there is a problem that the accuracy of the trained large language model is not high due to insufficient training sample quantity. Summary of the Invention

[0004] Embodiments of the present disclosure provide a method, apparatus, storage medium, and computer device for training a large language model, and this method can improve the accuracy of the trained large language model.

[0005] According to one aspect of the present disclosure, there is provided a method for training a large language model, which is applied to a first local terminal in a federated learning system. The federated learning system includes a federated central server and multiple local terminals, and the method includes:

[0006] Obtaining first training sample data encrypted by quantum based on a preset quantum key;

[0007] Receiving first large language model parameters sent by a second local terminal in the federated learning system;

[0008] Determining a teacher model according to the first large language model parameters and the local model;

[0009] Predicting the first training sample data based on the teacher model to obtain prediction probability data output by the teacher model;

[0010] Training the local model according to the first training sample data and the prediction probability data to obtain a student model;

[0011] Sending second large language model parameters of the student model to a third local terminal so that the third local terminal updates the parameters of the local model based on the second large language model parameters; and sending second large language model parameters of the student model to the federated central server so that the federated central server updates the parameters of the federated large language model based on the large language model parameters returned by each local terminal

[0012] According to one aspect of the present disclosure, there is provided a training device for a large language model, which is applied to a first local terminal in a federated learning system. The federated learning system includes a federated central server and multiple local terminals. The device includes:

[0013] An acquisition unit, configured to acquire quantum-encrypted first training sample data based on a preset quantum key;

[0014] A receiving unit, configured to receive first large language model parameters sent by a second local terminal in the federated learning system;

[0015] A determination unit, configured to determine a teacher model according to the first large language model parameters and a local model;

[0016] A prediction unit, configured to predict the first training sample data based on the teacher model to obtain prediction probability data output by the teacher model;

[0017] A training unit, configured to train the local model according to the first training sample data and the prediction probability data to obtain a student model;

[0018] A sending unit, configured to send second large language model parameters of the student model to a third local terminal so that the third local terminal updates the parameters of the local model based on the second large language model parameters; and send the second large language model parameters of the student model to the federated central server so that the federated central server updates the parameters of the federated large language model based on the large language model parameters returned by each local terminal.

[0019] Optionally, in some embodiments, the training device for a large language model provided by the present disclosure further includes:

[0020] A first acquisition subunit, configured to acquire a hierarchical label and a category label of each local terminal in the federated learning system;

[0021] A second acquisition subunit, configured to acquire a transmission record of large language model parameters among the multiple local terminals;

[0022] A first determination subunit, configured to determine a third local terminal according to the transmission record, the hierarchical label, and the category label;

[0023] Optionally, in some embodiments, the determination subunit includes:

[0024] An extraction module, configured to extract historical hierarchical labels and historical category labels of local terminals that have undergone transmission of large language model parameters in the transmission record;

[0025] The first determination module is configured to determine the target hierarchical label and the target category label corresponding to the local terminal to be transmitted based on the preset expected hierarchical label, expected category label, the historical hierarchical label, and the historical category label;

[0026] The second determination module is configured to determine the third local terminal according to the matching result between the target hierarchical label and the target category label and the hierarchical label and the category label.

[0027] Optionally, in some embodiments, the training device for the large language model provided by the present disclosure further includes:

[0028] The third acquisition subunit is configured to acquire the quantum entropy and the ecological coefficient of the first training sample data in response to a target operation on the first training sample data, where the target operation includes at least one of an access operation and a modification operation;

[0029] The first update subunit is configured to update the preset quantum key of the first training sample data according to the quantum entropy and the ecological coefficient.

[0030] Optionally, in some embodiments, the first training sample data includes a plurality of sub-sample data, and each sub-sample data is associated with a quantum state. The third acquisition subunit includes:

[0031] The first calculation module is configured to calculate the probability value of observing the quantum state based on the quantum state;

[0032] The second calculation module is configured to calculate the quantum entropy of the first training sample data according to the probability value corresponding to each sub-sample data;

[0033] The third determination module is configured to determine the interaction coefficient between any two sub-sample data;

[0034] The third calculation module is configured to calculate the ecological coefficient of the first training sample data according to the interaction coefficient.

[0035] Optionally, in some embodiments, the update subunit includes:

[0036] The fourth calculation module is configured to calculate the ratio between the quantum entropy and the ecological coefficient;

[0037] The fifth calculation module is configured to calculate a correction value for correcting the preset quantum key according to the ratio;

[0038] The update module is configured to update the preset quantum key based on the correction value.

[0039] Optionally, in some embodiments, the training unit includes:

[0040] A prediction subunit, configured to predict input data in the first training sample data based on the local model to obtain output data;

[0041] A second determination subunit, configured to determine a first loss according to the difference between the output data and the label data in the first training sample data;

[0042] A first calculation subunit, configured to calculate a second loss according to the difference between the output data and the corresponding prediction probability data;

[0043] An adjustment subunit, configured to adjust parameters of the local model according to the first loss and the second loss to obtain a student model.

[0044] Optionally, in some embodiments, the first calculation subunit includes:

[0045] An acquisition module, configured to acquire a temperature parameter corresponding to knowledge distillation;

[0046] A sixth calculation module, configured to calculate a cross entropy between the output data and the corresponding prediction probability data based on the temperature parameter;

[0047] A fourth determination module, configured to determine a second loss according to the cross entropy.

[0048] Optionally, the training device for the large language model provided by the present disclosure further includes a federated large language model update unit, and the federated large language model update unit includes:

[0049] A fourth acquisition subunit, configured to acquire weight parameters and sample quantities corresponding to multiple target local terminals that return large language model parameters during the current round of federated large language model parameter update;

[0050] A weighting subunit, configured to weight the large language model parameters returned by the multiple target local terminals based on the weight parameters and the sample quantities to obtain global model parameters;

[0051] A second update subunit, configured to update the parameters of the federated large language model to the global model parameters.

[0052] Optionally, in some embodiments, the acquisition unit includes:

[0053] A fifth acquisition subunit, configured to acquire quantum-encrypted document data based on a preset quantum key;

[0054] A first processing subunit, configured to perform document cleaning, word segmentation, and sequence padding on the document data to obtain a target text;

[0055] A generation subunit for generating first training sample data according to the target text.

[0056] Optionally, in some embodiments, the generation subunit includes:

[0057] A first generation module for processing the target text based on the generator in a preset generative adversarial network to obtain a generated text;

[0058] A second generation module for generating first training sample data based on the generated text and the target text;

[0059] Wherein, the training process of the generative adversarial network includes the following steps:

[0060] Obtain second training sample data, where the second training sample data includes data samples and label data of the data samples;

[0061] Update the Lagrange multiplier of a preset Lagrangian function based on the data samples and the label data;

[0062] Calculate the generation loss corresponding to the generator according to the Lagrange multiplier and random noise, and calculate the discrimination loss corresponding to the discriminator according to the data generated by the generator and the data samples;

[0063] Update the parameters of the generator and discriminator in the generative adversarial network based on the generation loss and the discrimination loss.

[0064] Optionally, in some embodiments, the structure of the local model includes a feature extraction network and a classification network. The training device for the large language model provided by the present disclosure further includes a feature extraction network construction unit, including:

[0065] A sixth acquisition subunit for acquiring multiple candidate model structures;

[0066] A second calculation subunit for calculating the gravitational fluctuation score corresponding to each candidate model structure;

[0067] A second processing subunit for performing regularization processing on each candidate model structure based on a preset regularization method to obtain a regularization factor;

[0068] A third calculation subunit for calculating the comprehensive evaluation score of each candidate model structure according to the gravitational fluctuation score and the regularization factor;

[0069] A third determination subunit for determining the structure of the large language model from the multiple candidate model structures according to the comprehensive evaluation score.

[0070] Optionally, in some embodiments, the training apparatus for the large language model provided by the present disclosure further includes a classification layer training unit, including:

[0071] An initialization subunit, configured to initialize the input weights and biases of the classification network;

[0072] A first encoding subunit, configured to encode the input weights based on a preset quantum encoding function to obtain weight initial values;

[0073] A second encoding subunit, configured to perform sparse encoding on the hidden layer data determined according to the input features, the weight initial values, and the biases to obtain a sparse encoding representation;

[0074] A third update subunit, configured to update the input weights and the biases based on the difference between the sparse encoding representation and the labels corresponding to the input features.

[0075] According to one aspect of the present disclosure, there is provided a storage medium storing a computer program, and when the computer program is executed by a processor, the training method of the large language model as described above is implemented.

[0076] According to one aspect of the present disclosure, there is provided a computer program product including a computer program, and the computer program is read and executed by a processor of a computer device, so that the computer device executes the training method of the large language model as described above.

[0077] The training method of the large language model provided by the embodiments of the present disclosure is applied to a first local terminal in a federated learning system. The federated learning system includes a federated central server and multiple local terminals. The method obtains first training sample data encrypted by quantum based on a preset quantum key; receives first large language model parameters sent by a second local terminal in the federated learning system; determines a teacher model according to the first large language model parameters and the local model; makes a prediction on the first training sample data based on the teacher model to obtain prediction probability data output by the teacher model; trains the local model according to the first training sample data and the prediction probability data to obtain a student model; sends second large language model parameters of the student model to a third local terminal so that the third local terminal updates the parameters of the local model based on the second large language model parameters; and sends second large language model parameters of the student model to the federated central server so that the federated central server updates the parameters of the federated large language model based on the large language model parameters returned by each local terminal.

[0078] In the embodiments of the present disclosure, by adopting the federated learning technology, the data of multiple local terminals with data barriers due to data security considerations are comprehensively applied to the training of the large language model without sharing the data in any local terminal to other terminals. In this way, the number of training corpora of the large language model can be greatly expanded, thereby improving the accuracy of the large language model. In addition, the embodiments of the present disclosure also adopt the method of knowledge distillation to continuously accumulate the knowledge learned from different local terminal data, thereby further improving the accuracy of the trained large language model.

[0079] Other features and advantages of the present disclosure will be described in the following specification, and partly will be obvious from the specification, or will be understood by implementing the present disclosure. The objectives and other advantages of the present disclosure can be realized and obtained by the structures specifically pointed out in the specification, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0080] The drawings are used to provide a further understanding of the technical solutions of the present disclosure, and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solutions of the present disclosure, and do not constitute a limitation to the technical solutions of the present disclosure.

[0081] Figure 1 It is a system architecture diagram applied to the training method of the large language model in the embodiments of the present disclosure;

[0082] Figure 2 It is a schematic flowchart of a training method of the large language model provided by the present disclosure;

[0083] Figure 3 It is a schematic diagram of the architecture of the federated learning system provided in the embodiments of the present disclosure;

[0084] Figure 4 It is a schematic diagram of multiple meta-federations in the federated learning system of the present disclosure;

[0085] Figure 5 It is a schematic diagram of local model training using the circular knowledge distillation method;

[0086] Figure 6 It is another schematic flowchart of the training method of the large language model provided by the present disclosure;

[0087] Figure 7 It is a schematic diagram of the structure of the training device of the large language model provided in the embodiments of the present disclosure;

[0088] Figure 8 It is a terminal structure diagram for implementing the methods according to an embodiment of the present disclosure;

[0089] Figure 9It is a server structure diagram for implementing the methods according to an embodiment of the present disclosure. Detailed implementation

[0090] In order to make the objectives, technical solutions and advantages of the present disclosure clearer and more understandable, the present disclosure will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure and are not used to limit the present disclosure.

[0091] Before further elaborating on the embodiments of the present disclosure, the nouns and terms involved in the embodiments of the present disclosure are described. The nouns and terms involved in the embodiments of the present disclosure are applicable to the following explanations:

[0092] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning and decision-making. Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the basic model, can be widely applied to the downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0093] Federated learning (FL): Also known as federated machine learning (FML), joint learning or coalition learning. Federated learning is a machine learning framework that can effectively help multiple institutions to use data and perform machine learning modeling while meeting the requirements of user privacy protection, data security and relevant regulations.

[0094] Knowledge Distillation (KD): A model compression method and a training method based on the "teacher-student network" concept. Here, the teacher network is the knowledge provider, and the student network is the knowledge recipient. The knowledge distillation process is divided into two stages: the training stage of the teacher model, which is trained based on the hard labels of the first training sample data; and the training stage of the student model, which is trained based on the hard labels of the training sample data and the soft labels output by the teacher model, enabling the student model to learn the knowledge acquired by the teacher model.

[0095] Quantum Encryption (QE): Quantum encryption technology is a series of encryption techniques that utilize quantum principles for key generation, plaintext confusion encryption, ciphertext restoration decryption, and ciphertext communication.

[0096] Generative Adversarial Network (GAN): A generative model that learns through the mutual game of two neural networks. GAN can perform learning for generation tasks without using labeled data. It consists of a generator and a discriminator. The generator randomly samples from the latent space as input, and its output should mimic the real samples in the training set as much as possible. The input of the discriminator is either real samples or the output of the generator, and its purpose is to distinguish the output of the generator from the real samples as much as possible. The generator and the discriminator oppose each other and continuously learn, with the ultimate goal of making the discriminator unable to determine whether the output of the generator is real. That is, the generator can generate realistic data. This solution is often used in the expansion of sample data for neural network models.

[0097] In the related art, when training a large language model, the training sample data is generally centralized in a central server, and then the training sample data is used in the central server to complete the training of the large language model. However, in some special fields, such as the financial field or the medical field, there may be a large amount of private data in the data set. In this scenario, if the data is centralized in the central server for processing, it may not only lead to data leakage, but also cause other data security problems. To ensure the security of private data, the large language model is generally trained locally. However, the number of local training samples is limited, and the large language model needs to be trained with a large amount of sample data to obtain good accuracy. Therefore, the accuracy of the large language model trained based on local data is currently not high. To solve the problem that the large language model can only be trained based on local data, resulting in low accuracy of the large language model, the present disclosure provides a training method for a large language model, which can improve the accuracy of the large language model.

[0098] System architecture and scenario description applied in the embodiments of the present disclosure

[0099] Figure 1 It is a system architecture diagram to which the training method of the large language model according to the embodiments of the present disclosure is applied. It includes a terminal 140, the Internet 130, a gateway 120, a server 110, etc.

[0100] The terminal 140 includes various forms such as a desktop computer, a laptop computer, a PDA (Personal Digital Assistant), a mobile phone, a vehicle-mounted terminal, a home theater terminal, a dedicated terminal, a smart voice interaction device, a smart home appliance, an aircraft, etc. In addition, it can be a single device or a collection of multiple devices. The terminal 140 can communicate with the Internet 130 in a wired or wireless manner to exchange data. In this embodiment, the terminal 140 can specifically be a local terminal in a federated learning system.

[0101] The server 110 refers to a computer system that can provide certain services to the terminal 140. Compared with an ordinary terminal 140, the server 110 has higher requirements in terms of stability, security, performance, etc. The server 110 can be a high-performance computer in a network platform, a cluster of multiple high-performance computers, a part (such as a virtual machine) allocated from a high-performance computer, a combination of parts (such as virtual machines) allocated from multiple high-performance computers, etc. In this embodiment, the server 110 can specifically be a federated central server in a federated learning system.

[0102] The gateway 120 is also known as an internetwork connector or protocol converter. The gateway realizes network interconnection at the transport layer and is a computer system or device that acts as a converter. Between two systems using different communication protocols, data formats, or languages, and even with completely different architectures, the gateway is a translator. At the same time, the gateway can also provide filtering and security functions. The messages sent by the terminal 140 to the server 110 need to be sent to the corresponding server 110 through the gateway 120. The messages sent by the server 110 to the terminal 140 also need to be sent to the corresponding terminal 140 through the gateway 120.

[0103] The training method of the large language model provided by the embodiments of the present disclosure can be partially implemented in the terminal 140 and partially implemented in the server 110.

[0104] Specifically, in the terminal 140 (local terminal) and the server 110 (federated central server) in the federated learning system, large language models to be trained with the same model structure can be deployed. Then, the terminal 140 can obtain the first training sample data encrypted by quantum locally, and then train the local model (i.e., the locally large language model to be trained) based on the first training sample data to obtain the trained local model. Then, the terminal 140 can send the model parameters of the trained local model to the server 110 and the next terminal 140. After receiving the model parameters sent by the previous terminal 140, the next terminal 140 can determine the teacher model based on the model parameters sent by the previous terminal 140, and perform knowledge distillation learning based on the local training sample data and the teacher model to train the local model, so as to obtain the trained local model (student model). In this way, the next terminal 140 can send the model parameters of the student model to the server 110 and another next terminal 140, and so on in a cycle until the model parameters are passed through all the terminals 140, completing one cycle of training of the federated large language model. At this time, the server 110 receives the model parameters sent by each terminal 140, and the server 110 can update the parameters of the federated large language model deployed therein accordingly. And further distribute the model parameters of the updated federated large language model to each terminal 140 for the next round of update of the federated large language model.

[0105] The large language model trained by the training method of the large language model provided in the embodiments of the present disclosure can be specifically applied to scenarios where there are data barriers between multiple different data terminals. Specifically, it can be applied to scenarios such as privacy-protected social media scenarios, medical data analysis scenarios, and financial risk control scenarios. In the privacy-protected social media scenario, an object can train a recommendation model based on local data and then combine the update of the federated model to obtain more accurate model parameters without uploading the local data to the server of the recommendation system, thereby improving the accuracy of the recommendation model on the basis of protecting the privacy and data security of the object. In the medical data analysis scenario, since medical data is highly private and sensitive, hospitals and medical institutions can train medical data analysis models locally without sharing the original patient data and then combine the update of the federated model to obtain more accurate model parameters, so as to provide more accurate medical data analysis results for doctors. Similarly, in the financial risk control scenario, banks and financial institutions can let the large language model learn rich knowledge by training the large language model locally without disclosing or sharing the financial data of customers, thereby improving the ability of the large language model to control financial risks. The above examples do not limit the scope of protection of this case.

[0106] General description of the embodiments of the present disclosure

[0107] According to an embodiment of the present disclosure, a training method for a large language model is provided. As Figure 2 shown, it is a schematic flowchart of a training method for a large language model provided by the present disclosure. This method can be applied to a training device for a large language model, and the training device for the large language model can be integrated in a computer device, and the computer device can specifically be a local terminal in a federated learning system. The training method for the large language model may include:

[0108] Step 210, obtaining the first training sample data encrypted by quantum based on a preset quantum key.

[0109] As introduced above, the large language model provided by the present disclosure can be applied to various scenarios, such as medical analysis scenarios, financial risk control scenarios, and social media scenarios, etc. In the embodiments of the present disclosure, the medical analysis scenario is taken as an example to introduce the specific solutions of the present disclosure in detail.

[0110] For large language models applied to medical analysis scenarios, a large amount of medical scenario corpora are required for training. The corpus data volume of a single medical institution or hospital is insufficient to ensure that the trained large language model has good generalization ability. Moreover, patients' medical data are highly private data, and it is impossible to train a large language model by aggregating the corpus data of multiple medical institutions or hospitals. Therefore, the embodiments of the present disclosure provide a method for training a large language model using federated learning, aiming to fully utilize the corpus data of multiple medical institutions or hospitals to train the large language model without sharing local data, so as to obtain a large language model with good generalization ability.

[0111] As Figure 3 shown, it is a schematic diagram of the architecture of the federated learning system provided in the embodiments of the present disclosure. As shown in the figure, the federated learning system includes a federated model 310 (i.e., the aforementioned federated large language model) and multiple local models 320 (i.e., the aforementioned local large language models). Among them, the federated model is deployed in a central server (i.e., the aforementioned federated central server), and the local models are deployed in local terminals. The federated model and the local models may have the same model structure. In the embodiments of the present disclosure, when training the large language model, the core is to train the federated model, and the training of the local models is only to learn knowledge from each local data to assist the training of the federated model. The local models can transmit model parameters to each other to achieve cyclic distillation learning. After the federated model is trained, the model parameters of the federated model can be shared to each local terminal, and then the large language model is deployed on each local terminal based on the final model parameters.

[0112] Regardless of the training process of the federated model or the local models, it is a process of continuously updating parameters in multiple cycles. In one federated model training cycle, the local models need to go through multiple cycles until convergence before outputting the parameters of the local models to the central server. Steps 210 to 260 in the embodiments of the present disclosure will introduce the training method of the large language model taking one federated model training cycle as an example.

[0113] As described above, the local model is deployed in a local terminal, where the local terminal is specifically the terminal for training the local model, rather than necessarily being the local data storage terminal. For example, for a hospital, a terminal can be designated as the terminal for training the local model corresponding to the hospital, but the medical data of the hospital can be stored in multiple doctor terminals or in the hospital database. That is, when training the local model on the local terminal, it is necessary to first obtain the training sample data from multiple doctor terminals or the hospital database of the hospital. In the embodiments of the present disclosure, in order to further prevent the privacy data from being leaked due to being accessed by other third parties during the process of transmitting the training sample data from the doctor terminal or the hospital database to the local terminal, the quantum encryption technology is adopted in the embodiments of the present disclosure to encrypt and transmit the training sample data.

[0114] Therefore, for any local terminal, when training the local model, it can first obtain the quantum-encrypted training sample data from different data storage terminals based on a preset quantum key. To distinguish it from other training sample data in the embodiments of the present disclosure, it is referred to as the first training sample data here. The local terminal and multiple data storage terminals here can form a meta-federation, and local data can be transmitted within the meta-federation, while local data cannot be transmitted between different meta-federations, and only model parameters can be transmitted. As Figure 4 shown, it is a schematic diagram of multiple meta-federations in the federated learning system of the present disclosure. There are data barriers between different meta-federations, and local data cannot be transmitted.

[0115] After obtaining the first training sample data, if the corresponding source data, such as the data in the doctor terminal, changes, the training sample data can be updated simultaneously. That is, the training sample data in the local terminal can be accessed and modified by other terminals in the same meta-federation. In some embodiments, the method for training the large language model provided by the present disclosure further includes:

[0116] In response to a target operation on the first training sample data, obtaining the quantum entropy and ecological coefficient of the first training sample data, where the target operation includes at least one of an access operation and a modification operation;

[0117] Updating the preset quantum key of the first training sample data according to the quantum entropy and the ecological coefficient.

[0118] In the embodiments of the present disclosure, when the training sample data in the local terminal is accessed or modified, the preset quantum key corresponding to the training sample data can be dynamically adjusted to further protect data security. Specifically, in response to a target operation on the first training sample data, the quantum entropy and ecological coefficient of the first training sample data can be obtained. Here, the target operation can be an access operation, a modification operation, or an operation that is both accessed and modified. After obtaining the quantum entropy and ecological parameters of the first training sample data, the preset quantum key of the first training sample data can be updated based on the quantum entropy and ecological parameters of the first training sample data.

[0119] In some embodiments, the first training sample data includes a plurality of sub-sample data, and each sub-sample data is associated with a quantum state. Obtaining the quantum entropy and ecological coefficient of the first training sample data includes:

[0120] Calculating the probability value of observing the quantum state based on the quantum state;

[0121] Calculating the quantum entropy of the first training sample data according to the probability value corresponding to each sub-sample data;

[0122] Determining the interaction coefficient between any two sub-sample data;

[0123] Calculating the ecological coefficient of the first training sample data according to the interaction coefficient.

[0124] In the embodiments of the present disclosure, a specific method for determining the quantum entropy and ecological coefficient is provided. That is, in the embodiments of the present disclosure, the first training sample data may include a plurality of sub-sample data, and each sub-sample data is associated with a quantum state. Thus, the probability value of observing each quantum state can be calculated based on the quantum state associated with each sub-sample data, and then the quantum entropy of the first training sample data can be calculated according to the probability value corresponding to each sub-sample data.

[0125] Furthermore, for the plurality of sub-sample data, the interaction coefficient between any two sub-sample data can be determined first, and then the ecological coefficient of the first training sample data can be calculated according to the interaction coefficient. Among them, the interaction coefficient between two sub-sample data can be specifically determined according to the dependence relationship between the sub-sample data.

[0126] In some embodiments, updating the preset quantum key of the first training sample data according to the quantum entropy and ecological coefficient includes:

[0127] Calculating the ratio between the quantum entropy and the ecological coefficient;

[0128] Calculating the correction value for correcting the preset quantum key according to the ratio;

[0129] Update the preset quantum key based on the correction value.

[0130] In the embodiments of the present disclosure, a specific method for updating the preset quantum key of the first training sample data is provided. In this method, the ratio between the quantum entropy and the ecological coefficient of the first training sample data can be calculated first, and then the correction value for correcting the preset quantum key can be calculated based on this ratio. Further, the preset quantum key of the first training sample data can be updated using this correction value.

[0131] In some embodiments, an ecological dependence relationship graph can be further constructed according to the dependence relationship between sub-sample data. When the first training sample data is accessed or modified, in order to maintain the balance of the data ecosystem, it is necessary to adjust the first training sample data to restore the balance of the overall data ecosystem. Specifically, the balance of the data ecosystem can be controlled by introducing a balance factor. Specifically, it can be calculated whether the balance factor is greater than a preset threshold. When the balance factor does not reach the preset threshold, no processing is required; when the balance factor is greater than the preset threshold, the data needs to be adjusted to restore the ecological balance. Specifically, the redundant data in the training sample data can be deleted to restore the ecological balance.

[0132] Step 220, receive the first large language model parameters sent by the second local terminal in the federated learning system.

[0133] In the embodiments of the present disclosure, a method of circular knowledge distillation is provided to train the local models in each local terminal to further improve the training effect of the local models. That is, in this embodiment, the knowledge of the model is transferred from one meta-federation to the next meta-federation through knowledge distillation, and then so on until it returns to the original meta-federation, forming a complete cycle. During the circular knowledge distillation process, the local model of the previous meta-federation acts as the teacher model, and its knowledge is distilled into the local model of the next meta-federation, that is, the local model of the next meta-federation acts as the student model. As Figure 5 shown, it is a schematic diagram of training the local model using the circular knowledge distillation method. As shown in the figure, each meta-federation receives the model parameters of the previous meta-federation, then determines the teacher model based on the model parameters of the previous meta-federation, and then uses this teacher model and local data to train the local model through knowledge distillation.

[0134] In some embodiments, in a training cycle of the federated model, not all meta-federations in the federated learning system need to be included in the circular knowledge distillation. For example, if the federated learning system contains ten meta-federations, six of them can be randomly selected to construct a circular knowledge distillation system, and then the parameters of the federated model are updated based on the large language model parameters returned by these six meta-federations to the central server.

[0135] Therefore, in the embodiments of the present disclosure, when the first local terminal needs to train the local model, it can first receive the first large language model parameters sent by the previous local terminal (i.e., the second terminal) of the ring knowledge distillation system.

[0136] Step 230: Determine the teacher model according to the first large language model parameters and the local model.

[0137] As can be known from the foregoing introduction, in the central server and multiple local terminals in the federated learning system, the deployed large language models have the same model structure. Therefore, after receiving the first large language model parameters sent by the second terminal, the model parameters of the local model can be replaced according to the first large language model parameters, and the large language model after the second terminal's training can be obtained.

[0138] According to the foregoing introduction, in this embodiment, a ring knowledge distillation method is provided to continuously transfer the knowledge learned by the model. Then, the local model trained in the previous local terminal can be used as the teacher model, that is, the model obtained by replacing the model parameters of the local model with the first large language model parameters is the teacher model.

[0139] Among them, in the knowledge distillation process, the teacher model is the outputter of knowledge, and the student model is the recipient of knowledge. The model structure of the student model can be simpler than that of the teacher model or the same as that of the teacher model. In this embodiment, the model structure of the student model can be the same as that of the teacher model.

[0140] Step 240: Predict the first training sample data based on the teacher model to obtain the predicted probability data output by the teacher model.

[0141] Among them, the training process of the large language model can specifically include an unsupervised pre-training process based on a large amount of corpus data and a fine-tuning process based on high-quality supervised corpus. That is, the training sample data can be divided into unsupervised training sample data and supervised training sample data.

[0142] In the data storage terminal of each meta-federation, the stored data does not necessarily meet the training requirements of the large language model. After the local terminal obtains the stored data from the data storage terminal, it is necessary to further process the stored data to obtain the training sample data.

[0143] Specifically, the generation process of the training sample data can include the following steps:

[0144] Obtain quantum-encrypted document data based on a preset quantum key;

[0145] Clean the document data, perform word segmentation, and sequence padding on the document data to obtain the target text;

[0146] Generate the first training sample data according to the target text.

[0147] That is, in the embodiments of the present disclosure, the data obtained by the local terminal from the data storage terminal in the meta-federation based on the preset quantum key can be encrypted document data. After obtaining the document data, the document data can be cleaned. For example, redundant information can be removed and the document content information can be retained. Then, word segmentation is performed on the document content information, and each word is sequence-padded after word segmentation. The sequence padding process may include truncation operations and filling operations to obtain a target text with a consistent length. Further, word vector conversion can be performed on the target text to obtain a corpus that can be used to train the large language model.

[0148] For a part of the above corpus, its corresponding label information can be determined by adopting manual annotation or machine automatic annotation to obtain the aforementioned high-quality labeled training sample data.

[0149] For these labeled training sample data, a teacher model can be used to process and predict them to obtain multiple output results output by the teacher model and the prediction probability value corresponding to each output result. These output probability values can also be referred to as the soft labels of the training sample data.

[0150] In some embodiments, the target text generates the first training sample data, including:

[0151] Process the target text based on the generator in the preset generative adversarial network to obtain the generated text;

[0152] Generate the first training sample data based on the generated text and the target text;

[0153] Among them, the training process of the generative adversarial network includes the following steps:

[0154] Obtain the second training sample data, where the second training sample data includes data samples and label data of the data samples;

[0155] Update the Lagrange multiplier of the preset Lagrangian function based on the data samples and the label data;

[0156] Calculate the generation loss corresponding to the generator according to the Lagrange multiplier and random noise, and calculate the discrimination loss corresponding to the discriminator according to the data generated by the generator and the data samples;

[0157] Update the parameters of the generator and the discriminator in the generative adversarial network based on the generation loss and the discrimination loss.

[0158] That is, in the embodiments of the present disclosure, by using the Lagrangian relaxation technique to train the generative adversarial network, and then based on the trained generative adversarial network to expand the training sample data, thereby increasing the quantity of the training data of the large language model, and further improving the generalization ability of the large language model.

[0159] Specifically, when training the generative adversarial network, the second training sample data can be obtained first, and the second training sample data includes data samples and the label data corresponding to the data samples. Then, the Lagrange multipliers of the preset Lagrangian function can be updated based on the data samples and the label data. Among them, the goal of Lagrangian relaxation is to maximize the preset Lagrangian function, specifically, the Lagrange multipliers can be updated by gradient ascent to find the maximum value of the preset Lagrangian function. Then, the generation loss of the generator in the generative adversarial network can be calculated according to the Lagrange multipliers and the random noise, and the discrimination loss corresponding to the discriminator can be calculated according to the difference between the data generated by the generator and the true value of the data samples. Further, the parameters of the generator and the discriminator in the adversarial network can be updated based on the above losses to complete the training of the adversarial generative network. After the training is completed, the generator can be used to generate training sample data to expand the data volume of the training sample data.

[0160] Step 250, training the local model according to the first training sample data and the prediction probability data to obtain a student model.

[0161] Further, after determining the prediction probability data output by the teacher model for the first training sample data, knowledge distillation learning can be further performed on the local model based on the first training sample data and the prediction probability data.

[0162] In some embodiments, training the local model according to the first training sample data and the prediction probability data to obtain a student model includes:

[0163] Predicting the input data in the first training sample data based on the local model to obtain output data;

[0164] Determining a first loss according to the difference between the output data and the label data in the first training sample data;

[0165] Calculating a second loss according to the difference between the output data and the corresponding prediction probability data;

[0166] Adjusting the parameters of the local model according to the first loss and the second loss to obtain a student model.

[0167] Among them, when training the local model based on the first training sample data and the prediction probability data, specifically, the input data in the first training sample data can be predicted based on the local model first to obtain the output data. Among them, the input data in the first training sample data here can specifically be the input data corresponding to the aforementioned labeled high-quality training sample data.

[0168] Then, the first loss can be calculated according to the difference between the output data and the label data in the first training sample data. The first loss here can specifically be the cross-entropy loss.

[0169] Then, the second loss between the output data and the prediction probability data can be calculated. The second loss here can specifically be called the distillation loss. Further, the global loss can be calculated according to the first loss and the distillation loss, and the model parameters of the local model can be further adjusted according to the global loss to achieve the local training of the local model.

[0170] In some embodiments, calculating the second loss according to the difference between the output data and the corresponding prediction probability data includes:

[0171] Obtain the temperature parameter corresponding to knowledge distillation;

[0172] Calculate the cross-entropy between the output data and the corresponding prediction probability data based on the temperature parameter;

[0173] Determine the second loss according to the cross-entropy.

[0174] Among them, in the embodiments of the present disclosure, the specific calculation process of the distillation loss can be to first obtain the temperature parameter corresponding to knowledge distillation, and then the cross-entropy between the output data and the prediction probability data can be calculated based on the temperature parameter, so as to obtain the distillation loss.

[0175] In some embodiments, before training the large language model, the model structure of the large language model can be determined first. In the embodiments of the present disclosure, the unit model can include a feature extraction network and a classification network. The model structure of the feature extraction network can be determined by the following method:

[0176] Obtain multiple candidate model structures;

[0177] Calculate the gravitational wave score corresponding to each candidate model structure;

[0178] Perform regularization processing on each candidate model structure based on a preset regularization method to obtain a regularization factor;

[0179] Calculate the comprehensive evaluation score of each candidate model structure according to the gravitational wave score and the regularization factor;

[0180] Determine the structure of the large language model among multiple candidate model structures according to the comprehensive evaluation score.

[0181] In the embodiments of the present disclosure, when determining the model structure of the feature extraction network, multiple candidate model structures can be obtained first. Among them, obtaining multiple candidate model structures can specifically be to initialize the search space first, that is, define a search space containing various possible network structures, and then perform optimization of the model structure in this search space.

[0182] Furthermore, the gravitational wave score corresponding to each candidate model structure can be calculated for each candidate model structure, and the gravitational wave score corresponding to each candidate model structure is obtained. Then, each candidate model structure is regularized by using a preset regularization method. Specifically, bio-inspired regularization can be used here, and this regularization will simulate the interaction between biological cells to further refine the performance of each neural network architecture. The regularization process can obtain the regularization factor corresponding to each candidate model.

[0183] Furthermore, the comprehensive evaluation score of each candidate model structure can be calculated based on the gravitational wave score and the regularization factor of each candidate model structure. And further determine that the candidate model structure with the highest comprehensive evaluation score is the structure of the aforementioned feature extraction network.

[0184] In this embodiment, the optimal model structure can be selected from many optional model structures, so that the accuracy of the trained large language model can be further improved.

[0185] Optionally, in some embodiments, when training the local model, the training process of the classification network includes the following steps:

[0186] Initialize the input weights and biases of the classification network;

[0187] Encode the input weights based on a preset quantum encoding function to obtain the initial weight value;

[0188] Perform sparse coding on the hidden layer data determined according to the input features, the initial weight value, and the bias to obtain a sparse coding representation;

[0189] Update the input weights and biases based on the difference between the sparse coding representation and the label corresponding to the input features.

[0190] In the embodiments of the present disclosure, the training of the classification network can be carried out by using an extreme learning machine algorithm that combines sparse coding and quantum coding. The extreme learning machine is a single-hidden-layer feedforward neural network, and its training speed is very fast. The main reason is that its weights are randomly initialized and remain unchanged, and the network is optimized only by minimizing the norm of the output weights. The present invention proposes a new sparse coding strategy that introduces sparsity into the hidden layer nodes, forcing the extreme learning machine to learn more representative and discriminative features.

[0191] Specifically, the input weights and biases of the classification network can be determined first and initialized. Then, the input weights are encoded based on a preset quantum coding function to obtain the initial weight values, and then the hidden layer data is sparsely encoded to obtain a sparse coding representation, where the hidden layer data is calculated based on the input features, the initial weight values, and the biases. Further, the input weights and biases are updated based on the difference between the sparse coding representation and the labels corresponding to the input features, so as to obtain the trained classification network.

[0192] Step 260, sending the second largest language model parameters of the student model to the third local terminal and sending the second largest language model parameters to the federal central server.

[0193] After the local training of the large language model is completed at the local terminal, the large language model parameters of the trained large language model, that is, the second largest language model parameters, can be further sent to other local terminals to continue the circular knowledge distillation training. At the same time, the local terminal can also send the second largest language model parameters to the federal central server for parameter update of the federal large language model.

[0194] In some embodiments, before sending the second largest language model parameters of the student model to the third local terminal so that the third local terminal updates the parameters of the local model based on the second largest language model parameters, it further includes:

[0195] Obtaining the hierarchical labels and category labels of each local terminal in the federated learning system;

[0196] Obtaining the transmission records of the large language model parameters among multiple local terminals;

[0197] Determining the third local terminal according to the transmission records, the hierarchical labels, and the category labels.

[0198] That is, in the embodiments of the present disclosure, by hierarchically and categorically processing multiple local terminals, different local terminals can be selected to participate in the training of the federated model in different rounds of training of the federated model. In this way, it is possible to avoid the problem that all local terminals need to participate in each round of training, resulting in low training efficiency of the large language model, and it is also possible to make full use of the local data of all local terminals to train the large language model, so that the large language model has better generalization ability.

[0199] Specifically, before sending the second large language model parameters to the third local terminal, the hierarchical label and category label of each local terminal in the federated learning system can be obtained first. Then, obtain the transmission record of the large language model parameters among multiple local terminals. This transmission record can specifically be the transmission record in this round of federated model training, that is, in which local terminals the transmission and local model training have been carried out. Further, the third local terminal to which the second large language model parameters need to be transmitted next can be determined according to the transmission record, hierarchical label, and category label.

[0200] In some embodiments, determining the third local terminal according to the transmission record, hierarchical label, and category label includes:

[0201] Extract the historical hierarchical label and historical category label of the local terminals that have undergone the transmission of the large language model parameters from the transmission record;

[0202] Based on the preset expected hierarchical label, expected category label, historical hierarchical label, and historical category label, determine the target hierarchical label and target category label corresponding to the local terminal to be transmitted;

[0203] Determine the third local terminal according to the matching result between the target hierarchical label and target category label and the hierarchical label and category label.

[0204] In the embodiments of the present disclosure, when determining the next third local terminal to which the second large language model parameters need to be transmitted. Specifically, the historical hierarchical label and historical category label of the local terminals that have been transmitted can be extracted from the transmission record first. Then, determine the preset expected hierarchical label and expected category label according to the transmission record. Further, compare the expected hierarchical label and expected category label with the historical hierarchical label and historical category label, and the difference value between the two is the hierarchical label and category label of the local terminal that can be transmitted. Then any one of them can be determined as the third local terminal.

[0205] In some embodiments, the process of the federated central server updating the parameters of the federated large language model based on the large language model parameters returned by each local terminal includes the following steps:

[0206] Obtain the weight parameters and the number of samples corresponding to multiple target local terminals that return large language model parameters during the parameter update process of the current round of the federated large language model;

[0207] Weight the large language model parameters returned by multiple target local terminals based on the weight parameters and the number of samples to obtain global model parameters;

[0208] Update the parameters of the federated large language model to the global model parameters.

[0209] In the embodiments of the present disclosure, when the federated central server receives the large language model parameters returned by all local terminals specified in the current round of training process, it can perform the parameter update of the current round of the federated large language model according to the returned large language model parameters. Specifically, it can first obtain the weight parameters and the number of samples corresponding to multiple target local terminals that return large language model parameters during the update process of the current round of the federated large language model. Then, it can perform a weighting process on the large language model parameters returned by multiple target local terminals based on the weight parameters and the number of samples to obtain global model parameters.

[0210] The training method of the large language model provided by the embodiments of the present disclosure is applied to a first local terminal in a federated learning system. The federated learning system includes a federated central server and multiple local terminals. The method obtains quantum-encrypted first training sample data based on a preset quantum key; receives first large language model parameters sent by a second local terminal in the federated learning system; determines a teacher model according to the first large language model parameters and the local model; predicts the first training sample data based on the teacher model to obtain predicted probability data output by the teacher model; trains the local model according to the first training sample data and the predicted probability data to obtain a student model; sends second large language model parameters of the student model to a third local terminal so that the third local terminal updates the parameters of the local model based on the second large language model parameters; and sends second large language model parameters of the student model to the federated central server so that the federated central server updates the parameters of the federated large language model based on the large language model parameters returned by each local terminal.

[0211] The embodiments of the present disclosure comprehensively apply the data of multiple local terminals with data barriers due to data security considerations to train the large language model by adopting federated learning technology without sharing the data in any local terminal to other terminals. In this way, the number of training corpora of the large language model can be greatly expanded, thereby improving the accuracy of the large language model. In addition, the embodiments of the present disclosure also adopt the method of knowledge distillation to continuously accumulate the knowledge learned from the data of different local terminals, thereby further improving the accuracy of the trained large language model.

[0212] Detailed description of the embodiments of the present disclosure in combination with specific application scenarios

[0213] As Figure 6 shown, it is another schematic flowchart of the training method of the large language model provided by the present disclosure. In this embodiment, the training method of the large language model will be introduced in detail in combination with the execution subject of each step. The method specifically includes the following steps:

[0214] Step 601, the central server of the federated learning system determines the model structure of the large language model.

[0215] Among them, the large language model training method provided by the embodiments of the present disclosure is applied in the federated learning system. For the architecture diagram of the federated learning system, please continue to refer to Figure 3 . In the embodiments of the present disclosure, before training the large language model, the model structure of the large language model can be determined first. Among them, the model structure of the large language model can specifically be composed of a feature extraction network and a classification network.

[0216] In the embodiments of the present disclosure, the feature extraction network can be determined by using a neural network algorithm based on optimal architecture search. The purpose of optimal architecture search is to automatically find the best neural network structure. Traditional neural network design requires a lot of domain knowledge and experiments. Optimal architecture search automates this process through a search algorithm to find the best network architecture that meets the specific task requirements. The process of determining the optimal structure of the feature extraction network by using the optimal architecture search method will be introduced in detail below.

[0217] Specifically, the initial search space can be determined according to the task requirements of the large language model: that is, a search space containing various possible network structures matching the task requirements is defined.

[0218] Then, the gravitational wave fluctuation score is calculated for each network architecture in the search space. This score is based on the simulated gravitational wave theory and is positively correlated with the expected performance of the architecture. Considering the gravitational wave theory in physics, the gravitational wave fluctuation score S(A) of each network architecture A is defined as:

[0219] S(A) = ∫F(A,t)dt

[0220] Among them, F(A,t) represents the gravitational wave fluctuation of network architecture A at time t, and t is the time parameter, simulating the performance of the architecture at different training times.

[0221] Furthermore, bio-inspired regularization is applied to each network architecture in the search space. This will simulate the interaction between biological cells and further refine the expected performance of each architecture. The interaction between biological cells can be described by a regularization factor R(A), which is proportional to the complexity of network architecture A:

[0222] R(A) = α × Complexity(A)

[0223] Among them, α is the regularization coefficient, and Complexity(A) represents the complexity of the network architecture A.

[0224] Then, the expected performance of each network architecture in the search space can be evaluated by combining the gravitational wave fraction and bio-inspired regularization. The comprehensive evaluation score E(A) of the network architecture A is:

[0225] E(A) = S(A) - λ × R(A)

[0226] Among them, λ is a weight parameter that determines the importance of regularization.

[0227] Then, the network architecture with the highest comprehensive evaluation score E(A) can be selected as the optimized network architecture. Further, fine-grained optimization can be performed on the selected network architecture, such as depth, width, connection mode, etc. Then, the optimized network architecture can be used for feature extraction. This step transforms the original data into a set of features highly relevant to the target task. The optimized network architecture is used to extract features from the input data X, and the feature extraction function is defined as Φ:

[0228] F = Φ(X; A * )

[0229] Among them, F is the extracted feature set, and A * is the network architecture with the highest score.

[0230] Then, the above score evaluations are respectively performed on multiple possible network structures in the initialized search space. When the search time reaches the preset search time or meets other search termination conditions, the A * with the highest evaluation score among all the currently evaluated network structures can be determined as the network architecture of the feature extraction network.

[0231] Step 602, the central server of the federated learning system sends the model structure of the large language to each local terminal.

[0232] After determining the model structure of the large language model according to the task requirements of the large language model, the model structure can be sent to all local terminals in the federated learning system to deploy the large language model with the above model structure in the central server and each local terminal of the federated learning system.

[0233] Among them, the local terminal is specifically the terminal for training the local model. Generally, one local terminal corresponds to one meta-federation. A meta-federation can include one local terminal and multiple clients, and the training corpus data can be stored in the clients. A certain number of clients form a single meta-federation, and there are differences in the data between different meta-federations.

[0234] Further, for each local terminal, a dynamic identifier can be assigned, and then multiple local terminals can be divided into different levels and different categories according to a multi-dimensional parameter space. Among them, the identifier of the local terminal can be calculated based on the attributes and characteristics of the members in the corresponding meta-federation. The formula for the dynamic representation can be expressed as:

[0235] ID i = F(a i ,b i ,c i ,...,z i )

[0236] Among them, ID i is the dynamic identifier of the i-th local terminal or the i-th meta-federation.

[0237] F() is a multi-dimensional characteristic function, and a i ,b i ,...,z i are the multi-dimensional characteristic parameters of the i-th meta-federation, including but not limited to network topology structure characteristics, data distribution characteristics, etc.

[0238] Further, classifying each meta-federation according to its multi-dimensional parameter space can be expressed as:

[0239] C i = G(ID i ,V ij )

[0240] Among them, C i is the classification of the i-th meta-federation; G is a classification function that classifies according to network topology structure characteristics, data distribution characteristics, etc.; V ij is the relationship vector between the i-th meta-federation and the j-th meta-federation, indicating the correlation between meta-federations.

[0241] Further, meta-federations are grouped into different levels according to their attributes and behavioral characteristics, and the organizational structures of different levels can be expressed using the following formula:

[0242] H i = H(C i ,W i )

[0243] Among them, H i is the level where the i-th meta-federation is located, and H is a hierarchical division function based on the classification categories of meta-federations; W i is a weight factor that adjusts its position in the federation structure based on the key attributes of the i-th meta-member.

[0244] Step 603, during one round of training of the federated model, the central server determines multiple target local terminals among multiple local terminals.

[0245] When training the federated model, the number of training rounds of the federated model can be set. For example, the number of training rounds can be set to T rounds. Also, the number of training rounds of the local model can be set during each round of training of the federated model. For example, the number of training rounds of the local model can be set to M rounds. In addition, during each round of training of the federated model, a certain number of target local terminals can be determined from all local terminals in the federated learning system. For example, K target local terminals can be determined, and then the training process of this round of the federated model can be based on the local data of these K target local terminals to train the federated model.

[0246] Among them, during each round of training of the federated model, the central server can uniformly select the local terminals of the corresponding meta-federation in each layer or each category as target local terminals according to the hierarchical information and category information of the meta-federation corresponding to the local terminal.

[0247] Step 604, the target local terminal obtains training sample data and trains the local model based on the training sample data and the model parameters sent by the previous target local terminal.

[0248] The target local terminal can first generate training sample data, which can be obtained by processing text data such as books, web pages, documents, and research papers in the open domain. These data include, but are not limited to, fields such as education, technology, art, economy, and news. All data are stored in pure text form (.txt). To ensure the integrity and completeness of the data, each piece of data is accompanied by its source information and collection date.

[0249] Specifically, the attributes of the obtained source data can include: document ID, source, collection date, content, author, title, field, keywords, references, etc. The dataset D can be expressed as:

[0250] D = {d1, d2,..., d n}

[0251] Among them, D represents the dataset, and d i represents a single document.

[0252] Then, redundant data can be removed from each document, and redundant and irrelevant attribute information can be removed to obtain: d′ i = content i; for the document data with redundant data removed, text cleaning can be further performed. Text cleaning can specifically include: unifying case, removing special symbols and numbers, removing stop words, etc. Text cleaning can be expressed as:

[0253] d″i = clean(d′ i )

[0254] where clean is a text cleaning function.

[0255] Furthermore, word segmentation is performed. Since Chinese is space-less, the document needs to be segmented into individual words, which can be expressed as:

[0256] d″′ i = tokenize(d″ i )

[0257] where tokenize is a word segmentation function that segments the document d″ i into a sequence of words w1, w2,..., w n .

[0258] Furthermore, word vector representation is performed. To convert text data into a numerical format that can be understood by machines, pre-trained word vectors are used to encode each word. That is, for each word w j , there is a corresponding vector v j , which can be expressed as:

[0259] v j = word2vec(w j )

[0260] where word2vec is a function that converts words into vectors.

[0261] Furthermore, sequence padding is performed. Since neural network models require fixed-length inputs, the word vector sequences of each document need to be padded or truncated to reach a fixed length L. The padding value is set to the zero vector, which can be expressed as:

[0262] d″′ i = pad(d″′ i , L)

[0263] where pad is a sequence padding function.

[0264] Thus, unlabeled training sample data corresponding to each local model can be obtained. Then, these unlabeled training sample data need to be further marked to obtain high-quality labeled training sample data.

[0265] Furthermore, in the embodiments of the present disclosure, sample data generation can also be further performed through a generative adversarial network based on the Lagrangian relaxation technique to achieve data augmentation.

[0266] Specifically, a generative adversarial network includes a generator and a discriminator. Embodiments of the present disclosure introduce Lagrangian relaxation techniques to transform the optimization problem by adding constraints to the objective function. Lagrangian relaxation is used to constrain the generator so that the generated data is not only real but also satisfies certain specific properties or distributions, which are provided by fine-grained labels (such as certain specific topics or sentiments).

[0267] Specifically, the parameters of the generator G and the discriminator D can be randomly initialized first. Then, fine-grained labels are set for each data sample, and these labels describe certain characteristics of the data. Define a label function L: x → y, where x is the input data and y is the fine-grained label of this data, which can be expressed as:

[0268] y = L(x)

[0269] Furthermore, based on the currently generated data and fine-grained labels, the Lagrange multipliers are updated. Ensure that the generated data meets the given constraints. The goal of Lagrangian relaxation is to maximize the following Lagrangian function:

[0270]

[0271] where λ is the Lagrange multiplier, constraints represent the constraint conditions, which can be the expected values of certain properties or distributions of the generated data. E is the expected value, p data (x) is the distribution of real data, p z (z) is the distribution of noise data, D(x) is the discriminator output, representing the probability that x is real data, G(z) is the generator output, representing the data generated from noise z, and λ is the Lagrange multiplier.

[0272] For the update of the Lagrange multiplier, the gradient ascent method is adopted:

[0273]

[0274] where α is the learning rate and t represents the current iteration number, is the gradient operator.

[0275] Then, the generation loss function of the generator can be calculated:

[0276]

[0277] Then, the gradient descent method is used to update the generator based on this generation loss function:

[0278]

[0279] where β is the learning rate, θ G ,θD are the parameters of the generator and discriminator respectively.

[0280] Furthermore, data x is generated using the noise z: ′ x = G(z), and the discriminator D is used to evaluate the authenticity of the data.

[0281] Furthermore, the generator G is updated based on the feedback of the discriminator D and the Lagrangian relaxation technique.

[0282] Then, the trained generator can be used to augment the training sample data.

[0283] In the embodiments of the present disclosure, during each round of training of the federated model, multiple target local terminals train the local model in a circular knowledge distillation manner.

[0284] Furthermore, in the data storage and access mechanism between local terminals corresponding to each meta-federation, the propagation process between local terminals is divided into two stages, specifically the common knowledge accumulation stage and the personalization stage. In these two stages, the model parameters are sequentially passed between local terminals, and adaptive information exchange is performed through knowledge distillation.

[0285] In the first stage, a dynamically changing quantum key can be assigned to each training sample data. When the data needs to be accessed or modified, the key will be dynamically adjusted. At the same time, a dependency relationship graph between data is established. When a certain data is accessed or changed, the corresponding adjustments need to be made to other associated data to maintain the balance of the overall data ecosystem.

[0286] Specifically, first, according to the characteristics of the training sample data and the principles of quantum mechanics, an initial quantum key is generated for each data block. Quantum entropy can be regarded as a measurement of the uncertainty of a quantum system. For a given training sample data D, its quantum entropy is defined as S(D). Suppose the training sample data D consists of n sub-data, that is, D = {d1, d2,..., d n}, and each sub-data d i is associated with a quantum state |ψ i >. The quantum entropy of the training sample data is determined by the sum of the quantum entropies of all its sub-data, and the quantum entropy can be expressed as:

[0287]

[0288] where p i is the probability of observing the quantum state |ψ i >. This probability can be determined by the Born rule:

[0289] p i = |<ψ i |ψi >| 2

[0290] Furthermore, the training sample data, its corresponding quantum key, and the ecological dependency relationship are stored together in a specific storage structure, such as a quantum blockchain or an ecological database. When there is an external request to access a certain data, first verify the identity and permissions of the requester, and then dynamically adjust its corresponding quantum key.

[0291] Furthermore, after the training sample data is accessed or modified, check the entire data ecosystem to determine whether it is necessary to adjust the associated data. Considering the dependency relationship between the training sample data, define the ecological coefficient E(D), which is determined by the interactions between all sub-data within the data block D. If there is a dependency relationship between d i and d j , there is an interaction coefficient α ij , otherwise α ij = 0. The calculation method of the ecological coefficient can be expressed as:

[0292]

[0293] When the training sample data D is accessed, its quantum key needs to be dynamically adjusted. Define the original key as K(D) and the adjusted key as K ′ (D), and their relationship can be expressed as:

[0294]

[0295] Furthermore, according to the ecological dependency relationship graph, process the data that needs to be adjusted to restore the balance of the overall data ecosystem. To maintain the balance of the data ecosystem, this embodiment introduces a balance factor β(D), which is jointly determined by the quantum entropy and the ecological coefficient of the training sample data and can be expressed as:

[0296]

[0297] When β(D) exceeds a predetermined threshold T, that is, β(D) > T, then it is necessary to adjust the training sample data to restore the ecological balance. Specifically, redundant data in the training sample data can be deleted.

[0298] In the second stage, the model can be sent to the next meta-federation without local training to prevent the loss of common knowledge due to excessive local training. After the local terminals in each meta-federation perform local training on the local model, the strategy of circular knowledge distillation can be adopted to transfer the model parameters from one meta-federation to the next, and then so on until the original meta-federation is returned.

[0299] Among them, when the target local terminal trains the local model according to the training sample data obtained by the above method and the model parameters sent by the previous target local terminal, the target loss function can be calculated as follows:

[0300] L = αL CE +(1 - α)T 2 L KD

[0301] Among them, L CE is the cross-entropy loss; L KD is the knowledge distillation loss, usually the KL divergence of the output probabilities of the teacher and student models; α is a weight parameter; T is a temperature parameter used to soften the probability distribution.

[0302] L KD can be expressed as:

[0303] L KD = KL(softmax(o student / T)||softmax(o teacher / T))

[0304] Among them, o student is the output of the student model, o teacher is the output of the teacher model, and T is the temperature parameter.

[0305] Among them, when updating the parameters of the local model based on the above target loss function, the extreme learning machine algorithm combining sparse coding and quantum coding can also be used to further train the classification network of the local model to obtain a more accurate classification network.

[0306] Specifically, the parameters of the classification network include the input weight parameter W in and the bias parameter b in , and the number of nodes in the hidden layer is h. The sample data is subjected to feature extraction to obtain the feature x i , and its corresponding label is y i . The quantum algorithm can be used to optimize the input weight parameter:

[0307] W q = Q(W in )

[0308] Among them, Q represents the quantum coding function. W q is the initial weight value optimized by the quantum algorithm.

[0309] Then, the input feature is transformed into the hidden layer space using the optimized initial weight value and bias parameter, which is expressed as:

[0310] H = σ(W q xi +b in )

[0311] Among them, σ is an activation function, such as the sigmoid function.

[0312] Since the training process expects the representations of the hidden layer to be sparse, this disclosure introduces a sparsity constraint such that the activation values of some nodes in the hidden layer are zero. The hidden layer after sparse coding is represented as:

[0313] H s = S(H, λ)

[0314] Among them, S is a sparsity function that sets some activation values to zero, and λ is a parameter of the sparsity constraint.

[0315] Furthermore, the output weight W out is trained by minimizing the difference between the output of the hidden layer after sparse coding and the true label. The formula is:

[0316]

[0317] By solving the above least squares problem, the output weight W out .

[0318] Then the performance of the model can be evaluated on the validation set. The output of the model is:

[0319]

[0320] The performance of the model on the validation set can be evaluated by comparing and the true label y i .

[0321] To further optimize the model, update the input weights and biases, which can be expressed as:

[0322]

[0323]

[0324] Among them, L is the mean squared error loss function, and α is the learning rate.

[0325] Through multiple rounds of iterative optimization until the convergence of the model parameters is achieved, the network parameters of a more accurate classification network can be obtained.

[0326] Step 605, the target local terminal sends the local model parameters to the next target local terminal.

[0327] After training the local model to obtain accurate feature extraction network parameters and classification network parameters, the local model parameters composed of the feature extraction network parameters and the classification network parameters can be further sent to the next target local terminal for distillation learning.

[0328] Among them, the next target local terminal here can specifically be the target local terminal among the foregoing multiple target local terminals that has not undergone model parameter transmission.

[0329] Step 606, the target local terminal sends the local model parameters to the central server.

[0330] In addition, after the target local terminal trains the local model to obtain the model parameters, it is also necessary to further send the trained local model parameters to the central server so that the central server can update the parameters of the federated model according to the local model parameters returned by the target local terminal.

[0331] Step 607, the central server obtains the weight parameters and sample data volumes of each target local terminal.

[0332] After receiving the local model parameters sent by each target local terminal, the central server can further obtain the weight parameters and sample data volumes of each target local terminal.

[0333] Step 608, the central server updates the federated model parameters according to the received local model parameters, the weight parameters of the target local terminal, and the sample data volume.

[0334] After the central server obtains the local model parameters sent by each target local terminal, the weight parameters of each target local terminal, and the sample data volume, it can further calculate the model parameters of the federated model based on the above data:

[0335]

[0336] where n k is the number of training data of each participating node, and and w t+1 is the model parameter of the federated model after the (t + 1)-th round of training, is the model parameter returned by the k-th local model in the (t + 1)-th round of training of the federated model.

[0337] Step 609, the central server repeatedly initiates multiple rounds of training of the federated model until the federated model converges.

[0338] Furthermore, the central server repeats the above steps, initiates multiple rounds of training of the federated model until the federated model converges, or the number of training rounds of the federated model reaches the above T rounds.

[0339] Step 610: The central server sends the trained federated model to multiple local terminals for local deployment.

[0340] After obtaining the above accurate federated model through training, the central server can send the model parameters of the trained federated model to each local terminal, so that the large language model can be deployed in each local terminal to perform corresponding task processing.

[0341] Device and equipment description of the embodiments of the present disclosure

[0342] It can be understood that although the steps in the above various flowcharts are sequentially shown according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this embodiment, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above flowcharts may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0343] It should be noted that in each specific implementation manner of the present disclosure, when it comes to performing relevant processing based on data related to the characteristics of the target object, such as target object attribute information or a set of attribute information, the permission or consent of the target object will be obtained first. Moreover, the collection, use, and processing of these data will comply with the relevant laws, regulations, and standards of the relevant region. In addition, when the embodiments of the present application need to obtain target object attribute information, the separate permission or separate consent of the target object will be obtained through pop-up windows or by jumping to a confirmation page. After clearly obtaining the separate permission or separate consent of the target object, the necessary data related to the target object for the normal operation of the embodiments of the present application will be obtained.

[0344] Figure 7 The following is a schematic structural diagram of a training device 700 for a large language model provided by an embodiment of the present disclosure. The device is applied to a first local terminal in a federated learning system. The federated learning system includes a federated central server and multiple local terminals. The device includes:

[0345] An acquisition unit 710, configured to obtain quantum-encrypted first training sample data based on a preset quantum key;

[0346] A receiving unit 720, configured to receive first large language model parameters sent by a second local terminal in the federated learning system;

[0347] A determination unit 730, configured to determine a teacher model according to first large language model parameters and a local model;

[0348] A prediction unit 740, configured to perform a prediction on first training sample data based on the teacher model to obtain prediction probability data output by the teacher model;

[0349] A training unit 750, configured to train the local model according to the first training sample data and the prediction probability data to obtain a student model;

[0350] A sending unit 760, configured to send second large language model parameters of the student model to a third local terminal so that the third local terminal updates parameters of the local model based on the second large language model parameters; and send the second large language model parameters of the student model to a federated central server so that the federated central server updates parameters of the federated large language model based on large language model parameters returned by each local terminal.

[0351] Optionally, in some embodiments, the training device for a large language model provided by the present disclosure further includes:

[0352] A first obtaining subunit, configured to obtain a hierarchical label and a category label of each local terminal in a federated learning system;

[0353] A second obtaining subunit, configured to obtain a transmission record of large language model parameters among multiple local terminals;

[0354] A first determination subunit, configured to determine a third local terminal according to the transmission record, the hierarchical label, and the category label.

[0355] Optionally, in some embodiments, the determination subunit includes:

[0356] An extraction module, configured to extract a historical hierarchical label and a historical category label of a local terminal that has undergone transmission of large language model parameters from the transmission record;

[0357] A first determination module, configured to determine a target hierarchical label and a target category label corresponding to the local terminal to be transmitted based on a preset expected hierarchical label, an expected category label, a historical hierarchical label, and a historical category label;

[0358] A second determination module, configured to determine the third local terminal according to a matching result between the target hierarchical label and the target category label and the hierarchical label and the category label.

[0359] Optionally, in some embodiments, the training device for a large language model provided by the present disclosure further includes:

[0360] A third acquisition subunit, configured to acquire the quantum entropy and the ecological coefficient of the first training sample data in response to a target operation on the first training sample data, where the target operation includes at least one of an access operation and a modification operation;

[0361] A first update subunit, configured to update a preset quantum key of the first training sample data according to the quantum entropy and the ecological coefficient.

[0362] Optionally, in some embodiments, the first training sample data includes a plurality of sub-sample data, and each sub-sample data is associated with a quantum state. The third acquisition subunit includes:

[0363] A first calculation module, configured to calculate the probability value of observing the quantum state based on the quantum state;

[0364] A second calculation module, configured to calculate the quantum entropy of the first training sample data according to the probability value corresponding to each sub-sample data;

[0365] A third determination module, configured to determine the interaction coefficient between any two sub-sample data;

[0366] A third calculation module, configured to calculate the ecological coefficient of the first training sample data according to the interaction coefficient.

[0367] Optionally, in some embodiments, the update subunit includes:

[0368] A fourth calculation module, configured to calculate the ratio between the quantum entropy and the ecological coefficient;

[0369] A fifth calculation module, configured to calculate a correction value for correcting the preset quantum key according to the ratio;

[0370] An update module, configured to update the preset quantum key based on the correction value.

[0371] Optionally, in some embodiments, the training unit includes:

[0372] A prediction subunit, configured to predict the input data in the first training sample data based on a local model to obtain output data;

[0373] A second determination subunit, configured to determine a first loss according to the difference between the output data and the label data in the first training sample data;

[0374] A first calculation subunit, configured to calculate a second loss according to the difference between the output data and the corresponding predicted probability data;

[0375] An adjustment subunit, configured to adjust the parameters of the local model according to the first loss and the second loss to obtain a student model.

[0376] Optionally, in some embodiments, the first computing subunit includes:

[0377] An acquisition module, configured to acquire the temperature parameter corresponding to knowledge distillation;

[0378] A sixth computing module, configured to calculate the cross entropy between the output data and the corresponding predicted probability data based on the temperature parameter;

[0379] A fourth determination module, configured to determine the second loss according to the cross entropy.

[0380] Optionally, the training device for the large language model provided by the present disclosure further includes a federated large language model update unit, and the federated large language model update unit includes:

[0381] A fourth acquisition subunit, configured to acquire the weight parameters and the number of samples corresponding to multiple target local terminals that return the large language model parameters during the current round of federated large language model parameter update;

[0382] A weighting subunit, configured to weight the large language model parameters returned by multiple target local terminals based on the weight parameters and the number of samples to obtain global model parameters;

[0383] A second update subunit, configured to update the parameters of the federated large language model to the global model parameters.

[0384] Optionally, in some embodiments, the acquisition unit includes:

[0385] A fifth acquisition subunit, configured to acquire quantum-encrypted document data based on a preset quantum key;

[0386] A first processing subunit, configured to perform document cleaning, word segmentation, and sequence padding on the document data to obtain a target text;

[0387] A generation subunit, configured to generate first training sample data according to the target text.

[0388] Optionally, in some embodiments, the generation subunit includes:

[0389] A first generation module, configured to process the target text based on a generator in a preset generative adversarial network to obtain a generated text;

[0390] A second generation module, configured to generate first training sample data based on the generated text and the target text;

[0391] Wherein, the training process of the generative adversarial network includes the following steps:

[0392] Acquire second training sample data, where the second training sample data includes data samples and label data of the data samples;

[0393] Update the Lagrange multipliers of the preset Lagrangian function based on data samples and labeled data;

[0394] Calculate the generation loss corresponding to the generator according to the Lagrange multipliers and the random noise generator, and calculate the discrimination loss corresponding to the discriminator according to the data generated by the generator and the data samples;

[0395] Update the parameters of the generator and discriminator in the generative adversarial network based on the generation loss and the discrimination loss.

[0396] Optionally, in some embodiments, the structure of the local model includes a feature extraction network and a classification network, and the training device for the large language model provided by the present disclosure further includes a feature extraction network construction unit, including:

[0397] A sixth acquisition subunit, configured to acquire a plurality of candidate model structures;

[0398] A second calculation subunit, configured to calculate the gravitational wave fluctuation score corresponding to each candidate model structure;

[0399] A second processing subunit, configured to perform regularization processing on each candidate model structure based on a preset regularization method to obtain a regularization factor;

[0400] A third calculation subunit, configured to calculate the comprehensive evaluation score of each candidate model structure according to the gravitational wave fluctuation score and the regularization factor;

[0401] A third determination subunit, configured to determine the structure of the large language model from the plurality of candidate model structures according to the comprehensive evaluation score.

[0402] Optionally, in some embodiments, the training device for the large language model provided by the present disclosure further includes a classification layer training unit, including:

[0403] An initialization subunit, configured to initialize the input weights and biases of the classification network;

[0404] A first encoding subunit, configured to encode the input weights based on a preset quantum encoding function to obtain weight initial values;

[0405] A second encoding subunit, configured to perform sparse encoding on the hidden layer data determined according to the input features, weight initial values, and biases to obtain a sparse encoding representation;

[0406] A third update subunit, configured to update the input weights and biases based on the difference between the sparse encoding representation and the labels corresponding to the input features.

[0407] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the functions of that module or unit.

[0408] Referring to Figure 8 , Figure 8 FIG. is a structural block diagram of a part of the terminal 140 for implementing the training method of the large language model in the embodiments of the present disclosure. The terminal 140 includes components such as a Radio Frequency (RF) circuit 1410, a memory 815, an input unit 830, a display unit 840, a sensor 850, an audio circuit 860, a wireless fidelity (WiFi) module 870, a processor 880, and a power supply 890. Those skilled in the art can understand that Figure 8 the shown structure of the terminal 140 does not limit a mobile phone or a computer, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0409] The RF circuit 810 can be used for receiving and transmitting information or signals during a call. Specifically, after receiving the downlink information from the base station, it is given to the processor 880 for processing; in addition, the designed uplink data is sent to the base station.

[0410] The memory 815 can be used to store software programs and modules. The processor 880 executes various functional applications and document editing of the terminal by running the software programs and modules stored in the memory 815.

[0411] The input unit 830 can be used to receive input digital or character information, and generate key signal inputs related to the settings and function controls of the terminal. Specifically, the input unit 830 may include a touch panel 831 and other input devices 832.

[0412] The display unit 840 can be used to display the input information or the provided information and various menus of the terminal. The display unit 840 may include a display panel 841.

[0413] The audio circuit 860, the speaker 861, and the microphone 862 can provide an audio interface.

[0414] In this embodiment, the processor 880 included in the terminal 140 can execute the training method of the large language model in the previous embodiments.

[0415] The terminal 140 in the embodiments of the present disclosure includes, but is not limited to, mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, etc.

[0416] Figure 9 It is a structural block diagram of a part of the server 110 for implementing the training method of the large language model in the embodiments of the present disclosure. The server 110 may vary greatly due to configuration or performance, and may include one or more central processing units (CPUs) 922 (for example, one or more processors) and a storage device 932, and one or more storage media 930 (for example, one or more mass storage devices) for storing application programs 942 or data 944. Among them, the storage device 932 and the storage media 930 may be transient storage or persistent storage. The program stored in the storage media 930 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server 110. Further, the central processing unit 922 may be configured to communicate with the storage media 930 and execute a series of instruction operations in the storage media 930 on the server 110.

[0417] The server 110 may further include one or more power supplies 926, one or more wired or wireless network interfaces 950, one or more input / output interfaces 958, and / or one or more operating systems 941, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0418] The central processing unit 922 in the server 110 may be used to execute the training method of the large language model in the embodiments of the present disclosure.

[0419] The embodiments of the present disclosure further provide a storage medium for storing program codes for executing the training method of the large language model in the foregoing respective embodiments.

[0420] The embodiments of the present disclosure further provide a computer program product, which includes a computer program. The processor of the computer device reads and executes the computer program, so that the computer device executes to implement the above-mentioned training method of the large language model.

[0421] In the description of the present disclosure and the above-mentioned drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "including" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0422] It should be understood that in the present disclosure, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0423] It should be understood that in the description of the embodiments of the present disclosure, the meaning of "a plurality (or multiple items)" is more than two. Understandings such as greater than, less than, exceeding, etc. do not include the present number, and understandings such as above, below, within, etc. include the present number.

[0424] In several embodiments provided by the present disclosure, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.

[0425] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0426] In addition, in each embodiment of the present disclosure, each functional unit may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0427] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present disclosure. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.

[0428] It should also be understood that the various embodiments provided in the present disclosure can be combined arbitrarily to achieve different technical effects.

[0429] The above is a specific description of the embodiments of the present disclosure, but the present disclosure is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present disclosure, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present disclosure.

Claims

1. A training method for a large language model, characterized in that, The method is applied to a first local terminal in a federated learning system, the federated learning system includes a federated central server and multiple local terminals, and the method includes: Obtain quantum-encrypted first training sample data based on a preset quantum key; Receive first large language model parameters sent by a second local terminal in the federated learning system; Determine a teacher model according to the first large language model parameters and the local model; Predict the first training sample data based on the teacher model to obtain prediction probability data output by the teacher model; Train the local model according to the first training sample data and the prediction probability data to obtain a student model; Send the second large language model parameters of the student model to a third local terminal so that the third local terminal updates the parameters of the local model based on the second large language model parameters; and send the second large language model parameters of the student model to the federated central server so that the federated central server updates the parameters of the federated large language model based on the large language model parameters returned by each local terminal.

2. The method according to claim 1, wherein Before sending the second large language model parameters of the student model to a third local terminal so that the third local terminal updates the parameters of the local model based on the second large language model parameters, it further includes: Obtain the hierarchical label and category label of each local terminal in the federated learning system; Obtain the transmission record of the large language model parameters among the multiple local terminals; Determine the third local terminal according to the transmission record, the hierarchical label, and the category label.

3. The method according to claim 2, wherein The determining the third local terminal according to the transmission record, the hierarchical label, and the category label includes: Extract the historical hierarchical label and historical category label of the local terminal that has undergone the transmission of the large language model parameters in the transmission record; Determine the target hierarchical label and target category label corresponding to the local terminal to be transmitted based on a preset expected hierarchical label, expected category label, the historical hierarchical label, and the historical category label; Determine the third local terminal according to the matching result between the target hierarchical label and target category label and the hierarchical label and the category label.

4. The method according to any one of claims 1 to 3, characterized in that, The method further includes: In response to a target operation on the first training sample data, obtain the quantum entropy and ecological coefficient of the first training sample data, where the target operation includes at least one of an access operation and a modification operation; Update the preset quantum key of the first training sample data according to the quantum entropy and the ecological coefficient.

5. The method according to claim 4, wherein The first training sample data includes multiple sub-sample data, and each sub-sample data is associated with a quantum state. The obtaining the quantum entropy and ecological coefficient of the first training sample data includes: Calculate the probability value of observing the quantum state based on the quantum state; Calculate the quantum entropy of the first training sample data according to the probability value corresponding to each sub-sample data; Determine the interaction coefficient between any two sub-sample data; Calculate the ecological coefficient of the first training sample data according to the interaction coefficient.

6. The method according to claim 4, characterized in that, Updating the preset quantum key of the first training sample data according to the quantum entropy and the ecological coefficient includes: Calculating the ratio between the quantum entropy and the ecological coefficient; Calculating a correction value for correcting the preset quantum key according to the ratio; Updating the preset quantum key based on the correction value.

7. The method according to claim 1, wherein Training the local model according to the first training sample data and the prediction probability data to obtain a student model, including: Predicting input data in the first training sample data based on the local model to obtain output data; Determining a first loss according to the difference between the output data and the label data in the first training sample data; Calculating a second loss according to the difference between the output data and the corresponding prediction probability data; Adjusting the parameters of the local model according to the first loss and the second loss to obtain a student model.

8. The method according to claim 7, wherein Calculating the second loss according to the difference between the output data and the corresponding prediction probability data includes: Obtaining the temperature parameter corresponding to knowledge distillation; Calculating the cross entropy between the output data and the corresponding prediction probability data based on the temperature parameter; Determining the second loss according to the cross entropy.

9. The method according to claim 1, characterized in that The process of the federal central server updating the parameters of the federated large language model based on the large language model parameters returned by each local terminal includes the following steps: Obtaining the weight parameters and the number of samples of multiple target local terminals that return the large language model parameters during the current round of federated large language model parameter update; Weighting the large language model parameters returned by the multiple target local terminals based on the weight parameters and the number of samples to obtain global model parameters; Updating the parameters of the federated large language model to the global model parameters.

10. The method according to claim 1, wherein Obtaining the first training sample data encrypted by quantum based on the preset quantum key includes: Obtaining the document data encrypted by quantum based on the preset quantum key; Performing document cleaning, word segmentation, and sequence padding on the document data to obtain a target text; Generating the first training sample data according to the target text.

11. The method according to claim 10, wherein Generating the first training sample data according to the target text includes: Processing the target text based on the generator in the preset generative adversarial network to obtain a generated text; Generating the first training sample data based on the generated text and the target text; Among them, the training process of the generative adversarial network includes the following steps: Obtaining second training sample data, where the second training sample data includes data samples and label data of the data samples; Updating the Lagrange multiplier of the preset Lagrangian function based on the data samples and the label data; Calculating the generation loss corresponding to the generator according to the Lagrange multiplier and random noise, and calculating the discrimination loss corresponding to the discriminator according to the data generated by the generator and the data samples; Updating the parameters of the generator and the discriminator in the generative adversarial network based on the generation loss and the discrimination loss.

12. The method according to claim 1, wherein The structure of the local model includes a feature extraction network and a classification network. The feature extraction network is constructed through the following steps: Obtaining multiple candidate model structures; Calculate the gravitational wave score corresponding to each of the candidate model structures; Perform regularization processing on each of the candidate model structures based on a preset regularization method to obtain a regularization factor; Calculate the comprehensive evaluation score of each candidate model structure according to the gravitational wave score and the regularization factor; Determine the structure of the large language model among the multiple candidate model structures according to the comprehensive evaluation score.

13. The method according to claim 12, characterized in that When training the local model, the training process of the classification network includes the following steps: Initialize the input weights and biases of the classification network; Encode the input weights based on a preset quantum encoding function to obtain initial weight values; Perform sparse coding on the hidden layer data determined according to the input features, the initial weight values, and the biases to obtain a sparse coding representation; Update the input weights and the biases based on the difference between the sparse coding representation and the labels corresponding to the input features.

14. A training device for a large language model, characterized in that, The device is applied to the first local terminal in a federated learning system. The federated learning system includes a federated central server and multiple local terminals. The device includes: An acquisition unit, configured to acquire quantum-encrypted first training sample data based on a preset quantum key; A receiving unit, configured to receive first large language model parameters sent by a second local terminal in the federated learning system; A determination unit, configured to determine a teacher model according to the first large language model parameters and the local model; A prediction unit, configured to predict the first training sample data based on the teacher model to obtain prediction probability data output by the teacher model; A training unit, configured to train the local model according to the first training sample data and the prediction probability data to obtain a student model; A sending unit, configured to send second large language model parameters of the student model to a third local terminal so that the third local terminal updates the parameters of the local model based on the second large language model parameters; and send the second large language model parameters of the student model to the federated central server so that the federated central server updates the parameters of the federated large language model based on the large language model parameters returned by each local terminal.

15. A storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the training method of the large language model according to any one of claims 1 to 13.

16. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the training method of the large language model according to any one of claims 1 to 13.

17. A computer program product, which includes a computer program. The computer program is read and executed by a processor of a computer device, so that the computer device executes the training method of the large language model according to any one of claims 1 to 13.

Citation Information

Cited By

  • Non-data-driven quantum federal learning method based on single communication

    CN121390213A