Model training method, text processing method and related equipment
By pre-assigning word units to the expert network with the probability and load threshold in the large language model, and selecting the appropriate expert network to process word units, the problem of expert network overload is solved, and the training and text processing effects are improved.
Patent Information
- Application Number
- CN202510279888.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-09-23
AI Technical Summary
During the training process of large language models, each expert network may be overloaded, affecting the training effect and text processing efficiency.
By determining the probability of word units being pre-assigned to multiple expert networks and combining the load and load threshold of the expert network, an appropriate expert network is selected for processing to avoid overload.
Ensure that the expert networks of large language models are not overloaded during training and text processing, thereby improving the accuracy and efficiency of training results and text processing.
Smart Images

Figure CN120687816A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and specifically to a model training method, a text processing method, and related equipment. Background Art
[0002] Large Language Models (LLMs) are a type of natural language processing model with a large number of parameters built using deep learning techniques. These models are trained on large text datasets to understand and generate natural language, enabling them to perform a variety of language tasks such as text generation, translation, and question answering.
[0003] However, in the training process of a large language model, a large amount of training data is usually used to train all the expert networks in the large language model, which leads to the problem of overloading of each expert network in the large language model, thereby affecting the training effect of the large language model. Summary of the Invention
[0004] The present application provides a model training method, a text processing method, and related equipment to solve the technical problem of expert network overload during the training of a large language model.
[0005] A first aspect of an embodiment of the present application provides a model training method, comprising: calling a large language model to process a training sample, determining the probability that a word in the training sample is pre-assigned to multiple expert networks of the large language model for processing; determining the loads of the multiple expert networks based on multiple probabilities corresponding to the word; determining a first expert network and a second expert network from the multiple expert networks based on the loads and load thresholds of the multiple expert networks, the second expert network being an expert network having a load greater than the corresponding load threshold; determining a third expert network to process the word based on the probability that the word is assigned to the second expert network; and training the large language model based on the output results of the first expert network for the word and the output results of the third expert network for the word.
[0006] A second aspect of an embodiment of the present application provides a text processing method, the method comprising: inputting text into a large language model to obtain a processing result of the text, wherein the large language model is obtained by the model training method described in the first aspect.
[0007] According to a third aspect of an embodiment of the present application, a model training device is provided, comprising: a determination unit for calling a large language model to process a training sample and determining the probability that a word in the training sample is pre-assigned to multiple expert networks of the large language model for processing; the determination unit is further configured to determine the loads of the multiple expert networks based on multiple probabilities corresponding to the word; the determination unit is further configured to determine a first expert network and a second expert network from the multiple expert networks based on the loads and load thresholds of the multiple expert networks, the second expert network being an expert network having a load greater than the corresponding load threshold; the determination unit is further configured to determine a third expert network to process the word based on the probability that the word is assigned to the second expert network; and a training unit is configured to train the large language model based on the output results of the first expert network for the word and the output results of the third expert network for the word.
[0008] A fourth aspect of an embodiment of the present application provides a text processing device, comprising: an input module for inputting text into a large language model to obtain a processing result of the text, wherein the large language model is obtained by the model training method described in the first aspect.
[0009] A fifth aspect of an embodiment of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method provided in the first or second aspect above when executing the computer program.
[0010] A sixth aspect of an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the method provided in the first or second aspect are implemented.
[0011] A seventh aspect of the embodiments of the present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the steps in the method provided in the first or second aspect above.
[0012] In the model training method of this embodiment, by determining the probability of a word being pre-assigned to the multiple expert networks for processing, the load of the multiple expert networks can be estimated. Furthermore, combining the load and load threshold of the multiple expert networks, a second expert network can be determined. The probability of a word being assigned to the second expert network can be used to determine the third expert network to process the word, thereby avoiding overloading the second expert network. By training a large language model using the output of the first expert network for the word and the output of the third expert network for the word, the large language model can be trained while ensuring the training effect of the large language model while avoiding overloading of the individual expert networks within the large language model.
[0013] In the text processing method of this embodiment, the text is processed by a large language model. Since the large language model is obtained through load training of multiple expert networks, it can ensure that the expert networks will not be overloaded when the large language model processes the text, thereby improving the processing effect of the large language model on the text, thereby improving the accuracy of the processing results and the determination efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0015] Figure 1 This is a schematic diagram of a large language model processing word units; Figure 2 This is a schematic diagram of an application scenario of a model training method and a text processing method provided in an embodiment of the present application; Figure 3 This is a flow chart of a model training method provided in an embodiment of the present application; Figure 4 This is a schematic diagram of the structure of the large language model provided in the embodiment of the present application; Figure 5 This is a detailed flowchart of training a large language model provided in an embodiment of the present application; Figure 6 This is a flowchart of a text processing method provided by an embodiment of the present application; Figure 7 This is another structural diagram of the large language model provided in the embodiment of the present application; Figure 8 This is another structural diagram of the large language model provided in the embodiment of the present application; Figure 9 This is a functional module diagram of a model training device provided in an embodiment of the present application; Figure 10 This is a functional module diagram of a text processing device provided by an embodiment of the present application; Figure 11 It is a structural diagram of an electronic device for implementing a model training method and a text processing method provided in an embodiment of the present application. DETAILED DESCRIPTION
[0016] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0017] In the following, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or design schemes. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete way.
[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs. The terms used in the specification of this application are for the purpose of describing specific embodiments only and are not intended to limit this application. It should be understood that, unless otherwise specified in this application, " / " means or. For example, A / B can mean A or B. "And / or" in this application is merely a way to describe the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. "At least one" means one or more. "Multiple" means two or more than two. For example, at least one of a, b or c can mean: a, b, c, a and b, a and c, b and c, a, b and c.
[0019] Some terminology explanations: Large Language Models (LLMs) are an AI technology that uses deep learning models to learn linguistic rules and semantic information from large amounts of text data to generate natural language text. LLMs typically involve large neural network models with billions or even tens of billions of parameters, hence the name "large language model." LLMs can be used to handle a variety of natural language tasks, such as text classification, intelligent question answering, and text extraction.
[0020] The basic idea of the Mixture-of-Expert (MOE) model is to train multiple expert networks, each responsible for processing only a portion of the input data. A MOE model typically includes a gating network and N expert networks, which generally have the same network structure, with N being an integer greater than 1. Given an input x, the MOE model first selects several expert networks for that input through the gating network. The input is then assigned to the selected expert networks for separate computations, and the results of these expert networks are aggregated to produce the final result for that input.
[0021] Dynamic Routing: When the load of any expert network in the MOE model is too high, the dynamic routing mechanism will probabilistically input the word unit to the expert network with the next highest priority. If the expert network with the next highest priority is also overloaded, the dynamic routing mechanism will continue to probabilistically input the word unit to the expert network with the lower priority until an expert network without overload is found.
[0022] Regular Gate Logic Control (RGLC): Conventional MOE logic gate control structures tend to produce uniformly distributed expert networks, resulting in nearly identical functionality and making it difficult for these networks to truly become specialized, specialized experts in their respective fields. By introducing regularization and random dropout neurons to optimize the expert networks, the discriminability of each expert network is improved.
[0023] Supervised Fine-Tuning (SFT): refers to the process of performing additional training on a pre-trained model on a specific task to improve its performance.
[0024] Low-Rank Adaptation (LoRA): This technique is used to fine-tune large language models by training a low-rank matrix and then injecting the parameters into the original model. This approach reduces computational requirements, making training resources much smaller than directly training the original model, making it suitable for use in resource-constrained environments.
[0025] like Figure 1 As shown, Figure 1 This is a diagram of a large language model processing word units. Figure 1 In the illustrated solution, the text can be divided into N word-unit tokens, namely Token1, Token2, ..., TokenN. The gating network in the MOE model determines the expert network to process each word-unit token. For example, the expert network to process Token1, Token2, and Token5 is Expert3. However, because Expert3 is under high load when processing Token5, it discards Token5. This approach results in the loss of word-unit tokens, which is detrimental to model training and text processing performance.
[0026] In addition, in order to ensure the effectiveness of model training and the model's processing of text, related technologies usually set a larger load threshold to reduce the overload of the expert network to a certain extent. However, this method will reduce the efficiency of model training and the model's processing of text.
[0027] Based on the above problems, in order to solve the problem of expert network overload during the training process of a large language model while ensuring the training effect of the model and the model's processing effect on text, an embodiment of the present application provides a model training method that can solve the problem of overload of each expert network during the training process of a large language model while ensuring the training effect of the large language model. In addition, an embodiment of the present application also provides a text processing method that can ensure that each expert network will not be overloaded when the large language model processes text, thereby improving the processing effect of the large language model on text, thereby improving the accuracy of the processing results and the determination efficiency.
[0028] See also Figure 2 , Figure 2 Schematic diagram of an application scenario of a model training method and a text processing method provided in an embodiment of the present application. The scenario may include various electronic devices 100 and a server 200.
[0029] The electronic device 100 can be a mobile phone, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a cellular phone, a personal digital assistant (PDA), an augmented reality (AR) device, a virtual reality (VR) device, an artificial intelligence (AI) device, a wearable device, an in-vehicle device, a smart home device and / or a smart city device. The embodiments of the present application do not impose any special restrictions on the specific type of the electronic device 100.
[0030] Server 200 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, i.e., Content Delivery Network (CDN), as well as big data and artificial intelligence platforms, but is not limited to these.
[0031] It should be noted that the method in the embodiment of the present application can be performed independently by the electronic device 100 or the server 200, or can be performed jointly by the server 200 and the electronic device 100. When performed independently by the electronic device 100 or the server 200, the model training and application processes can be implemented independently by the electronic device 100 or the server 200. For example, the large language model can be fine-tuned on the electronic device 100 to obtain a trained large language model. Accordingly, after training, the electronic device 100 can use the trained large language model to recognize text. The above process can also be performed independently by the server 200. When performed jointly by the server 200 and the electronic device 100, the server 200 can train the large language model and then deploy the trained large language model to the electronic device 100, and the electronic device 100 can implement the text processing process. Alternatively, part of the model training or application process can be implemented by the electronic device 100, and part of the process can be implemented by the server 200, and the two can cooperate to implement the model training or application process. In actual application, specific configuration can be made according to the situation and is not specifically limited here.
[0032] It should be noted that when the model training method and text processing method provided in the embodiment of the present application are executed separately by the server 200 or the electronic device 100, the above-mentioned application scenario may also only include a single device of the server 200 or the electronic device 100, or the server 200 and the electronic device 100 may be considered to be the same device. In actual application, when the model training method and text processing method provided in the embodiment of the present application are jointly executed by the server 200 and the electronic device 100, the server 200 and the electronic device 100 may also be the same device, that is, the server 200 and the electronic device 100 may be different functional modules of the same device, or virtual devices virtualized from the same physical device.
[0033] In a possible implementation, the user may provide text through the electronic device 100, and the server 200 may use the text processing method of the embodiment of the present application to determine the processing result corresponding to the text and return it to the electronic device 100 for presentation.
[0034] In the embodiment of the present application, the electronic device 100 and the server 200 can be directly or indirectly connected to each other through one or more networks. The network can be a wired network or a wireless network. For example, the wireless network can be a mobile cellular network or a Wireless Fidelity (WIFI) network. Of course, it can also be other possible networks, and the embodiment of the present application does not limit this. It should be noted that Figure 1 The examples shown are just for illustration. In fact, the number of terminal devices and servers is not limited and is not specifically limited in the embodiments of this application.
[0035] The model training method and text processing method provided by the embodiments of the present application are described below in combination with the application scenarios described above and with reference to the accompanying drawings. It should be noted that the above application scenarios are only shown to facilitate understanding of the spirit and principles of the embodiments of the present application, and the embodiments of the present application are not limited in this respect.
[0036] like Figure 3 FIG. 1 is a flow chart of a model training method provided by an embodiment of the present application. The model training method is applied to electronic devices, for example, Figure 2 The electronic device 100. According to different requirements, the order of the steps in the flowchart can be changed, and some steps can be omitted.
[0037] S301 , calling a large language model to process a training sample, and determining the probability that a word in the training sample is pre-assigned to a plurality of expert networks of the large language model for processing.
[0038] In at least one embodiment of the present application, the training sample can be text data in any field, for example, text data in the financial services field, or text data in the healthcare field. The training sample includes multiple word units, which can be words in the text data, but is not limited to this in actual applications.
[0039] The large language model includes a gated network and multiple expert networks. Multiple expert networks can have the same network structure, and the specific processing processes of each expert network can be the same. Figure 4 As shown in FIG, it is a schematic diagram of a large language model provided in an embodiment of the present application. Figure 4 As shown in Figure 1, the large language model includes a gating network and 8 expert networks.
[0040] In at least one embodiment of the present application, a gating network can be used to determine the assignment information of each word in a training sample to each expert network. Specifically, the electronic device can obtain the encoding vector of the word in the training sample and then, based on the encoding vector of the word, invoke the gating network to perform expert network assignment processing on the word, thereby obtaining the probability of the word being pre-assigned to multiple expert networks for processing.
[0041] The encoding vector of a word unit can be a vector output by the embedding layer in the large language model after processing the word unit. For example, the electronic device inputs multiple word units in the training sample into the embedding layer in the large language model, and encodes each word unit through the embedding layer to obtain the encoding vector of each word unit. The encoding vector of a word unit can also be obtained through the embedding layer and attention layer of the large language model. For example, the electronic device inputs multiple word units in the training sample into the embedding layer in the large language model, and encodes each word unit through the embedding layer to obtain the embedding representation of each word unit, and inputs the embedding representation of each word unit into the attention layer, and performs attention operation on the embedding representation of each word unit through the attention layer to obtain the encoding vector of each word unit. The embedding layer can include a word embedding layer and a position embedding layer.
[0042] The gating network may include, but is not limited to, a normalization layer, a fully connected layer, and an activation layer. In one example, the electronic device may operate on the encoding vector of each word unit using the activation function of the activation layer in the gating network to obtain the probability that the word unit is pre-assigned to multiple expert networks for processing. In another example, the electronic device may normalize the encoding vector of the word unit using the normalization layer of the gating network to obtain a first vector, where the first vector can be expressed as: ,in, Can indicate the The first vector of the word, Can indicate the The encoding vector of word units, It can represent the mean vector of the encoding vectors of all word units in the training sample. The mean vector can be the mean of the corresponding elements in multiple encoding vectors. For example, if one encoding vector is (0, 1, 0) and the other encoding vector is (1, 1, 1), then the mean vector can be (0.5, 1, 0.5). It can represent the standard deviation vector of the encoding vectors of all word units in the training sample. The standard deviation vector can be the standard deviation of the corresponding elements in multiple encoding vectors. The first vector is integrated by the fully connected layer in the gating network to obtain the second vector. For example, the second vector can be expressed as The second vector is processed by the activation function in the activation layer to obtain the probability that the word unit is pre-assigned to multiple expert networks for processing. For example, the probability that the word unit is pre-assigned to the expert network for processing can be expressed as In addition, the gating network can output the probability that the word is pre-assigned to K expert networks for processing, for example, , where K can be less than or equal to the total number of expert networks in the large language model. This embodiment can improve the computational efficiency of the load by outputting the probability that a word is pre-assigned to K expert networks for processing.
[0043] This embodiment analyzes the encoding vector of a word unit through a gated network, and can quickly obtain the probability that the word unit is pre-assigned to multiple expert networks for processing.
[0044] S302 : Determine the loads of multiple expert networks based on multiple probabilities corresponding to word units.
[0045] In at least one embodiment of the present application, after determining the probability that a word-gram is pre-assigned to multiple expert networks for processing, the load of each expert network may be determined based on the multiple probabilities corresponding to the word-gram.
[0046] In at least one embodiment of the present application, the electronic device assigns a word to the expert network corresponding to the highest probability based on multiple probabilities corresponding to the word. For example, for word Token1, the probabilities pre-assigned to expert networks Expert1 to Expert8 for processing are 0.1, 0.1, 0.3, 0.1, 0.1, 0.1, 0.1, and 0.1, respectively. Since the probability of word Token1 being pre-assigned to expert network Expert3 is the highest, expert network Expert3 is determined to be the expert network corresponding to word Token1.
[0047] The electronic device determines the load corresponding to each expert network based on the word elements assigned to each expert network. Specifically, for any expert network, the load of any expert network can be determined based on the number of word elements assigned to any expert network. For example Figure 4As shown, it is determined that the word units allocated to the expert network Expert3 include Token1, Token2 and Token5, and the load of the expert network Expert3 can be determined to be 3.
[0048] In this embodiment, the expert network to which the word-unit is assigned can be determined by using multiple probabilities corresponding to the word-unit. The load corresponding to each expert network can be further quantified by using the word-units assigned to each expert network.
[0049] S303 : Determine a first expert network and a second expert network from the multiple expert networks according to the loads and load thresholds of the multiple expert networks.
[0050] In at least one embodiment of the present application, each expert network corresponds to a load threshold, and the load threshold can be used to detect whether the corresponding expert network is overloaded. Specifically, the electronic device determines an expert network with a load less than or equal to the corresponding load threshold as a first expert network, and determines an expert network with a load greater than the corresponding load threshold as a second expert network.
[0051] S304 : Determine a third expert network for processing the word-unit according to the probability of the word-unit being assigned to the second expert network.
[0052] In at least one embodiment of the present application, each expert network corresponds to a probability threshold, and the importance of the expert network to the word-unit can be detected by the probability threshold. The electronic device determines the third expert network to process the word-unit based on the probability of the word-unit being assigned to the second expert network and the probability threshold corresponding to the second expert network. Specifically, the method for determining the third expert network can be found in Figure 5 The process shown.
[0053] S3041 , determining whether the probability of the word being assigned to the second expert network is less than a probability threshold corresponding to the second expert network.
[0054] In at least one embodiment of the present application, if the probability of the word being assigned to the second expert network is greater than or equal to the probability threshold corresponding to the second expert network, step S3042 is executed; if the probability of the word being assigned to the second expert network is less than the probability threshold corresponding to the second expert network, step S3043 is executed.
[0055] S3042: Determine the target network layer in the second expert network as the third expert network.
[0056] In at least one embodiment of the present application, if the probability of a word being assigned to the second expert network is greater than or equal to a probability threshold corresponding to the second expert network, it may indicate that the second expert network is of greater importance to the word. To reduce the load on the second expert network, the electronic device may determine the target network layer in the second expert network as the third expert network.
[0057] In at least one embodiment of the present application, when determining the target network layer, the electronic device may calculate the score of the second expert network based on the load of the second expert network, the probability of word units being assigned to the second expert network, and the load threshold and probability threshold corresponding to the second expert network. The score calculation formula may be expressed as: ,in, can represent the score of the second expert network, can represent the load of the second expert network, It can represent the load threshold corresponding to the second expert network, It can represent the probability of word unit assigned to the second expert network, It can represent the probability threshold corresponding to the second expert network, The weights can be set and adjusted according to needs.
[0058] The electronic device selects a target network layer from the second expert network based on the score. Specifically, the electronic device can select all network layers between the i-th network layer and the N-th network layer from the second expert network as target network layers based on the score of the second expert network. Where N is the total number of layers in the second expert network, and i can be determined based on the score of the second expert network and the total number of layers in the second expert network. For example, , and for example, .
[0059] In this embodiment, when the load of the second expert network is greater than the corresponding load threshold and the probability of the word being assigned to the second expert network is greater than or equal to the corresponding probability threshold, it can be indicated that the second expert network is overloaded, and the second expert network is more important for the word. Therefore, by inputting the encoding vector of the word into the target network layer in the second expert network, the load of the second expert network can be reduced, and the processing of the word by the second expert network can also be realized.
[0060] S3043: Determine a third expert network from the first expert network.
[0061] In at least one embodiment of the present application, if the probability of a word-unit being assigned to the second expert network is less than a probability threshold corresponding to the second expert network, it may indicate that the second expert network is less important to the word-unit. The electronic device may select any expert network from the first expert network as the third expert network.
[0062] In this embodiment, when the load of the second expert network is greater than the corresponding load threshold and the probability of a word being assigned to the second expert network is less than the corresponding probability threshold, it can be indicated that the second expert network is overloaded, and the second expert network is less important to the word. Therefore, selecting any expert network from the first expert network as the third expert network can reduce the load of the second expert network.
[0063] In at least one embodiment of the present application, to reduce the load on a single graphics card and improve processing efficiency for large language models, multiple expert networks for a large language model are typically stored on different graphics cards. To further improve word unit processing efficiency, an electronic device may use a first expert network stored in the same storage location (e.g., graphics card) as a second expert network as a third expert network.
[0064] like Figure 4 As shown, the large language model includes expert networks Expert1-Expert8, wherein expert network Expert6 and expert network Expert3 are in the same storage location, and expert network Expert4 and expert network Expert5 are in different storage locations from expert network Expert3. Based on the load and load threshold of the multiple expert networks, the electronic device determines from the multiple expert networks that the first expert network includes expert network Expert4, expert network Expert5, and expert network Expert6, and determines from the multiple expert networks that the second expert network includes expert network Expert1, expert network Expert2, expert network Expert3, expert network Expert7, and expert network Expert8. For the second expert network Expert3, since the probability of 0.5 assigned to the second expert network Expert3 by the word unit Token5 is less than the probability threshold of 0.65 corresponding to the second expert network Expert3, the electronic device can determine the first expert network Expert6 in the same storage location as the expert network Expert3 as the third expert network.
[0065] By selecting the third expert network from the first expert network, which is stored in the same location as the second expert network, the above embodiment ensures that word units are forwarded to the expert network in the same location for processing, thus avoiding the inefficiencies caused by cross-device communication. Furthermore, by selecting the third expert network from the first expert network, the third expert network is prevented from being overloaded.
[0066] In another embodiment, if there are multiple first expert networks in the same storage location as the second expert network, the electronic device selects the expert network corresponding to the maximum probability as the third expert network based on the probability of word units being assigned to the expert networks. For example, the load threshold of expert network Expert5 is 2, and the load of expert network Expert5 is 3, indicating that expert network Expert5 is overloaded. Since the probability of word unit Token7 being assigned to expert network Expert5, 0.5, is less than the probability threshold of 0.75 corresponding to expert network Expert5, the electronic device determines that the first expert networks in the same storage location as expert network Expert5 include: expert network Expert4 and expert network Expert6. Since the probability of word unit Token7 being assigned to expert network Expert4 is 0.1 and the probability of word unit Token7 being assigned to expert network Expert6 is 0.2, the electronic device can use expert network Expert6 as the third expert network.
[0067] In addition, in other embodiments, if there are multiple first expert networks in the same storage location as the second expert network, and the maximum probability of word units being assigned to the expert networks is two or more, one of the two or more expert networks corresponding to the same maximum probability can be selected as the third expert network.
[0068] In this embodiment, when there are multiple first expert networks in the same storage location as the second expert network, the probability of assigning word units to expert networks is used to select an expert network that is more important to the word unit as the third expert network, thereby ensuring the processing effect of the third expert network on the word unit.
[0069] S305 , training a large language model based on the output results of the first expert network on the word-unit and the output results of the third expert network on the word-unit.
[0070] In at least one embodiment of the present application, the expert network in the large language model can be a feedforward neural network. For example, the expert network in the large language model can include a first fully connected layer, an activation layer, and a second fully connected layer. The expert network in the large language model can also be a Transformer network. The expert network in the large language model can also include other network structures. For example, the expert network in the large language model can include multiple hidden layers.
[0071] In at least one embodiment of the present application, the electronic device inputs the encoding vector of the word unit assigned to the first expert network into the first expert network to obtain the output result of the first expert network.
[0072] In at least one embodiment of the present application, the electronic device inputs the encoding vector of the word unit assigned to the third expert network into the third expert network to obtain the output result of the third expert network. In one example, the third expert network includes the i-th network layer to the N-th network layer of the second expert network. For example, if the second expert network includes a first fully connected layer, an activation layer, and a second fully connected layer, and i is 2, the encoding vector of the word unit can be input into the activation layer to obtain a third vector, and the third vector can be input into the second fully connected layer to obtain the output result of the third expert network.
[0073] In at least one embodiment of the present application, the network parameters of the large language model include but are not limited to: a load threshold of each expert network and a probability threshold of each expert network. The electronic device can adjust the network parameters of the large language model to optimize the large language model.
[0074] In at least one embodiment of the present application, the training sample may be any text, the annotation result may be the semantic information of the training sample, the annotation result may also be the reply sentence of the training sample, the annotation result may also be the keywords in the training sample, etc.
[0075] In at least one embodiment of the present application, when training a large language model, the electronic device may determine a loss value for the large language model based on the output of the first expert network for the word unit, the output of the third expert network for the word unit, and the labeling results of the training samples. Based on the loss value, the electronic device adjusts the network parameters of the large language model until the loss value meets a preset condition. The preset condition may be set as: the loss value is less than or equal to a preset loss value, the loss value no longer decreases, the number of training times reaches a preset number, etc.
[0076] Specifically, the large language model also includes an output layer, through which the electronic device analyzes the output results of the first expert network and the output results of the third expert network to obtain prediction information of the large language model. Based on the prediction information and the annotation results, the electronic device calculates a loss value for the large language model. The loss value can be determined based on a function such as a cross-entropy loss function or a Hinge loss function.
[0077] In at least one embodiment of the present application, when adjusting the network parameters of a large language model, the electronic device can freeze the hyperparameters of other networks in the large language model, except for the gated network. This embodiment can quickly adjust the hyperparameters of the gated network by freezing the hyperparameters of other networks.
[0078] In the model training method of this embodiment, by determining the probability of word units being pre-assigned to multiple expert networks for processing, the load of the multiple expert networks can be estimated. Furthermore, combining the load and load threshold of the multiple expert networks, a second expert network can be determined. The probability of word units being assigned to the second expert network can be used to determine the third expert network that processes the word units, thereby avoiding overloading the second expert network. By training a large language model using the output of the first expert network for the word units and the output of the third expert network for the word units, the large language model can be trained while ensuring the training effect of the large language model while avoiding overloading of the individual expert networks within the large language model.
[0079] like Figure 6 FIG. 1 is a flowchart of a text processing method provided by an embodiment of the present application. The text processing method is applied in electronic devices, for example, Figure 2 The electronic device 100. According to different requirements, the order of the steps in the flowchart can be changed, and some steps can be omitted.
[0080] S601: Input the text into the large language model to obtain the processing result of the text.
[0081] In at least one embodiment of the present application, a large language model can be used to determine the processing result of the text, referring to Figure 7 As shown, Figure 7 This is a schematic diagram of the structure of the large language model provided in the embodiment of this application. Figure 7 As shown, the large language model includes an embedding layer, M attention layers, and an output layer. Each attention layer includes a multi-head attention layer, a normalization layer, a gating network, and multiple expert networks. M can be a positive integer greater than or equal to 1. Among them, the embedding layer can be used to encode multiple words in the text, the multi-head attention layer can be used to perform aggregation operations on the encoding vector of each word, the normalization layer can be used to standardize the output of the multi-head attention layer, the gating network can be used to determine the probability of words in the text being assigned to multiple expert networks for processing. Multiple expert networks can have the same network structure. The output layer can determine the processing result of the text based on the output result of each expert network.
[0082] Assume that M is 1. The network structure of the large language model consists of an embedding layer, a multi-head attention layer, a normalization layer, a gating network, multiple expert networks, and an output layer. The electronic device encodes each word in the text using the embedding layer in the large language model, obtaining an embedded representation for each word. The multi-head attention layer and the normalization layer perform an attention operation on the embedded representation of each word, obtaining an encoding vector for each word.
[0083] The electronic device normalizes the encoding vector of each word through the normalization layer in the gated network, performs full-connection processing on the information output by the normalization layer in the gated network through the fully-connected layer in the gated network, and operates on the information output by the fully-connected layer in the gated network through the activation function in the gated network to obtain the probability of each word being pre-assigned to multiple expert networks for processing.
[0084] The electronic device determines the load of each expert network based on the probability of each word in the text being pre-assigned to multiple expert networks for processing. Each expert network corresponds to a load threshold and a probability threshold. The electronic device determines an expert network with a load less than or equal to the corresponding load threshold as a first expert network, and determines an expert network with a load greater than the corresponding load threshold as a second expert network.
[0085] The electronic device determines whether the probability of the word being assigned to the second expert network is less than the probability threshold corresponding to the second expert network. If the probability of the word being assigned to the second expert network is greater than or equal to the probability threshold corresponding to the second expert network, the electronic device determines the target network layer in the second expert network to be the third expert network, where the target network layer may be the i-th network layer in the second expert network, i is a positive integer greater than 1 and less than N, and N is the total number of layers in the second expert network. If the probability of the word being assigned to the second expert network is less than the probability threshold corresponding to the second expert network, the electronic device determines a third expert network from the first expert network, where the third expert network may be the first expert network in the same storage location as the second expert network, the third expert network may also be the first expert network in the same storage location as the second expert network and having the highest probability, or another expert network other than the second expert network.
[0086] For example Figure 4As shown, the large language model includes expert networks Expert1-Expert8, wherein expert network Expert6 and expert network Expert3 are in the same storage location, and expert network Expert4 and expert network Expert5 are in different storage locations from expert network Expert3. Based on the load and load threshold of the multiple expert networks, the electronic device determines from the multiple expert networks that the first expert network includes expert network Expert4, expert network Expert5, and expert network Expert6, and determines from the multiple expert networks that the second expert network includes expert network Expert1, expert network Expert2, expert network Expert3, expert network Expert7, and expert network Expert8. For the second expert network Expert3, since the probability of 0.5 assigned to the second expert network Expert3 by the word unit Token5 is less than the probability threshold of 0.65 corresponding to the second expert network Expert3, the electronic device can determine the first expert network Expert6 in the same storage location as the expert network Expert3 as the third expert network.
[0087] Another example Figure 8 As shown, the load threshold of expert network Expert3 is 2, and the load of expert network Expert3 is 3, indicating that expert network Expert3 is overloaded. Since the probability of vocabulary Token5 assigned to expert network Expert3 is 0.25, which is less than the probability threshold of 0.65 corresponding to expert network Expert3, the electronic device determines that the probability of vocabulary Token5 being assigned to expert network Expert1, expert network Expert2, expert network Expert4, expert network Expert5, expert network Expert7, and expert network Expert8 is 0.1, and the probability of vocabulary Token5 being assigned to expert network Expert6 is 0.15. The electronic device determines that the expert network with the highest probability except expert network Expert3 is expert network Expert6, and thus determines expert network Expert6 as the third expert network.
[0088] The electronic device transforms the output results of the vocabulary through the output layer of the large language model to obtain a text processing result. The processing result can be semantic information of the text, a response sentence of the text, or a key word in the text, etc. This application does not limit the examples of the processing result.
[0089] In the text processing method of this embodiment, the text is processed by a large language model. Since the large language model is obtained through load training of multiple expert networks, it can ensure that the expert networks will not be overloaded when the large language model processes the text, thereby improving the processing effect of the large language model on the text, thereby improving the accuracy of the processing results and the determination efficiency.
[0090] like Figure 9 , is a functional module diagram of a model training device provided by an embodiment of the present application. The model training device 91 includes a determination unit 910, a training unit 911, a calculation unit 912 and a selection unit 913. The module / unit referred to in this application refers to a unit that can be processed by a processor (e.g. Figure 11 The processor 1101 shown in FIG. 1 is obtained and is capable of performing a series of computer-readable instruction segments that are stored in a memory (eg, Figure 11 1102).
[0091] Determination unit 910 is used to call the large language model to process training samples and determine the probability that the word units in the training samples are pre-assigned to multiple expert networks of the large language model for processing; determination unit 910 is also used to determine the loads of multiple expert networks based on multiple probabilities corresponding to the word units; determination unit 910 is also used to determine a first expert network and a second expert network from multiple expert networks based on the loads and load thresholds of the multiple expert networks, where the second expert network is an expert network with a load greater than the corresponding load threshold; determination unit 910 is also used to determine a third expert network to process the word unit based on the probability that the word unit is assigned to the second expert network; training unit 911 is used to train the large language model based on the output results of the first expert network for the word unit and the output results of the third expert network for the word unit.
[0092] In one embodiment, the large language model also includes a gating network, which includes a normalization layer, a fully connected layer, and an activation layer. The determination unit 910 is specifically used to: use the normalization layer to normalize the encoding vector of the word unit to obtain a first vector; based on the fully connected layer, perform feature integration on the first vector to obtain a second vector; based on the second vector, use the activation function in the activation layer to calculate the probability that the word unit is pre-assigned to multiple expert networks for processing.
[0093] In one embodiment, the determining unit 910 is specifically configured to: assign a word-gram to an expert network corresponding to the maximum probability according to multiple probabilities corresponding to the word-gram; and determine a load corresponding to each expert network according to the word-gram assigned to each expert network.
[0094] In one embodiment, the determination unit 910 is specifically configured to: if the probability of the word being assigned to the second expert network is greater than or equal to the probability threshold corresponding to the second expert network, determine the target network layer in the second expert network as the third expert network; if the probability of the word being assigned to the second expert network is less than the probability threshold corresponding to the second expert network, determine the third expert network from the first expert network.
[0095] In one embodiment, the calculation unit 912 is used to calculate the score of the second expert network based on the load of the second expert network, the probability of the word unit being assigned to the second expert network, and the load threshold and probability threshold corresponding to the second expert network; the selection unit 913 is used to select the target network layer from the second expert network based on the score.
[0096] In one embodiment, the determining unit 910 is specifically configured to determine a first expert network stored in the same storage location as the second expert network as the third expert network.
[0097] In one embodiment, the training unit 911 is specifically used to: determine the loss value of the large language model based on the output results of the first expert network on the word unit, the output results of the third expert network on the word unit, and the labeling results of the training samples; and adjust the network parameters of the large language model based on the loss value, the network parameters including the load threshold and probability threshold of each expert network.
[0098] In various embodiments of the present application, by determining the probability of a word being pre-assigned to multiple expert networks for processing, the load of the multiple expert networks can be estimated. Furthermore, combining the load and load threshold of the multiple expert networks, a second expert network can be determined. The probability of a word being assigned to the second expert network can be used to determine the third expert network that processes the word, thereby avoiding overloading the second expert network. By training a large language model using the output of the first expert network for the word and the output of the third expert network for the word, it is possible to avoid overloading the individual expert networks within the large language model while ensuring the training effectiveness of the large language model.
[0099] like Figure 10 , is a functional module diagram of a text processing device provided by an embodiment of the present application. The text processing device 101 includes an input module 1010. The module / unit referred to in this application refers to a unit that can be processed by a processor (e.g. Figure 11 The processor 1101 shown in FIG. 1 is obtained and is capable of performing a series of computer-readable instruction segments that are stored in a memory (eg, Figure 11 1102).
[0100] Input module 1010 is used to input text into the large language model to obtain the processing result of the text. The large language model is Figure 3The model training method is obtained.
[0101] In multiple embodiments of the present application, text is processed by a large language model. Since the large language model is obtained through load training of multiple expert networks, it can ensure that the various expert networks will not be overloaded when the large language model processes the text, thereby improving the processing effect of the large language model on the text, thereby improving the accuracy of the processing results and the determination efficiency.
[0102] Figure 11 Schematic diagram of the structure of an electronic device for implementing the model training method and text processing method provided in the embodiment of the present application. Figure 11 The electronic device 100 is used to perform Figure 3 、 Figure 5 and Figure 6 The method shown.
[0103] The electronic device 100 includes at least one processor 1101 , a memory 1102 , and at least one network interface 1103 .
[0104] The processor 1101 is, for example, a general-purpose central processing unit (CPU), a network processor (NP), a graphics processing unit (GPU), a neural-network processing unit (NPU), a data processing unit (DPU), a microprocessor, or one or more integrated circuits for implementing the solution of the present application. For example, the processor 1101 includes an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The PLD is, for example, a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0105] The memory 1102 may be, for example, a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, a random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, an optical disc storage (including a compact disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto. Optionally, the memory 1102 exists independently and is connected to the processor 1101 via the internal connection 1104. Alternatively, the memory 1102 and the processor 1101 may be integrated together.
[0106] The network interface 1103 uses any transceiver-like device for communicating with other devices or communication networks. For example, the network interface 1103 includes at least one of a wired network interface and a wireless network interface. For example, the wired network interface is an Ethernet interface. For example, the Ethernet interface is an optical interface, an electrical interface, or a combination thereof. For example, the wireless network interface is a wireless local area network (WLAN) interface, a cellular network interface, or a combination thereof.
[0107] In some embodiments, the processor 1101 includes one or more CPUs, such as Figure 11 CPU0 and CPU1 are shown in the figure.
[0108] In some embodiments, the electronic device 100 optionally includes multiple processors, such as Figure 11 1 and 1105. Each of these processors is, for example, a single-CPU or a multi-CPU. A processor herein optionally refers to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).
[0109] In some embodiments, electronic device 100 further includes internal connections 1104. Processor 1101, memory 1102, and at least one network interface 1103 are connected via internal connections 1104. Internal connections 1104 include pathways for transmitting information between these components. Internal connections 1104 may optionally be a single board or bus. Internal connections 1104 may optionally be divided into an address bus, a data bus, a control bus, and the like.
[0110] In some embodiments, the electronic device 100 further includes an input / output interface 1106 , which is connected to the internal connection 1104 .
[0111] Optionally, the processor 1101 implements the method in the above embodiment by reading the program code 1107 stored in the memory 1102, or the processor 1101 implements the method in the above embodiment by internally stored program code. In the case where the processor 1101 implements the method in the above embodiment by reading the program code 1107 stored in the memory 1102, the memory 1102 stores the program code that implements the method provided in the embodiment of the present application.
[0112] For more details on how the processor 1101 implements the above functions, please refer to the descriptions in the previous method embodiments, which will not be repeated here.
[0113] This embodiment also provides a computer storage medium, which stores computer instructions. When the computer instructions are executed on an electronic device, the electronic device executes the above-mentioned related method steps to implement the model training method and text processing method in the above-mentioned embodiment.
[0114] This embodiment also provides a computer program product. When the computer program product is run on an electronic device, the electronic device executes the above-mentioned related steps to implement the model training method and text processing method in the above-mentioned embodiment.
[0115] In addition, an embodiment of the present application also provides a device, which can specifically be a chip, component or module, and the device may include a connected processor and memory; wherein the memory is used to store computer-executable instructions, and when the device is running, the processor can execute the computer-executable instructions stored in the memory to enable the chip to execute the model training method and text processing method in the above-mentioned method embodiments.
[0116] Among them, the electronic device, computer storage medium, computer program product or chip provided in this embodiment is used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, and will not be repeated here.
[0117] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0118] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0119] The units described as separate components may or may not be physically separate, and the components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple places. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0120] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0121] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a device (which can be a single-chip microcomputer, chip, etc.) or a processor to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program code.
[0122] The above are only specific embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any changes or replacements within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A model training method, characterized in that: The method comprises: Invoking a large language model to process a training sample, and determining a probability that a word in the training sample is pre-assigned to a plurality of expert networks of the large language model for processing; determining the loads of the plurality of expert networks based on the plurality of probabilities corresponding to the word-grams; determining a first expert network and a second expert network from the plurality of expert networks according to the loads and load thresholds of the plurality of expert networks, wherein the second expert network is an expert network having a load greater than a corresponding load threshold; determining a third expert network for processing the word-unit according to the probability of the word-unit being assigned to the second expert network; The large language model is trained based on the output results of the first expert network on the word-unit and the output results of the third expert network on the word-unit.
2. The model training method according to claim 1, characterized in that The large language model further includes a gating network, which includes a normalization layer, a fully connected layer, and an activation layer. Determining the probability that a word in the training sample is pre-assigned to the multiple expert networks of the large language model for processing includes: Using the normalization layer to normalize the encoding vector of the word unit to obtain a first vector; Performing feature integration on the first vector based on the fully connected layer to obtain a second vector; Based on the second vector, an activation function in the activation layer is used to calculate the probability that the word unit is pre-assigned to the multiple expert networks for processing.
3. The model training method according to claim 1, characterized in that The determining the loads of the multiple expert networks based on the multiple probabilities corresponding to the word-grams includes: According to the multiple probabilities corresponding to the word-unit, the word-unit is assigned to the expert network corresponding to the maximum probability; The load corresponding to each expert network is determined according to the word units allocated to each expert network.
4. The model training method according to claim 1, characterized in that The step of determining a third expert network for processing the word-unit according to the probability of the word-unit being assigned to the second expert network includes: If the probability of the word being assigned to the second expert network is greater than or equal to the probability threshold corresponding to the second expert network, determining the target network layer in the second expert network to be the third expert network; If the probability of the word being assigned to the second expert network is less than a probability threshold corresponding to the second expert network, the third expert network is determined from the first expert network.
5. The model training method according to claim 4, characterized in that The method further comprises: Calculating a score of the second expert network according to the load of the second expert network, the probability of the word being assigned to the second expert network, and a load threshold and a probability threshold corresponding to the second expert network; The target network layer is selected from the second expert network according to the score.
6. The model training method according to claim 4, characterized in that The determining the third expert network from the first expert network includes: A first expert network stored in the same storage location as the second expert network is determined as the third expert network.
7. The model training method according to claim 1, characterized in that The training of the large language model based on the output result of the first expert network on the word-unit and the output result of the third expert network on the word-unit includes: Determining a loss value of the large language model according to an output result of the first expert network on the word-gram, an output result of the third expert network on the word-gram, and an annotation result of the training sample; Based on the loss value, network parameters of the large language model are adjusted, where the network parameters include a load threshold and a probability threshold of each expert network.
8. A text processing method, characterized in that: The method comprises: Inputting text into a large language model to obtain a processing result of the text, wherein the large language model is obtained by the model training method according to any one of claims 1 to 7.
9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the model training method according to any one of claims 1 to 7 or the text processing method according to claim 8 is implemented.
10. A computer program product, characterized in that The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the model training method according to any one of claims 1 to 7 or the text processing method according to claim 8.