Dense language model sparse upgrading method and sparse language model text processing method

By using tasks and context to represent the weight of the initialized routing network in the dense language model, the problem that existing sparse upgrade methods are difficult to significantly improve performance in instruction fine-tuning scenarios is solved, and more efficient computing and better task adaptability are achieved.

CN119940533APending Publication Date: 2025-05-06INST OF SOFTWARE - CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411937914.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing sparse upgrade method is difficult to significantly improve the performance of a hybrid expert model with billions of parameters in the instruction fine-tuning scenario, and there are instability problems during the training process.

Method used

By using task representation and contextual representation of the weight of the initialized routing network, it is possible to convert the dense language model into a sparse activation model without increasing the computational cost, improve the computing efficiency of the model and give the expert network professional processing capabilities.

Benefits of technology

While keeping the consumption of computing resources basically unchanged, the model's performance in complex inference, multi-task processing and other aspects will be significantly improved, and the model's adaptability to tasks and context will be improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940533A_ABST
    Figure CN119940533A_ABST
Patent Text Reader

Abstract

The invention relates to a dense language model sparse upgrading method and a sparse language model text processing method, and belongs to the technical field of artificial intelligence. According to the method, the weight of the routing network is initialized by using the task representation and the context representation, so that the dense language model is efficiently converted into the sparse activation model on the premise of not increasing the calculation cost, the calculation efficiency of the model is improved, the specialized processing capability of each expert network for different tasks is also endowed, and the calculation efficiency is improved. On the premise of keeping computing resource consumption basically unchanged, the performance of the model in the aspects of complex reasoning, multi-task processing and the like is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology, and in particular relates to a dense language model sparse upgrading method and a sparse language model text processing method. Background Art

[0002] With the rapid development of artificial intelligence technology, large language models have made significant breakthroughs in the field of natural language processing. As an important means to improve model performance, instruction fine-tuning technology has been proven to significantly improve the performance of models in zero-sample and few-sample scenarios through targeted training on specific task data sets. Existing studies have shown that improving the diversity and quality of instruction data can produce better results than simply increasing the amount of data. For example, some research teams use large language models such as ChatGPT and GPT-4 to build high-quality data sets to teach smaller-scale models to master reasoning and problem-solving skills. However, due to the size of parameters, such models often find it difficult to achieve ideal results when dealing with complex tasks.

[0003] The sparsely activated hybrid expert model adopts a technical solution that splits some parameters into expert modules and selectively activates some experts for different inputs during training and inference. This sparsity enables the hybrid expert model to have a large parameter scale while maintaining a moderate amount of computation. Compared with dense language models with comparable computational complexity, hybrid expert models generally show superior performance. However, hybrid expert models often face instability issues during training. Although researchers have proposed a variety of techniques to alleviate this problem, verifying these techniques on large-scale language models requires a lot of computing resources and time costs. Therefore, building a hybrid expert model based on a pre-trained dense language model is more feasible than training from scratch.

[0004] Sparse upscaling technology can transform existing dense language models into hybrid expert models. This technology has proven its effectiveness in the scenario of continued pre-training. However, continued pre-training requires a lot of computing resources, and as the scale of the basic model increases, the performance improvement effect gradually weakens. In particular, when building a hybrid expert model with billions of parameters in the scenario of instruction fine-tuning, the existing sparse upscaling methods are difficult to achieve significant performance improvements. Summary of the invention

[0005] In view of the above technical problems, the present invention proposes a method for sparse upgrading of dense language models. The present invention uses task representation and context representation to initialize the weights of the routing network, and realizes the efficient conversion of dense language models into sparse activation models without increasing the computational cost. It not only improves the computational efficiency of the model, but also gives each expert network specialized processing capabilities for different tasks. Under the premise of keeping the consumption of computing resources basically unchanged, it significantly improves the performance of the model in complex reasoning, multi-task processing, etc. Based on this, the present invention also proposes a sparse language model text processing method.

[0006] The technical solution adopted by the present invention to solve the above technical problems is as follows:

[0007] A method for sparsely updating a dense language model comprises the following steps:

[0008] Keep the attention network in the dense language model unchanged and replicate the feedforward network into multiple independent expert networks;

[0009] Add a routing network for activating the expert network. The routing network consists of a two-dimensional linear network. Each expert network corresponds to an embedding vector, and the activated expert network is selected based on the matching degree with the attention network features.

[0010] During the training phase, the routing network parameters are optimized.

[0011] Furthermore, for the multi-task learning scenario, the steps of initializing the routing network include:

[0012] 1) According to the perplexity of the dense language model on the data, the top multiple items with the highest perplexity are selected from each task dataset;

[0013] 2) Input the filtered data into the dense language model for forward calculation to extract the features of each word;

[0014] 3) Calculate the average of the word-unit features of each data to form the overall feature representation of each task;

[0015] 4) Use a clustering method to cluster the task features, generate feature vector clusters, and use them to initialize the routing network of the sparse language model.

[0016] Furthermore, for a general training scenario, the steps of initializing the routing network include:

[0017] 1) According to the perplexity of the dense language model on the data, the top multiple items with the highest perplexity are selected from each task dataset;

[0018] 2) Input the filtered data into the dense language model for forward calculation, extract the features of each word, and form a context feature set;

[0019] 3) Cluster all context features to generate feature vector clusters, which are used to initialize the routing network of the sparse language model.

[0020] Furthermore, the routing network adopts the Top-1 routing strategy to select the expert network, and the steps include:

[0021] 1) The routing network calculates the matching score between the input features and the embedding vector corresponding to each expert network;

[0022] 2) Use the softmax method to convert the matching scores of all expert networks into matching degree distribution;

[0023] 3) According to the matching degree distribution, several expert networks with the highest matching degree are selected as the target expert networks;

[0024] 4) Pass the input features to the selected target expert network for processing and generate the final output.

[0025] Furthermore, the matching score is calculated according to the following formula:

[0026] r(x)=W r ·x

[0027] Where r(x) represents the matching score, x represents the input feature; W r represents the routing network parameter matrix, each row vector of which represents the embedding vector of the corresponding expert network.

[0028] Furthermore, the matching degree distribution is calculated according to the following formula:

[0029]

[0030] where p i (x) represents the matching degree distribution, r(x) represents the matching score, x represents the input feature, and N is the number of assigned expert networks.

[0031] Furthermore, the input features are passed to the selected target expert network for processing, and the weighted combination of the outputs of each target expert network is used as the final output. The weighted calculation is as follows:

[0032]

[0033] Where y represents the final output, p i (x) represents the matching degree distribution, E i (x) represents the output of a single target expert network, and x represents the input feature.

[0034] Furthermore, during the training phase, the load balancing loss function is used to calculate the loss, and the routing network parameters are optimized with the goal of minimizing the loss so that all expert networks are load balanced;

[0035] The load balancing loss function is:

[0036]

[0037] Among them, α is the auxiliary loss function coefficient; N is the number of expert networks; is the batch size of the input features; L is the length of each input feature; f i is the proportion of words assigned to expert network i, P i is the proportion of expert network i selected, p(x),p i (x) represents the matching degree distribution, and x represents the input feature.

[0038] A sparse language model text processing method, wherein the sparse language model is obtained by upgrading the above method based on a dense language model, and the method comprises the following steps:

[0039] Enter the text to be processed;

[0040] The sparse language model models the word units in the text sentence through the attention network and generates word unit features;

[0041] The routing network of the sparse language model matches the embedded vectors of each expert network based on the extracted word features, finds the best matching expert network and activates it;

[0042] The activated expert network processes the input word features and outputs the processing results.

[0043] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.

[0044] A computer-readable storage medium stores a computer program, which implements the steps of the above method when executed.

[0045] The present invention has the following technical effects:

[0046] 1) The present invention uses the representation of task data in a dense language model to initialize the routing network of the upgraded sparse language model, effectively guiding the routing network to achieve task specialization allocation of the expert network, and improving the adaptability and processing accuracy of the model to the task;

[0047] 2) The present invention uses the representation of context data in the dense language model to initialize the routing network of the upgraded sparse language model, thereby realizing the specialized allocation of the expert network to different context scenarios and enhancing the model's adaptability to diverse contexts;

[0048] 3) The present invention activates the task processing capabilities of each expert network through multi-task instruction fine-tuning training, improves the performance of the model in multi-task scenarios, and meets diverse needs;

[0049] 4) The present invention enhances the expert network's ability to understand complex contexts through complex instruction fine-tuning training, significantly improving the model's performance and processing efficiency in complex tasks;

[0050] 5) The present invention adopts a sparse activation mechanism to activate only a small number of expert networks that best match the input features, thereby reducing the amount of computation and storage requirements and significantly improving the resource utilization and operating efficiency of the model;

[0051] 6) The present invention realizes feature-driven dynamic expert network selection by combining the attention network features with the allocation mechanism of routing network matching, ensuring that the model can intelligently process different tasks and contexts. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 Schematic diagram of the dense language model and sparse language model structure;

[0053] Figure 2 Schematic diagram for initializing routing network parameters. DETAILED DESCRIPTION

[0054] In order to make the various technical features and advantages or technical effects in the above technical solutions of the present invention more obvious and easy to understand, they are described in detail below in conjunction with embodiments and drawings.

[0055] The embodiment of the present invention specifically proposes a language model sparse upgrade method based on routing representation, which can efficiently construct a sparsely activated hybrid expert network model (i.e., sparse language model) based on a pre-trained dense language model (i.e., dense language model). The structures of the dense language model and the sparse language model are as follows: Figure 1 As shown in the figure, the dense language model is mainly composed of an attention network and a feedforward network, while the upgraded sparse language model is mainly composed of an attention network, a routing network and an expert network, wherein a sparse activation structure composed of a routing network and a multi-expert network is used to replace the feedforward network in the dense language model, and the expert network has the same network architecture as the feedforward network. The following is a specific description of this method.

[0056] The process of converting a pre-trained dense language model to a sparse language model is as follows:

[0057] S1: Keep the attention network of the dense language model unchanged and copy the feedforward network into multiple independent expert networks; here we choose to convert all layers of the feedforward network of the original dense language model into sparse layers of the expert network. This is to overcome the problem that the sparse layers in the existing methods will significantly increase the parameter scale of the entire model, resulting in an increase in computational complexity and consumption of video memory resources.

[0058] S2: A new routing network is added to decide which expert networks to activate; the routing network consists of a two-dimensional linear network composed of a set of representations. Each representation is regarded as an embedding vector of the corresponding expert network, which is used to calculate the matching degree with the features of the attention network and select the expert network accordingly.

[0059] In the above upgrade process, except for the weight of the routing network which is a newly added parameter, other weights are all from the initial dense language model.

[0060] 1. Parameter initialization, the process is as follows Figure 2 As shown:

[0061] In order to avoid the parameter mismatch problem caused by random initialization of new parameters of the routing network, two parameter initialization schemes are proposed:

[0062] 1. Multi-task learning scenario

[0063] For multi-task learning scenarios, assuming that the training dataset It consists of N clearly divided task data, and the parameter initialization steps are as follows:

[0064] 1) Filter a fixed proportion of data from each task data set based on the high perplexity of the initial dense language model to form a feature data set

[0065] 2) Input the feature data set into the dense language model for forward calculation. Each layer of attention network models the features of each word and generates word features;

[0066] 3) Calculate the average value of the word-unit features of each data as the feature representation of the data to obtain the task feature set

[0067] 4) Cluster the feature data of each task to obtain a feature vector set Used to initialize each layer of the routing network of the sparse language model.

[0068] The advantages of this initialization scheme are:

[0069] a) The training data of common language model training tasks (such as math problem solving, code generation, logical reasoning, etc.) often form vector clusters based on tasks in the feature space generated by the language model. This solution follows this rule;

[0070] b) The allocation mechanism of the routing network is essentially the similarity matching between the output features of the attention network and the representation of the expert network in the routing network. Using the features of the attention network to initialize the parameters of the routing network conforms to this mechanism and guides the router to perform specialized task allocation to the expert network.

[0071] c) Task-oriented allocation helps guide the expert network to handle different tasks and alleviates the homogeneity problem caused by copying from the feedforward network to the expert network.

[0072] 2. General training scenarios

[0073] For general training scenarios without explicit task division, the parameter initialization steps are as follows:

[0074] 1) Filter a fixed proportion of data from the complete data set. The basis for the selection is that the initial dense language model produces a high degree of confusion on the data to form a feature data set.

[0075] 2) Input the feature data set into the dense language model for forward calculation. Each layer of the attention network models the features of each word unit and generates word unit features. Each word unit feature is used as an independent context feature to obtain the feature set

[0076] 3) Perform K-means clustering on all context features (K is equal to the number of expert networks) to obtain a feature vector set Used to initialize each layer of the routing network of the sparse language model.

[0077] The advantages of this initialization scheme are:

[0078] a) Applicable to a wider range of general training scenarios;

[0079] b) Token-level context features provide more fine-grained data modeling capabilities than task features, guiding the router to perform specialized context assignments on the expert network.

[0080] 2. Routing Strategy

[0081] The routing network adopts the Top-1 routing strategy, that is, each input feature is matched to only one expert network to control the computational complexity of the sparse language model to be comparable to that of the dense language model. The specific implementation steps are as follows:

[0082] The routing network receives feature input Distribute it to N expert networks The expert network with the highest matching degree in . Routing network parameter matrix Each row vector in Represents the embedding vector of the corresponding expert network. The matching degree calculation process is:

[0083] a) Calculate the matching score between the feature and the expert network embedding vector: r(x) = W r x;

[0084] b) The expert network matching degree distribution is obtained through softmax normalization:

[0085]

[0086] c) Select the expert network with the highest matching degree to form an expert network set

[0087] d) The output of the hybrid expert network is calculated by weighted combination of the selected expert networks:

[0088]

[0089] Expert Network Load Balancing

[0090] To promote uniform routing among expert networks, a differentiable load balancing loss function is adopted:

[0091]

[0092] Where N is the number of expert networks; is the batch size of the input feature; L is the length of each input feature; α is the auxiliary loss function coefficient; f i is the proportion of words assigned to expert network i:

[0093]

[0094] P i is the proportion of expert network i selected:

[0095]

[0096] The following is an example of the application of this method:

[0097] The input data is a question entered by the user: "What is the weather like today?"

[0098] The specific processing steps are as follows:

[0099] 1) Input feature extraction: The language model receives input text, models the words in the sentence (such as "today", "weather", "how") through the attention network, and generates semantic feature representation.

[0100] 2) Routing network assignment: The routing network matches the embedding vectors of each expert network based on the extracted semantic features. Assuming that there are multiple expert networks focusing on different tasks (such as weather query, news understanding, and daily conversation), the routing network recognizes that the input text best matches the features of the "weather query" expert network, so it activates the expert network.

[0101] 3) Expert network processing: The activated expert network refines the input features, such as identifying key information "today" and "weather", and associating them with the weather data in the knowledge base.

[0102] 4) Output: The expert network generates the output: "Today's weather is sunny and the temperature is 25℃."

[0103] The advantages of this method are:

[0104] Efficiency: The sparse activation mechanism allows only a few expert networks to participate in the calculation, saving resources.

[0105] Task specialization: The activated expert network focuses on a specific domain, such as weather-related tasks, thus improving the accuracy of the results.

[0106] Another embodiment of the present invention provides a computer device (computer, server, smart phone, etc.), which includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing each step in the method of the present invention.

[0107] Another embodiment of the present invention provides a computer-readable storage medium (such as ROM / RAM, magnetic disk, optical disk), wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a computer, the steps of the method of the present invention are implemented.

[0108] The specific embodiments of the present invention disclosed above are intended to help understand the content of the present invention and implement it accordingly. It can be understood by those skilled in the art that various replacements, changes and modifications are possible without departing from the spirit and scope of the present invention. The present invention should not be limited to the contents disclosed in the embodiments of this specification, and the scope of protection of the present invention shall be subject to the scope defined in the claims.

Claims

1. A method for sparse upgrading of dense language models, characterized in that: The following steps are involved: Keep the attention network in the dense language model unchanged and replicate the feedforward network into multiple independent expert networks; Add a routing network for activating the expert network. The routing network consists of a two-dimensional linear network. Each expert network corresponds to an embedding vector, and the activated expert network is selected based on the matching degree with the attention network features. During the training phase, the routing network parameters are optimized.

2. The method according to claim 1, characterized in that For multi-task learning scenarios, the steps to initialize the routing network include: 1) According to the perplexity of the dense language model on the data, the top multiple items with the highest perplexity are selected from each task dataset; 2) Input the filtered data into the dense language model for forward calculation to extract the features of each word; 3) Calculate the average of the word-unit features of each data to form the overall feature representation of each task; 4) Use a clustering method to cluster the task features, generate feature vector clusters, and use them to initialize the routing network of the sparse language model.

3. The method according to claim 1, characterized in that For general training scenarios, the steps to initialize the routing network include: 1) According to the perplexity of the dense language model on the data, the top multiple items with the highest perplexity are selected from each task dataset; 2) Input the filtered data into the dense language model for forward calculation, extract the features of each word, and form a context feature set; 3) Cluster all context features to generate feature vector clusters, which are used to initialize the routing network of the sparse language model.

4. The method according to claim 3, characterized in that K-means clustering is performed on all context features, where K is equal to the number of expert networks.

5. The method according to claim 1, characterized in that The routing network uses the Top-1 routing strategy to select the expert network. The steps include: 1) The routing network calculates the matching score between the input features and the embedding vector corresponding to each expert network; 2) Use the softmax method to convert the matching scores of all expert networks into matching degree distribution; 3) According to the matching degree distribution, several expert networks with the highest matching degree are selected as the target expert networks; 4) Pass the input features to the selected target expert network for processing and generate the final output.

6. The method according to claim 5, characterized in that The matching score is calculated according to the following formula: r(x)=W r ·x Where r(x) represents the matching score and x represents the input feature; W r represents the routing network parameter matrix, each row vector of which represents the embedding vector of the corresponding expert network.

7. The method according to claim 5, characterized in that The matching degree distribution is calculated according to the following formula: where p i (x) represents the matching degree distribution, r(x) represents the matching score, x represents the input feature, and N is the number of assigned expert networks.

8. The method according to claim 5, characterized in that The input features are passed to the selected target expert network for processing, and the weighted combination of the outputs of each target expert network is used as the final output. The weighted calculation is as follows: Where y represents the final output, p i (x) represents the matching degree distribution, E i (x) represents the output of a single target expert network, and x represents the input feature.

9. The method according to claim 1, characterized in that In the training phase, the load balancing loss function is used to calculate the loss, and the routing network parameters are optimized with the goal of minimizing the loss so that all expert networks are load balanced; The load balancing loss function is: Among them, α is the auxiliary loss function coefficient; N is the number of expert networks; is the batch size of the input features; L is the length of each input feature; f i is the proportion of words assigned to expert network i, P i is the proportion of expert network i selected, p(x),p i (x) represents the matching degree distribution, and x represents the input feature.

10. A sparse language model text processing method, wherein the sparse language model is obtained by upgrading the dense language model based on the method described in any one of claims 1 to 9, characterized in that: The method comprises the following steps: Enter the text to be processed; The sparse language model models the word units in the text sentence through the attention network and generates word unit features; The routing network of the sparse language model matches the embedded vectors of each expert network based on the extracted word features, finds the best matching expert network and activates it; The activated expert network processes the input word features and outputs the processing results.