Model construction and training method and device and storage medium
By multiplexing the network structure and parameters of pre-trained dense models, building sparse models is solved, and a more efficient training process and model performance is achieved.
Patent Information
- Application Number
- CN202311636797.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-30
- Publication Date
- 2025-05-30
AI Technical Summary
The training cost of existing sparse models is high, mainly due to the need to pre-train based on large-scale sample data, resulting in increased resource and time costs.
By multiplexing the network structure of the target dense model completed by pre-training, the parameter amount is expanded to obtain the sparse units of the sparse model, and the parameter information of the target dense model is used to initialize the initial sparse model to reduce the dependence on large-scale sample data.
The training cost of sparse models is reduced, and by directly using the samples of professional tasks for fine-tuning, the pre-training of large-scale sample data is avoided, and the training efficiency and the convergence speed of the model are improved.
Smart Images

Figure CN120067557A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a method, device, and storage medium for model construction and training. Background Art
[0002] A large model refers to a machine learning model with a large number of parameters and a complex structure. These models can be applied to process large-scale data and complex problems. Although large models have shown good performance in many scenarios, there is still significant room for improvement in some professional fields. Models with a larger parameter scale require more resources for training and deployment, which limits the further optimization and use of large models.
[0003] During the training and inference of a sparse model, only some of the parameters of the sparse model are used when calculating each sample. While expanding the model parameter scale, high training and inference efficiency can be maintained. However, the training cost of existing sparse models is relatively high. Summary of the Invention
[0004] Multiple aspects of this application provide a method, device, and storage medium for model construction and training to reduce the training cost of sparse models.
[0005] An embodiment of this application provides a model construction method, including:
[0006] Obtain a target dense model and parameter information of the target dense model; the target dense model is a pre-trained dense model or a fine-tuned dense model;
[0007] From the network structure of the target dense model, obtain the network structure of the dense units of the target dense model and the connection relationship between the dense units;
[0008] Expand the number of parameters based on the network structure of the dense units to obtain the sparse units of the sparse model; the number of sparse units is the same as that of the dense units;
[0009] Connect the sparse units according to the connection relationship between the dense units to obtain an initial sparse model;
[0010] Use the parameter information of the target dense model to initialize the parameters of the initial sparse model to obtain a sparse model to be trained.
[0011] An embodiment of this application also provides a model training method, including:
[0012] Obtain training sample pairs for a target task; each training sample pair includes: a first text and a second text corresponding to the first text in the target task;
[0013] Predict a third text corresponding to the first text in the target task by using a sparse model to be trained;
[0014] Adjust model parameters of the sparse model to be trained according to the difference between the third text and the second text, so as to obtain a target sparse model;
[0015] Wherein, the sparse model to be trained is obtained by initializing parameters of an initial sparse model by using parameter information of a target dense model; the initial sparse model is obtained by expanding the number of parameters on the basis of the network structure of dense units of the target dense model; the target dense model is a pre-trained dense model or a fine-tuned dense model.
[0016] An embodiment of this application further provides a computing device, including: a memory and a processor; wherein, the memory is used to store a computer program;
[0017] The processor is coupled to the memory and is used to execute the computer program to perform the steps in the above model construction method and / or model training method.
[0018] An embodiment of this application further provides a computer-readable storage medium storing computer instructions, which when executed by one or more processors, cause the one or more processors to perform the steps in the above model construction method and / or model training method.
[0019] In the embodiment of this application, the network structure of a pre-trained target dense model or a fine-tuned target dense model can be reused, and the number of parameters is expanded on the basis of the network structure of dense units of the target dense model to obtain sparse units of the sparse model; and the parameters of the initial sparse model are initialized by using the parameter information of the target dense model, reusing the training and learning results of the target dense model. Therefore, when training the sparse model to be trained, samples of professional tasks can be directly used for fine-tuning, without having to pre-train with a large amount of sample data again, which can reduce the training cost of the sparse model. Description of the Drawings
[0020] The drawings described herein are used to provide a further understanding of this application and constitute a part of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:
[0021] Figure 1 It is a comparison schematic diagram of a traditional dense model and a sparse model provided by an embodiment of this application;
[0022] Figure 2a It is a flow schematic diagram of a model construction method provided by an embodiment of this application;
[0023] Figure 2b This is a schematic structural diagram of the sparse model provided by the embodiments of the present application;
[0024] Figures 3 - 10 This is a schematic structural diagram of some other sparse models provided by the embodiments of the present application;
[0025] Figures 11 - 13 This is a schematic flowchart of the model training method provided by the embodiments of the present application;
[0026] Figure 14 This is a schematic structural diagram of the computing device provided by the embodiments of the present application. Detailed implementation manners
[0027] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0028] Figure 1 This is a comparison schematic diagram of the traditional dense model and the sparse model provided by the embodiments of the present application. During the training and inference processes of the dense model, the calculation of each sample requires the use of all the parameters of the dense model. During the training and inference processes of the sparse model, only some of the parameters of the sparse model are used when calculating each sample, which can maintain high training and inference efficiency while expanding the scale of the model parameters. Therefore, the sparse model is widely used in fields such as natural language processing (NLP) and computer vision (CV).
[0029] In the present application, the dense model and the sparse model are neural network models. In some embodiments, when the number of model parameters of the dense model and the sparse model is large, for example, when the number of parameters of the dense model and the sparse model is in the millions, hundreds of millions, billions or even more, the dense model and the sparse model can also be referred to as large models, such as large language models (LLMs). In the embodiments of the present application, a large model is defined as a neural network model whose number of model parameters conforms to a preset parameter number range. Among them, the number of model parameters corresponding to the preset parameter number range is very large, which can be in the millions, hundreds of millions, billions or even more, and the specific value can be determined by the standards in the field of artificial intelligence (AI).
[0030] In the embodiments of the present application, the specific implementation frameworks of the dense model and the sparse model are not limited. The dense model and the sparse model can be based on frameworks such as Convolutional Neural Network (CNN), Recurrent Neural Network (RNN), or Transformer framework, etc. Figure 1 Only the neural network models in which the dense model and the sparse model are based on the Transformer framework are taken as examples for illustration, but this does not constitute a limitation.
[0031] As Figure 1 As shown in the left figure, the dense model may include at least one dense unit. Generally, the dense model includes multiple dense units. Multiple means two or more. Among them, the dense unit is the basic module of the dense model and is used to extract features from the input information of the dense model; through the feature extraction of multiple dense units, the feature information of the input information of the dense model is finally obtained. The dense unit includes an expert layer and other network architectures. The number of expert layers can be one or more. Multiple means two or more.
[0032] The expert layer can be implemented as modules such as Feed Forward Network Layer (FFN), Multilayer Perceptron (MLP), Multi-head Attention, Support Vector Machines (SVM), Gaussian processes (GP), Hidden Markov Models (HMM), CNN, or RNN, etc., but not limited to this. For example, as Figure 1 shown, for the dense model implemented by the Transformer framework, the dense unit can be a Transformer Block. The expert layer of the Transformer Block can be implemented as an FFN module; other network architectures can include: multi-head attention layer and residual connection and normalization (Add&Normalize, Add&Norm) layer, etc.
[0033] Among them, Attention maps the query statement (Query) and the key (Key) into the same high-dimensional space to calculate the similarity, while Multi-head Attention maps the query statement (Query) and the key (Key) into different subspaces of the high-dimensional space to calculate the similarity. The essence of Multi-head Attention is to map the same query statement (Query), key (Key), and value (Value) into different subspaces of the original high-dimensional space to calculate the attention while keeping the total number of parameters unchanged, and finally merge the attention information in different subspaces. This reduces the dimension of each vector when calculating the attention of each head and can prevent overfitting.
[0034] The expert layer (such as FFN) is used for spatial transformation and may include multiple linear transformation layers and activation function layers. For example, it may include 2 linear transformation layers. The activation function can be the ReLu function, sigmoid function, tanh function, etc. The expert layer (such as FFN) introduces non-linearity (activation function layer) and transforms the output space of the multi-head attention layer, thereby increasing the performance of the model.
[0035] The residual connection and normalization layer are used to add residual connection and normalization operations between the multi-head self-attention mechanism and the expert layer (such as FFN). The role of this layer is to add the output of the previous layer to the input of the previous layer and perform normalization to better transmit information and control the gradient. The operations of the residual connection and normalization layer can be divided into the following steps: (1) Residual connection: Add the output of the previous layer to the input of the previous layer to obtain a residual vector; (2) Normalization: Normalize the residual vector to better transmit information and control the gradient; (3) Linear transformation: Perform a linear transformation on the normalized vector to better adapt to the input of the next layer. The residual connection and normalization layer can avoid the problems of gradient disappearance or explosion while maintaining information fluency, thereby improving the training efficiency and performance of the model.
[0036] In addition, Figure 1 the hidden states in
[0037] such as Figure 1 As shown in the right figure, the sparse model may include at least one sparse unit. Generally, the sparse model includes multiple sparse units. Multiple means 2 or more. The sparse unit includes an expert layer and other network architectures. The number of expert layers can be 1 or more. Multiple means 2 or more. Such as Figure 1As shown, the main change of the sparse model compared with the dense model is that the expert layer of the sparse model includes multiple expert modules (Experts). Multiple means two or more. Figure 1 The right figure only shows the expert layer including Z expert modules, but it does not constitute a limitation. Z ≥ 2 and is an integer. The expert module has the same parameters and network structure as the expert layer of the dense model. For example, Figure 1 In the right figure, the expert layer of the sparse model includes multiple FFN modules. The FFN module and Figure 1 the expert layer (FFN layer) shown in the left figure have the same parameters and network structure.
[0038] Among them, the sparse unit is the basic module of the sparse model, which is used to extract features from the input information of the sparse model; after the feature extraction of multiple sparse units, the feature information of the input information of the sparse model is finally obtained. When the sparse unit extracts features from the input information of the sparse model, when passing through the expert layer, some of the multiple expert modules are used to extract features from the input information, which can maintain a high processing efficiency equivalent to that of the dense model.
[0039] According to Figure 1 the schematic diagrams of the network structures of the dense model and the sparse model shown, it can be seen that the sparse model has an increased number of model parameters compared with the dense model. Therefore, in an ideal situation, the sparse model has a better processing effect than the dense model. On the other hand, the sparse model only uses some of the parameters of the sparse model when calculating each sample. For example, when selecting some expert modules of the expert layer for calculation, it can maintain high training efficiency while expanding the scale of the model parameters.
[0040] However, the inventors of this application have found through research that if a traditional sparse model wants to achieve the processing effect of a dense model in the processing of tasks in a professional field, or obtain a better processing effect than the dense model, it needs to start from the pre-training stage, learn and train the initial sparse model based on a large amount of sample data to obtain a pre-trained sparse model. Then, use high-quality data to perform instruction fine-tuning on the pre-trained sparse model. After that, use the samples of professional tasks to fine-tune the sparse model after instruction fine-tuning. The scale of the sample data in the early pre-training is extremely large, resulting in a very high training cost in the pre-training stage of the sparse model, and thus a relatively high training cost of the sparse model. The training cost of the sparse model includes but is not limited to: the consumed resource cost and time cost, etc.
[0041] In some embodiments of the present application, in order to reduce the training cost of the sparse model, the network structure of the pre-trained target dense model or the fine-tuned target dense model can be reused, and the number of parameters can be expanded on the basis of the dense units of the target dense model to obtain the sparse units of the sparse model; and the initial sparse model is parameter-initialized using the parameter information of the target dense model, reusing the training learning results of the target dense model. Therefore, when training the sparse model to be trained, the samples of the professional task can be directly used for fine-tuning without reusing a large amount of sample data for pre-training, which can reduce the training cost of the sparse model.
[0042] The following will describe in detail the technical solutions provided by the embodiments of the present application with reference to the accompanying drawings.
[0043] It should be noted that the same reference numerals represent the same object in the following drawings and embodiments. Therefore, once an object is defined in one drawing or embodiment, it does not need to be further discussed in the subsequent drawings and embodiments.
[0044] Figure 2a It is a schematic flowchart of the model construction method provided by the embodiments of the present application. This model construction method can be applied to any electronic device with information processing functions. The electronic device can be a terminal device such as a desktop computer, a laptop computer, a mobile phone, or an Internet of Things device; it can also be various server devices such as a traditional server, a cloud server, or a server cluster. As Figure 2a shown, this model construction method mainly includes:
[0045] 201. Obtain the target dense model and the parameter information of the target dense model; the target dense model is a pre-trained dense model or a fine-tuned dense model.
[0046] 202. Obtain the network structure of the dense units of the target dense model and the connection relationship between the dense units from the network structure of the target dense model.
[0047] 203. Expand the number of parameters on the basis of the network structure of the dense units to obtain the sparse units of the sparse model; the number of sparse units is the same as that of the dense units.
[0048] 204. Connect the sparse units according to the connection relationship between the dense units to obtain the initial sparse model.
[0049] 205. Use the parameter information of the target dense model to perform parameter initialization on the initial sparse model to obtain the sparse model to be trained.
[0050] The inventors of the present application have found through research that if a traditional sparse model wants to achieve the processing effect of a dense model in task processing in a professional field, or obtain a better processing effect than a dense model, it is necessary to start from the pre-training stage, learn and train the initial sparse model based on a large amount of sample data to obtain a pre-trained sparse model. Then, use high-quality data to perform instruction fine-tuning on the pre-trained sparse model. After that, use the samples of tasks in the professional field to fine-tune the sparse model after instruction fine-tuning. The scale of the sample data in the early pre-training is extremely large, resulting in a high training cost in the pre-training stage of the sparse model, and thus a relatively high training cost of the sparse model.
[0051] Among them, instruction fine-tuning is a specific fine-tuning method. The model for instruction adjustment receives a pair of input and output, which describes the task of guiding the model. Among them, the input is the instruction, and the output is the output prompt information for guiding the model. For example, the instruction (Instruction) is: Write a list of interesting weekend activities; the output (Output) is: Hiking, spending a day in the park, having a picnic, watching a movie at night. The instruction can be any text, such as writing an email, editing a sentence, etc. It is expected that the model can generalize well in many instruction-driven tasks, so that the model learns to follow the pattern of instructions and outputs, and becomes more valuable when answering questions that people are often interested in. In other words, instruction training improves the quality of the answers given by the language model by training the model in a way that is consistent with the format in which humans tend to give and receive instructions. The instructions used in the instruction fine-tuning stage are generally sample pairs in the general field.
[0052] Fine-tuning the sparse model after instruction fine-tuning using the samples of tasks in the professional field means using a small amount of specific domain data to fine-tune the pre-trained model or the model after instruction fine-tuning, so as to further optimize the model's ability in the specific domain.
[0053] Since the large-scale pre-training data used in the early stage of training the sparse model and the high-quality instruction fine-tuning data used in the subsequent instruction fine-tuning are not open source, it is difficult to train the sparse model from scratch. Even if the sample data in the pre-training stage is open source, due to the large amount of sample data in the pre-training stage, the cost of reconstructing a larger-scale sparse model based on these data will also be very high.
[0054] In this embodiment, in order to reduce the training cost of the sparse model, a sparse model to be trained is constructed based on the pre-trained dense model or the dense model after fine-tuning training. In this way, in the sparse model training stage, the previous training and learning results of the dense model can be directly reused, and there is no need to retrain the sparse model from the pre-training stage, which can reduce the training cost of the sparse model. The construction method of the sparse model will be specifically described below.
[0055] In each embodiment of the present application, the sparse model may be a mixture of experts (MoE). For the convenience of description, the pre-trained dense model or the fine-tuned dense model is defined as the target dense model. In order to reuse the previous training and learning results of the dense model, in step 201, the target dense model and the parameter information of the target dense model are obtained. The parameter information of the target dense model may include: the model parameters of the target dense data and the parameter values of the model parameters. The parameter values of the model parameters are the parameter values of the model parameters when the dense model is pre-trained; or the parameter values of the model parameters when the dense model is fine-tuned. In the embodiments of the present application, the fine-tuned dense model may be a dense model after instruction fine-tuning; or a dense model fine-tuned using professional tasks in a specific field.
[0056] Further, in step 202, the network structure of the dense units of the target dense model and the connection relationship between the dense units can be obtained from the network structure of the target dense model. The target dense model includes at least one dense unit. Generally, the target dense model includes multiple dense units. Multiple means two or more. For the description of the dense units, reference can be made to the relevant content above Figure 1 which will not be elaborated here.
[0057] In this embodiment, the network structure of the dense unit includes: the connection relationship between the layers in the dense unit and the model parameters of the dense unit. Among them, the connection relationship between the layers in the dense unit can be represented by the mathematical relationship between the model parameters of the dense unit.
[0058] Based on the comparison chart of the dense model and the sparse model shown above Figure 1 it can be seen that the difference between the sparse units of the sparse model and the dense units of the dense model is that: the expert layer of the sparse unit includes multiple expert modules with the same parameters and network structure as the expert layer of the dense unit. Multiple means two or more. That is, the expert layer of the sparse model includes multiple expert layers of the dense unit. The other network structures of the sparse unit except the expert layer are the same as those of the dense unit. Therefore, in step 203, the number of parameters can be expanded on the basis of the network structure of the dense unit to obtain the sparse units of the sparse model; the number of sparse units is the same as that of the dense units.
[0059] In the embodiments of the present application, the specific implementation form of expanding the number of parameters on the basis of the network structure of the dense unit is not limited.
[0060] In some embodiments, based on the network structure of the dense unit, the parameter quantity of the expert layer of the dense unit can be expanded according to the network structure of the expert layer of the dense unit; and a routing module can be added before the expert layer of the dense unit after expanding the parameter quantity to obtain a sparse unit. Among them, the expert layer of the dense unit after expanding the parameter quantity is the expert layer of the sparse unit.
[0061] In some embodiments, based on the network structure of the dense unit, the expert layer of the dense unit can be modified to expand the parameter quantity of the expert layer of the dense unit. Specifically, based on reusing the network structure of other parts of the dense unit and the connection relationship between the expert layer of the dense unit and the network structure of other parts of the dense unit, the parameter quantity of the expert layer of the dense unit is expanded according to the network structure of the expert layer of the dense unit, so that the expert layer of the dense unit after expanding the parameter quantity includes at least one expert module, and the network structure of each expert module is the same as that of the expert layer of the dense unit before expanding the parameter quantity, and the total number of expert modules of the dense unit after expanding the parameter quantity is greater than the total number of the expert layer of the dense unit before expanding the parameter quantity. Since the expert modules of the dense unit after expanding the parameter quantity have the same network structure as the expert layer of the dense unit before expanding the parameters, the total number of expert modules of the dense unit after expanding the parameter quantity is greater than the total number of the expert layer of the dense unit before expanding the parameters, which can make the model parameter quantity of the dense unit after expanding the parameter quantity greater than the model parameter quantity of the dense unit before expanding the parameters, realizing the expansion of the parameter quantity of the expert layer of the dense unit.
[0062] After that, a routing module can be added before the expert layer of the dense unit after expanding the parameter quantity to obtain a sparse unit. Among them, the expert layer of the dense unit after expanding the parameter quantity is the expert layer of the sparse unit. The routing module is used to select a target expert module from at least one expert module included in each expert layer of the sparse unit to process the input of the expert layer of the sparse unit.
[0063] In some other embodiments, the parameter quantity of the dense unit can be expanded by copying the network structure of the dense unit. Specifically, the network structure of the expert layer of the dense unit, the network structure of other parts of the dense unit, and the connection relationship between the expert layer of the dense unit and the network structure of other parts of the dense unit can be obtained from the network structure of the dense unit. The network structure of other parts of the dense unit refers to the network structure of other layers in the dense unit except the expert layer, such as Figure 1 the multi-head attention layer and the residual connection and normalization layer of the dense unit in
[0064] Further, according to the network structure of the expert layer of the dense unit, the number of parameters of the expert layer of the dense unit can be expanded to obtain the expert layer of the sparse unit. Each expert layer of the sparse unit includes at least one expert module, and each expert module has the same network structure as the expert layer of the dense unit. Moreover, the total number of expert modules of the sparse unit is greater than the total number of expert layers of the dense unit. Since the expert module of the sparse unit has the same network structure as the expert layer of the dense unit, the total number of expert modules of the sparse unit is greater than the total number of expert layers of the dense unit, which can make the number of model parameters of the sparse unit greater than that of the dense unit, realizing the expansion of the number of parameters of the expert layer of the dense unit.
[0065] Optionally, as Figure 2b shown, for any expert layer i of the dense unit, the network structure of any expert layer i of the dense unit is used as an expert module and copied to the corresponding expert layer of the sparse model (i.e., copied to the i-th expert layer of the sparse unit) to obtain the expert layer of the sparse unit. At least one expert layer in the expert layer of the sparse unit includes multiple expert modules. Multiple means two or more. i represents the i-th expert layer (i.e., expert layer i) in the dense unit or the sparse unit. i = 1, 2,..., N. N represents the total number of expert layers of the dense unit or the sparse unit. N ≥ 1 and is an integer. Generally, N ≥ 2 and is an integer. The number of expert layers of the sparse unit is the same as that of the dense unit.
[0066] Since at least one expert layer in the expert layer of the sparse unit includes multiple expert modules and the expert module of the sparse unit has the same network structure as the expert layer of the dense unit, the total number of expert modules included in the sparse unit is greater than the total number of expert layers of the dense module, which can make the number of model parameters of the sparse unit greater than that of the dense unit, realizing the expansion of the number of parameters of the expert layer of the dense unit.
[0067] In the embodiments of the present application, the number of expert modules included in each expert layer of the sparse unit is not limited, that is, the specific structure of the expert layer of the sparse unit is not limited. As long as the total number of expert modules of the sparse unit is greater than the number of model parameters of the dense unit, it can be implemented as the expert layer of the sparse unit provided in the embodiments of the present application.
[0068] For example, as Figure 2b , Figure 3 and Figure 4 shown, the sparse unit may include multiple expert layers. Each expert layer may include multiple expert modules. In the embodiments of the present application, multiple means two or more. Figures 2b - 4 In, the number of expert layers is illustrated as N, but it does not constitute a limitation. N ≥ 2 and is an integer. Figures 2b - 4In [description], the "E" in the expert layer of the sparse unit represents the expert module. The expert module is replicated from the network structure of the expert layer of the dense unit.
[0069] In the embodiments of the present application, the number of expert modules in each expert layer of the sparse unit is not limited. Optionally, as Figure 2b shown, the number of expert modules included in multiple expert layers of the sparse unit is the same. Optionally, multiple expert layers of the sparse unit may each include multiple expert modules, and the number of expert modules included is the same. The number of expert modules included in multiple expert layers of the sparse unit can be flexibly set according to actual needs. For example, multiple expert layers of the sparse unit may each include 3, 4, 5, or 6 or even more expert modules. Figure 2b In [description], only the case where multiple expert layers of the sparse unit each include 2 expert modules is illustrated as an example, but it does not constitute a limitation.
[0070] For another example, as Figure 3 shown, the number of expert modules included in some of the multiple expert layers of the sparse unit is the same, and each expert layer includes multiple expert modules. The number of expert modules included in each expert layer can be flexibly set according to actual needs. The embodiments of the present application do not limit which specific expert layers have the same number of included expert modules. Generally, the expert layers with the same number of included expert modules are adjacent expert layers. For example, Figure 3 in [description], the number of expert modules included in the 2nd expert layer (i.e., expert layer 2) to the (N - 2)th expert layer (i.e., expert layer (N - 2)) is the same, and the number of expert modules included in expert layer (N - 1) and expert layer N is the same, etc.
[0071] Since the content closer to the output layer of the sparse model has a greater impact on the final output result of the model, the larger the number of model parameters in the expert layer closer to the output layer of the sparse model, the more beneficial it is to improve the quality of the final output result of the model. Therefore, for other expert layers in multiple expert layers of the sparse unit except those with the same number of included expert modules, the number of expert modules included in the expert layer closer to the output layer of the sparse model can be set to be larger.
[0072] In some other embodiments, the number of expert modules included in multiple expert layers of the sparse unit is different, and each expert layer includes multiple expert modules. The number of expert modules included in each expert layer can be flexibly set according to actual needs. The embodiments of the present application do not limit which specific expert layers have the same number of expert modules. Generally, in order to improve the quality of the output result of the sparse model, the closer the expert layer is to the output layer of the sparse model, the more expert modules it may include. That is, the number of expert modules included in multiple expert layers of the sparse unit can increase layer by layer. Among them, the closer the expert layer is to the output layer of the sparse model, the larger the layer number. In this embodiment, the number of expert modules increased layer by layer in the expert layer is not limited. The number of expert modules increased in each layer can be the same or different, and can be specifically and flexibly set according to actual needs. For example, the number of expert modules included in multiple expert layers of the sparse unit can increase by 1, 2, or 3 layer by layer, etc. For example, as Figure 4 shown, the first expert layer of the sparse unit may include 2 expert modules, and the other expert layers increase by 1 expert module layer by layer as the layer number increases, that is, the i-th expert layer (i.e., expert layer i) includes (i + 1) expert modules.
[0073] The above embodiments take the sparse unit including multiple expert layers, and each expert layer including multiple expert modules as an example to exemplarily illustrate the network structure of the expert layer of the sparse unit, but do not constitute a limitation. Of course, in combination with Figures 5 - 10 , the sparse unit may include multiple expert layers, some expert layers include multiple expert modules, and other expert layers include a single (i.e., 1) expert module. In Figures 5 - 10 , it is illustrated with the number of expert layers being N, but it does not constitute a limitation. N ≥ 2 and is an integer. Figures 5 - 10 In the sparse unit of
[0074]
[0075] Figures 5 - 10 In the embodiments of the present application, the specific layout of the first expert layer and the second expert layer is not limited. In some embodiments, as Figures 5 - 10 shown, at least one second expert layer is connected between two adjacent first expert layers. For example, 1 second expert layer is connected between two adjacent first expert layers to achieve increasing expert modules at intervals.
[0076] In some embodiments, the first expert layer may be a plurality of consecutive expert layers close to the sparse model; the second expert layer is the other expert layers in the sparse unit except the second expert layer. This situation is not illustrated.
[0077] The following will combine Figures 5 - 10 with the schematic diagram of the network structure of the sparse unit shown in FIG., and exemplarily illustrate the specific implementation manner in which at least one second expert layer is connected between two adjacent first expert layers.
[0078] Embodiment 1: As Figures 5 - 7 shown, for a sparse unit in which at least one second expert layer is connected between two adjacent first expert layers, the number of expert modules included in the K first expert layers close to the output layer of the sparse model is different, and the closer the first expert layer is to the output layer of the sparse model among the K first expert layers, the more expert modules it includes. Among them, the number of expert modules included in the K first expert layers close to the output layer of the sparse model is greater than the number of expert modules included in the other first expert layers.
[0079] The other first expert layers refer to the other first expert layers in the sparse unit except the K first expert layers. Among them, 2 ≤ K < (N - Q), and K is an integer, N is the total number of expert layers included in the sparse unit, and Q is the number of second expert layers included in the sparse unit. The number of the other first expert layers may be one or more.
[0080] For multiple other first expert layers, as Figure 5 shown, the number of expert modules included in the other multiple first expert layers may be the same. Or, as Figure 6 shown, among the other multiple first expert layers, the closer the other first expert layer is to the output layer of the sparse model, the more expert modules it includes. Or, as Figure 7 shown, the number of experts included in the M other first expert layers close to the K first expert layers among the multiple other first expert layers is the same, and is greater than the number of experts in the first expert layers except the M other first expert layers among the multiple other first expert layers. Among them, 2 ≤ M < (N - Q - K), and M is an integer.
[0081] Embodiment 2: Combining Figures 8 - 10 with FIG., for a sparse unit in which at least one second expert layer is connected between two adjacent first expert layers, the number of expert modules included in the K first expert layers close to the output layer of the sparse model is the same, and is greater than the number of expert modules included in the other first expert layers. The other first expert layers refer to the other first expert layers in the sparse unit except the K first expert layers. Among them, 2 ≤ K < (N - Q), and K is an integer, N is the total number of expert layers included in the sparse unit, and Q is the number of second expert layers included in the sparse unit. The number of the other first expert layers may be one or more.
[0082] For multiple other first expert layers, such as Figure 8 shown, among the multiple other first expert layers, the closer an other first expert layer is to the output layer of the sparse model, the more expert modules it contains. Or, as Figure 9 shown, among the multiple other first expert layers, the M other first expert layers close to the K first expert layers have the same number of experts, and this number is greater than the number of experts in the first expert layers among the multiple other first expert layers except for the M other first expert layers. Among them, 2 ≤ M < (N - Q - K), and M is an integer. Or, as Figure 10 shown, the number of expert modules contained in the multiple other first expert layers can be the same.
[0083] The above Figures 5 - 10 shown network structure of the sparse unit is only for illustrative purposes and does not constitute a limitation. No matter which network structure of the sparse unit is adopted Figures 2b - 10 shown, the network structure of any expert layer i of the dense unit can be used as an expert module and copied to the corresponding expert layer of the sparse model (that is, copied to the i-th expert layer of the sparse unit) to obtain the expert layer of the sparse unit, and at least one expert layer in the expert layer of the sparse unit includes multiple expert modules, so that the total number of expert modules contained in the sparse unit is greater than the total number of expert layers of the dense module. Since the expert module of the sparse unit has the same network structure as the expert layer of the dense unit, therefore, any of the above network structures of the sparse unit can make the model parameter quantity of the sparse unit greater than that of the dense unit, realizing the expansion of the parameter quantity of the expert layer of the dense unit.
[0084] It should be noted that Figures 2b - 10 is a simplified structural schematic diagram of the dense unit and the sparse unit. Only the dense unit and the sparse unit including expert layers are illustrated as examples, but it does not constitute a limitation. The dense unit also includes other network structures in addition to the expert layer, such as Figure 1 the multi-head attention layer, residual connection, and normalization layer of the dense unit in , etc. Therefore, other network structures of the sparse unit can also be constructed according to other network structures of the dense unit. Specifically, the other network structures of the dense unit can be copied to the sparse unit to obtain the other network structures of the sparse unit.
[0085] Furthermore, according to the connection relationship between the expert layer of the dense unit and other network structures of the dense unit, the expert layer of the sparse unit and other network structures of the sparse unit can be connected. For the expert layer of the sparse unit containing multiple expert modules, according to the working principle of the sparse unit, during the training or inference process of the sparse model, it is also necessary to select some expert modules from multiple expert modules in the same expert layer to process the input of the expert layer. Therefore, in this embodiment, a routing module (Router) can also be added before the expert layer of the sparse unit to obtain the sparse unit. Optionally, as Figures 2b - 10 shown, a routing module can be added before each expert layer of the sparse unit to obtain the sparse unit. Here, before each expert layer refers to the transmission direction of the input of the expert layer in the sparse unit. The input of the expert layer passes through the routing module first and then through the expert layer, which can be said that the routing module is before the expert layer.
[0086] Among them, the routing module can be used to select a target expert module from at least one expert module of the expert layer to process the input of the expert layer of the sparse unit. The target expert module is a part of the expert modules in the expert layer of the expert layer. For example, the target expert module can be an expert module in the expert layer. The routing module generally adopts a lightweight network structure. For example, the routing module can adopt an MLP. Optionally, the MLP can include: 2 fully connected layers and 1 classifier.
[0087] Based on the above-described specific implementation manner of expanding the number of parameters of the dense unit to obtain the sparse unit, the increased number of parameters ΔX of the sparse unit compared to the dense unit is: the increment (X1 - X2) of the total number X1 of the expert modules of the expert layer of the sparse unit compared to the number of layers X2 of the expert layer of the dense unit, multiplied by the number of parameters X0 of a single expert layer of the dense unit, plus the number of parameters of N routing modules. That is, ΔX = (X1 - X2) * X0 + N * Y0. Wherein, Y0 represents the total number of layers of the expert layer of the dense unit or the sparse unit; Y0 represents the number of parameters of a single routing module; ΔX represents the increased number of parameters ΔX of the sparse unit compared to the dense unit.
[0088] For example, as Figure 2b shown, each expert layer of the sparse unit includes 2 expert modules, and each expert module has the same network structure as the expert layer of the dense unit. Then for Figure 2b the sparse unit and the dense unit shown, X1 - X2 = N; the increased number of parameters ΔX of the sparse unit compared to the dense unit is ΔX = N * X0 + N * Y0. Since the number of parameters of the routing module is much smaller than the number of parameters of the expert module, therefore, Figure 2b the sparse model composed of the sparse unit shown has expanded the number of parameters by about 2 times compared to the target dense model.
[0089] In the embodiments of the present application, the routing module can be used to select a target expert module from at least one expert module in the expert layer to process the input of the expert layer of the sparse unit. The target expert module is a part of the expert modules in the expert layer. For example, the target expert module can be an expert module in the expert layer.
[0090] In the embodiments of the present application, the specific implementation manner of the routing module for selecting the target expert module from at least one expert module in the expert layer is not limited. In some embodiments, the routing module can use the Top-k strategy to select k expert modules from at least one expert module in the expert layer as the target expert modules. Wherein, k is an integer, and 1 ≤ k < P. P represents the total number of expert modules in the expert layer. Optionally, k = 1.
[0091] Specifically, the routing module can calculate the probability of each expert module in the expert layer i for processing the input of this expert layer; and according to the probability of each expert module in the expert layer i for processing the input of the expert layer i, select the top k expert modules in terms of probability ranking (sorted from large to small) from at least the expert modules in the expert layer i as the target expert modules. For example, when k = 1, the expert module with the highest probability can be selected from at least the expert modules in the expert layer i as the target expert module. In this way, during the subsequent model training process of the sparse model, only one expert module is selected from the expert layer for calculation each time, making the computational amount of the sparse model after expanding the number of parameters almost equal to that of the dense model during the training process, and thus making the training efficiency of the sparse model almost the same as that of the dense model.
[0092] For the implementation manner in which the routing module adopts the Top-1 strategy to select the expert module with the highest probability from at least one expert module in the expert layer as the target expert module, the routing module does not perform probability scaling on the probability of each expert module for processing the input of the expert layer i. Here, not performing probability scaling means directly using the probability calculated by the routing module for each expert module to process the input of the expert layer i to select the target expert module; instead of multiplying the calculated probability by a certain coefficient for scaling. In this way, the result in the initialization stage of the sparse model can be made consistent with the target dense model. For example, the expert layer i of the sparse unit includes 4 expert modules, and the 4 expert modules reuse the parameter information of the expert layer of the dense model during the sub-initialization stage. Therefore, the parameter values of the 4 expert modules in the expert layer i of the sparse unit are the same. Without probability scaling, routing the input of the expert layer i to any one of the expert modules will result in the same calculation result as the dense model. If the calculated probability is multiplied by a certain coefficient for scaling, the calculation result will be different from that of the dense model.
[0093] On the other hand, in the traditional Transformer architecture, the gating MLP (Switch MLP) layer scales the probability of each expert module in the Switch MLP layer for processing the input of expert layer i by multiplying it with max(p). Here, max(p) refers to the maximum value of the probabilities of each expert module for processing the input of expert layer i. In this embodiment, however, the probabilities of each expert module for processing the input of expert layer i are not scaled. Instead, by simply comparing the magnitudes of the probabilities corresponding to each expert module, the target expert layer with the highest probability can be determined. In the Switch MLP layer of the traditional Transformer architecture, it is not necessary to obtain the specific maximum probability value max(p), but more sample data is required for training, resulting in slower convergence of model training. Therefore, in this embodiment, since the probabilities of each expert module for processing the input of expert layer i are not scaled and there is no need to obtain the specific maximum probability value, the model convergence speed can be improved.
[0094] After expanding the number of parameters of the dense unit according to the network structure of the dense unit in step 203 above to obtain the sparse units of the sparse module, in step 204, the sparse units can be connected according to the connection relationship between the dense units to obtain the initial sparse model. Specifically, the sparse units can be connected according to the connection relationship between the dense units to obtain the initial sparse model. The connection relationship between the sparse units is the same as that between the dense units.
[0095] Since the sparse units of the sparse model are obtained by expanding the number of parameters of the dense unit according to the network structure of the dense unit, and the sparse units reuse the network structure of the dense unit, in step 205, the parameter information of the target dense model can be used to initialize the parameters of the initial sparse model to obtain the sparse model to be trained.
[0096] According to the above Figures 2b - 10 From the comparison diagram of the sparse unit and the dense unit shown, it can be seen that except for the routing module, the sparse unit reuses the network structure of the dense unit, and the expert modules of the expert layer of the sparse unit reuse the network structure of the expert layer of the dense unit. Therefore, step 205 above can be specifically implemented as follows: the parameter information of the expert layer of the dense unit and the parameter information of other network structures of the dense unit can be obtained from the parameter information of the target dense model. Further, the parameter information of the expert layer of the dense unit can be used to assign values to the parameters of the expert modules in the expert layer of the sparse unit to initialize the parameters of the expert layer of the sparse unit; and the parameter information of other network structures of the dense unit can be used to assign values to the parameters of other network structures of the sparse unit to initialize the parameters of other network structures of the sparse unit; and, the parameters of the routing module are randomly initialized to complete the parameter initialization of the initial sparse model.
[0097] In the embodiments of the present application, the network structure of the target dense model after pre-training or fine-tuning can be reused, and the number of parameters of the dense unit can be expanded based on the expert layer of the dense unit of the target dense model to obtain the sparse unit of the sparse model; and the initial sparse model can be initialized with the parameter information of the target dense model, reusing the training and learning results of the target dense model. Therefore, when training the sparse model to be trained, it is possible to directly fine-tune with the samples of the professional task without reusing a large amount of sample data for pre-training, which can reduce the training cost of the sparse model.
[0098] On the other hand, since the training and learning results of the target dense model are reused, when fine-tuning the sparse model with the samples of the professional task, it can converge quickly without reusing a large amount of sample data for pre-training. Therefore, compared with the training process of the existing sparse model, the convergence speed of the sparse model can be improved, that is, the training efficiency of the sparse model can be improved.
[0099] The following is an exemplary description of the training method for fine-tuning the above-mentioned sparse model to be trained on a professional task provided by the embodiments of the present application.
[0100] Figure 11 It is a schematic flowchart of the model training method provided by the embodiments of the present application. As Figure 11 shown, the model training method mainly includes:
[0101] 1101. Obtain the training sample pairs of the target task; each training sample pair includes: the first text and the second text corresponding to the first text in the target task.
[0102] 1102. Use the sparse model to be trained to predict the third text corresponding to the first text in the target task.
[0103] 1103. Adjust the model parameters of the sparse model to be trained according to the difference between the third text and the second text to obtain the target sparse model.
[0104] In this embodiment, the sparse model to be trained can be the sparse model constructed by using the model construction method shown in the foregoing embodiments. The target task refers to a professional task in a specific field. For example, tasks for medical diagnosis or medication recommendation in the medical field; or, tasks for article abstract summarization or article recommendation in the academic field; or, tasks for video recommendation in the video field; or, tasks for object recognition in the field of image processing, etc.
[0105] In this embodiment, the training sample pairs are the training sample pairs corresponding to the target task. Each training sample pair includes a first text and a second text corresponding to the first text in the target task. Among them, the first text can be used as the input of the sparse model to be trained; the second text is the theoretical output result corresponding to the first text during the supervised task learning process. For example, in a medical diagnosis task, the first text can be symptom description information; the second text can be diagnosis conclusion information. Another example is that in a summary task, the first text can be the content of an article, and the second text can be the summary of the article content, and so on.
[0106] In this embodiment, when the first text is input into the sparse model to be trained, the sparse model to be trained can be used to predict the third text corresponding to the first text in the target task, that is, the output of the sparse model to be trained is the third text.
[0107] For the above Figures 2b - 10 shown model architecture of the sparse model to be trained, the expert layer of the sparse unit includes at least one expert module; a routing module is connected before the expert layer. Then, in the process of using the sparse model to be trained to predict the third text corresponding to the first text in the target task, for any expert layer i of the sparse unit, the routing module can be used to select a target expert module from at least one expert module of the expert layer i. At least one expert module of the expert layer i refers to all expert modules included in the expert layer i, which may be 1 or multiple.
[0108] Specifically, in some embodiments, the routing module can be used to calculate the probability of each expert module in the expert layer i for processing the input of the expert layer; and according to the probability of each expert module in the expert layer i for processing the input of the expert layer i, select the top k expert modules with the highest probability (sorted from largest to smallest) from at least the expert modules in the expert layer i as the target expert modules. Where k is an integer, and 1 ≤ k < P. P represents the total number of expert modules in the expert layer.
[0109] Optionally, k = 1. Correspondingly, the expert module with the highest probability can be selected from at least the expert modules in the expert layer i as the target expert module. In this way, during the subsequent model training process of the sparse model, only one expert module is selected from the expert layer for calculation each time, so that the computational complexity of the sparse model after expanding the number of parameters is almost equal to that of the dense model training process, and thus the training efficiency of the sparse model is almost the same as that of the dense model.
[0110] Furthermore, the model parameters of the sparse model to be trained can be adjusted according to the difference between the third text and the second text to obtain the target sparse model.
[0111] Among them, the specific calculation method of the difference between the third text and the second text is determined by the loss function of the sparse model. In some embodiments, the loss function of the sparse model is expressed as the difference between the third text predicted by the model and the second text in the training sample pair. For example, the loss function of the sparse model can be the cross-entropy between the third text predicted by the model and the second text in the training sample pair; or, the loss function of the sparse model can be the distance between the third text predicted by the model and the second text in the training sample pair, such as the Euclidean distance or the cosine distance, etc. Or, the loss function of the sparse model can be the mean square error between the third text predicted by the model and the second text in the training sample pair, etc.
[0112] Correspondingly, adjusting the model parameters of the sparse model to be trained according to the difference between the third text and the second text can be implemented as adjusting the model parameters of the sparse model to be trained with the goal of minimizing the loss function to obtain the target sparse model.
[0113] In this embodiment, since the sparse model to be trained reuses the training and learning results of the target dense model. Therefore, when training the sparse model to be trained, it can be directly fine-tuned using the samples of the professional task, without the need to pre-train using a large amount of sample data again, which can reduce the training cost of the sparse model.
[0114] For Figures 2b - 10 For the shown sparse model, since only one expert module is used for calculation each time during the training of the sparse model to be trained, and the number of parameters of the routing module is much smaller than that of the expert module, the training efficiency of the sparse model constructed in the embodiments of the present application is the same as that of the dense network. The inventor of the present application used 8 graphics cards with a video memory of 80GB each to train the sparse model to be trained constructed in the embodiments of the present application and found that a sparse model with about 30 billion (30B) parameters can be trained at most, and the training speed is almost the same as that of a dense model with a scale of 13 billion (13B) parameters. When inferring using the target sparse model trained in the embodiments of the present application, compared with the dense model, only the computational amount of the N-layer routing module is increased. When inferring a sparse model with a scale of 30 billion (30B) parameters, it can still reach a speed of processing about 15 tokens per second.
[0115] Next, in combination with the scenarios of medical diagnosis tasks and abstract summary tasks, the model training method provided in the embodiments of the present application will be exemplarily described.
[0116] Figure 12 It is a schematic flowchart of another model training method provided in the embodiments of the present application. As Figure 12 shown, the model training method mainly includes:
[0117] 1201. Obtain training sample pairs for the medical selection task; each training sample pair includes: symptom description information and diagnostic conclusion information.
[0118] 1202. Use the sparse model to be trained to predict the diagnostic conclusion information corresponding to the symptom description information.
[0119] 1203. According to the difference between the diagnostic conclusion information included in the training sample pair and the predicted diagnostic conclusion information, adjust the model parameters of the sparse model to be trained to obtain a target sparse model for medical diagnosis.
[0120] Figure 13 It is a schematic flowchart of another model training method provided by an embodiment of the present application. As Figure 13 shown, the model training method mainly includes:
[0121] 1301. Obtain training sample pairs for the abstract summarization task; each training sample pair includes: article content and the abstract of the article content.
[0122] 1302. Use the sparse model to be trained to predict the abstract corresponding to the article content.
[0123] 1303. According to the difference between the abstract included in the training sample pair and the predicted abstract, adjust the model parameters of the sparse model to be trained to obtain a target sparse model for abstract summarization.
[0124] In Figure 12 and Figure 13 the model training methods shown, the sparse model to be trained can be the sparse model constructed by the model construction method of the foregoing embodiment. Using Figure 12 and Figure 13 the model training methods shown, the inventors of the present application tested and found that the diagnostic quality of the target sparse model obtained by using Figure 12 the model training method shown during medical diagnosis can be improved by more than 3% compared with the diagnostic quality of the dense model. The quality of the abstract summarized by the target sparse model obtained by using Figure 13 the model training method shown is also significantly improved compared with the quality of the abstract summarized by the dense model. Therefore, the sparse model constructed by the embodiments of the present application can achieve better effects in subsequent downstream task executions in professional fields, at least better than the dense model.
[0125] It should be noted that, in order to compare the effects of the sparse models provided in the above Figure 2b and Figure 5 , the inventors of the present application set Figure 2b and Figure 5 to rely on the same target dense model, and Figure 2band Figure 5 expert layers with the same number of layers, that is, there are N expert layers in total. Further, the number of Figure 5 expert modules can be set to be the same as that of Figure 2b . Among them, each expert layer of the sparse unit in Figure 2b includes 2 expert modules, that is, the sparse unit shown in Figure 2b . Compared with the target dense model on the Figure 2b or Figure 5 left, the total number of additional parameters of the sparse unit shown is ΔX = N*X0 + N*Y0. Correspondingly, Figure 5 for the sparse unit shown, compared with the target dense model on the Figure 2b or Figure 5 left, the total number of additional parameters is also ΔX = N*X0 + N*Y0. Among them, N represents the total number of layers of the expert layers of the dense unit or the sparse unit; Y0 represents the number of parameters of a single routing module.
[0126] Among them, Figure 5 by the method of increasing the number of parameters layer by layer, the number of expert modules of the 2 expert layers close to the output layer of the sparse model, that is, expert layer N and expert layer (N - 2), is greater than that of other expert layers with increased parameters, such as expert layer 2 and expert layer 4, etc. Among them, other expert layers with increased parameters can include 2 expert modules, and the number of expert modules of the 2 expert layers close to the output layer of the sparse model with increased parameters is specifically determined by the number of additional parameters of the sparse unit shown in Figure 2b . In short, ensure that the number of additional parameters of the sparse unit shown in Figure 5 is equal to the number of additional parameters of the sparse unit shown in Figure 2b . For example, when N = 16, for the sparse unit shown in Figure 2b , compared with the target dense model on the Figure 2b or Figure 5 left, the total number of additional parameters is ΔX = 16*X0 + N*Y0; if the number of additional parameters of the sparse unit shown in Figure 5 is to be equal to the number of additional parameters of Figure 2b , then the number of expert modules of the 2 expert layers close to the output layer of the sparse model with increased parameters can include 8 and 4 expert modules respectively, etc. The inventors of the present application found through testing that when the sparse units shown in Figure 2b and Figure 5 use the same number of parameters, the sparse model shown in Figure 5 can make more efficient use of the output of the expert layer close to the output of the sparse model, and has a better output effect compared with the sparse model shown in Figure 2b .
[0127] It should be noted that the execution entity of each step of the method provided in the above embodiments can be the same device, or the method can also be executed by different devices as the execution entity. For example, the execution entities of steps 201 and 202 can be device A; for another example, the execution entity of step 201 can be device A, and the execution entity of step 202 can be device B; and so on.
[0128] In addition, in some of the processes described in the above embodiments and the accompanying drawings, there are multiple operations that appear in a specific order. However, it should be clearly understood that these operations can be executed not in the order in which they appear in this text or in parallel. The serial numbers of the operations, such as 201, 202, etc., are only used to distinguish between different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes can include more or fewer operations, and these operations can be executed in sequence or in parallel.
[0129] Correspondingly, the embodiments of the present application also provide a computer-readable storage medium storing computer instructions. When the computer instructions are executed by one or more processors, one or more processors are caused to execute the steps in the model construction method and / or the model training method provided in the foregoing embodiments. For the specific implementation manners of the steps in the model construction method and / or the model training method, reference may be made to the relevant content of the foregoing embodiments, which will not be elaborated herein.
[0130] Figure 14 It is a schematic structural diagram of a computing device provided by an embodiment of the present application. As Figure 14 shown, the electronic device includes: a memory 14a and a processor 14b. Among them, the memory 14a is used to store computer programs.
[0131] The processor 14b is coupled to the memory 14a and is used to execute the computer program to execute the steps in the model construction method and / or the model training method provided in the foregoing embodiments. For the specific implementation manners of the steps in the model construction method and / or the model training method, reference may be made to the relevant content of the foregoing embodiments, which will not be elaborated herein.
[0132] In some alternative embodiments, as Figure 14 shown, the computing device may further include: a communication component 14c, a power supply component 14d, a display component 14e, an audio component 14f and other optional components. Figure 14 Only some components are schematically shown, which does not mean that the computing device must include Figure 14 all the components shown, nor does it mean that the computing device can only include Figure 14 the components shown.
[0133] In addition, Figure 14The components within the dashed box are optional components, rather than mandatory components, and can be determined according to the product form of the specific visual electronic device. The computing device in this embodiment can be implemented as a terminal device such as a desktop computer, a laptop computer, a mobile phone, or an Internet of Things device; it can also be various server devices such as a traditional server, a cloud server, or a server cluster.
[0134] In the embodiment of the present application, the memory is used to store computer programs and can be configured to store various other data to support operations on the device where it is located. Among them, the processor can execute the computer programs stored in the memory to implement corresponding control logic. The memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read only memory (EEPROM), electrically programmable read only memory (EPROM), programmable read only memory (PROM), read only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc.
[0135] In the embodiment of the present application, the processor can be any hardware processing device capable of executing the above method logic. Optionally, the processor can be a central processing unit (CPU), a graphics processing unit (GPU), or a microcontroller unit (MCU); it can also be a programmable device such as a field-programmable gate array (FPGA), a programmable array logic device (PAL), a general array logic device (GAL), or a complex programmable logic device (CPLD); or an advanced reduced instruction set (RISC) processor (Advanced RISC Machines, ARM) or a system on chip (SoC), etc., but not limited thereto.
[0136] In an embodiment of the present application, the communication component is configured to facilitate communication between the device where it is located and other devices in a wired or wireless manner. The device where the communication component is located can access a wireless network based on communication standards, such as Wireless Fidelity (WiFi), 2G or 3G, 4G, 5G, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component can also be implemented based on Near Field Communication (NFC) technology, Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wide Band (UWB) technology, Bluetooth (BT) technology, or other technologies.
[0137] In an embodiment of the present application, the display component may include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the display component includes a touch panel, the display component can be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation.
[0138] In an embodiment of the present application, the power component is configured to provide power to various components of the device where it is located. The power component may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the device where the power component is located.
[0139] In an embodiment of the present application, the audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC). When the device where the audio component is located is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive external audio signals. The received audio signals can be further stored in a memory or transmitted via the communication component. In some embodiments, the audio component further includes a speaker for outputting audio signals. For example, for a device with a language interaction function, voice interaction with a user can be implemented through the audio component.
[0140] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0141] It should also be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types.
[0142] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, read-only compact disc (CD-ROM), optical storage, etc.) that contain computer-usable program code.
[0143] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, as well as the combination of flows and / or blocks in the flowchart and / or block diagram. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or multiple flows and / or blocks
[0144] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implement the functions specified in Figure 1 one or more of the flows Figure 1 or multiple flows and / or blocks
[0145] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide for implementing the process Figure 1 in one process or multiple processes and / or blocks Figure 1 steps of the functions specified in one block or multiple blocks.
[0146] In a typical configuration, a computing device includes one or more processors (such as CPUs), input / output interfaces, network interfaces, and memory.
[0147] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.
[0148] The storage medium of a computer is a readable storage medium, also known as a readable medium. Readable storage media include permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media do not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0149] It should also be noted that the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent in such process, method, commodity or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, commodity or device including the above elements.
[0150] The above content is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A method for constructing a model, characterized in that, it includes: Obtain a target dense model and parameter information of the target dense model; The target dense model is a pre-trained dense model or a fine-tuned dense model; From the network structure of the target dense model, obtain the network structure of the dense units of the target dense model and the connection relationship between the dense units; On the basis of the network structure of the dense units, expand the number of parameters of the dense units to obtain the sparse units of the sparse model; The number of the sparse units is the same as that of the dense units; According to the connection relationship between the dense units, connect the sparse units to obtain an initial sparse model; Use the parameter information of the target dense model to initialize the parameters of the initial sparse model to obtain a sparse model to be trained.
2. The method according to claim 1, characterized in that, The expanding the number of parameters on the basis of the network structure of the dense units to obtain the sparse units of the sparse model includes: On the basis of the network structure of the dense units, according to the network structure of the expert layer of the dense units, expand the number of parameters of the expert layer of the dense units; Add a routing module before the expert layer of the dense units after expanding the number of parameters to obtain the sparse units; wherein, the expert layer of the dense units after expanding the number of parameters is the expert layer of the sparse units.
3. The method according to claim 2, characterized in that, The expanding the number of parameters of the expert layer of the dense units according to the network structure of the expert layer of the dense units on the basis of the network structure of the dense units includes: From the network structure of the dense units, obtain the network structure of the expert layer of the dense units, other network structures of the dense units, and the connection relationship between the expert layer of the dense units and other network structures of the dense units; According to other network structures in the network structure of the dense units, construct other network structures of the sparse units; According to the network structure of the expert layer of the dense units, expand the number of parameters of the expert layer of the dense units to obtain the expert layer of the sparse units; each expert layer of the sparse units includes at least one expert module, and the network structure of the expert module is the same as that of the expert layer of the dense units; the total number of expert modules of the sparse units is greater than the total number of expert layers of the dense units; According to the connection relationship between the expert layer of the dense units and other network structures of the dense units, connect the expert layer of the sparse units and other network structures of the sparse units.
4. The method according to claim 3, characterized in that, The routing module is used to select a target expert module from the at least one expert module to process the input of the expert layer of the sparse units.
5. The method according to claim 3, characterized in that, The expanding the number of parameters of the expert layer of the dense units according to the network structure of the expert layer of the dense units to obtain the expert layer of the sparse units includes: For any expert layer of the dense unit, obtain the network structure of any expert layer of the dense unit from the network structure of the expert layers of the dense unit; Use the network structure of any expert layer of the dense unit as the expert module and copy it to the corresponding expert layer of the sparse unit to obtain the expert layer of the sparse unit; wherein, at least one expert layer in the expert layer of the sparse unit includes multiple expert modules; the number of expert layers of the sparse unit is the same as the number of expert layers of the dense unit.
6. The method according to claim 5, wherein, the sparse unit includes multiple expert layers, and each expert layer includes multiple expert modules; the number of expert modules included in the multiple expert layers of the sparse unit is the same; or, the number of expert modules included in some expert layers of the sparse unit is the same; or, the number of expert modules included in the multiple expert layers of the sparse unit is different from each other.
7. The method according to claim 6, wherein, in the case where the number of expert modules included in some expert layers of the sparse unit is the same, for the other expert layers except the part of the expert layers with the same number of included expert modules, the closer the expert layer is to the output layer of the sparse model, the more expert modules it includes; or, in the case where the number of expert modules included in the multiple expert layers of the sparse unit is different from each other, the closer the expert layer is to the output layer of the sparse model, the more expert modules it includes.
8. The method according to claim 5, wherein, the sparse unit includes multiple expert layers, the first expert layer in the multiple expert layers includes multiple expert modules, and the second expert layer in the multiple expert layers includes a single expert module; at least one second expert layer is connected between two adjacent first expert layers.
9. The method according to claim 8, wherein, the number of expert modules included in the K first expert layers close to the output layer of the sparse model is the same and greater than the number of expert modules included in the other first expert layers; or, the number of expert modules included in the K first expert layers is different, and the closer the first expert layer is to the output layer of the sparse model in the K first expert layers, the more expert modules it includes; the number of expert modules included in the K first expert layers is greater than the number of expert modules included in the other first expert layers; wherein, 2≤K<(N-Q), and K is an integer, N is the total number of expert layers included in the sparse unit, and Q is the number of the second expert layers.
10. The method according to claim 9, wherein, the other first expert layers are multiple; the number of expert modules included in the multiple other first expert layers is the same; or, the closer the other first expert layer is to the output layer of the sparse model, the more expert modules it includes; or, the number of expert modules included in the M other first expert layers close to the K first expert layers in the multiple other first expert layers is the same and greater than the number of experts included in the first expert layers except the M other first expert layers in the multiple other first expert layers; where 2 ≤ M < (N - Q - K) and M is an integer.
11. The method according to claim 5, wherein, the parameter initialization of the initial sparse model by using the parameter information of the target dense model to obtain the sparse model to be trained includes: obtaining the parameter information of the expert layer of the dense unit and the parameter information of other network structures of the dense unit from the parameter information of the target dense model; using the parameter information of the expert layer of the dense unit to assign values to the parameters of the expert module in the expert layer of the sparse unit to perform parameter initialization on the expert layer of the sparse unit; using the parameter information of other network structures of the dense unit to assign values to the parameters of other network structures of the sparse unit to perform parameter initialization on other network structures of the sparse unit; randomly initializing the parameters of the routing module.
12. The method according to any one of claims 3 - 11, wherein, the routing module is used to calculate the probability that at least one expert module processes the input of the expert layer of the sparse unit; and select, from the at least one expert module, the expert module with the highest probability of processing the input of the expert layer of the sparse unit as the target expert module.
13. The method according to any one of claims 3 - 11, wherein, further includes: obtaining a training sample pair of the target task; each training sample pair includes: a first text and a second text corresponding to the first text in the target task; using the sparse model to be trained to predict a third text corresponding to the first text in the target task; adjusting the model parameters of the sparse model to be trained according to the difference between the third text and the second text to obtain a target sparse model; wherein, in the process of using the sparse model to be trained to predict the third text corresponding to the first text in the target task, for any expert layer in the sparse unit, using the routing module to select a target expert module from at least one expert module of the any expert layer; using the target expert module to process the input of the any expert layer; the input of the any expert layer is the feature information of the first text.
14. A model training method, wherein, includes: obtaining a training sample pair of the target task; each training sample pair includes: a first text and a second text corresponding to the first text in the target task; using a sparse model to be trained to predict a third text corresponding to the first text in the target task; adjusting the model parameters of the sparse model to be trained according to the difference between the third text and the second text to obtain a target sparse model; Among them, the sparse model to be trained is obtained by initializing the parameters of the initial sparse model using the parameter information of the target dense model; the initial sparse model is obtained by expanding the number of parameters based on the network structure of the dense units of the target dense model; the target dense model is a pre-trained dense model or a fine-tuned dense model.
15. The method according to claim 14, wherein, the sparse units of the sparse model to be trained include at least one expert layer; the expert layer of the sparse unit includes: at least one expert module; the expert module has the same network structure as the expert layer of the dense unit; a routing module is provided before the expert layer of the sparse unit; in the process of using the sparse model to be trained to predict the third text corresponding to the first text in the target task, it includes: for any expert layer in the sparse unit, using the routing module to select a target expert module from at least one expert module of the any expert layer; using the target expert module to process the input of the any expert layer; the input of the any expert layer is the feature information of the first text.
16. A computing device, wherein, it includes: a memory and a processor; among them, the memory is used to store a computer program; the processor is coupled to the memory and is used to execute the computer program to execute the steps in the method according to any one of claims 1-15.
17. A computer-readable storage medium storing computer instructions, wherein, when the computer instructions are executed by one or more processors, the one or more processors are caused to execute the steps in the method according to any one of claims 1-15.