Natural language processing method, device and equipment for large language model knowledge distillation
By aggregating and reducing the intermediate layer features of the teacher model, constructing an alignment loss function, and conducting supervised fine-tuning and reinforcement learning, the student model is trained to learn the reasoning logic and decision-making path of the teacher model, thereby improving the accuracy of natural language processing and reducing resource consumption.
Patent Information
- Application Number
- CN202510865556.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2045-06-26
AI Technical Summary
The student model obtained by the existing large language model staged knowledge distillation method has low accuracy in natural language processing and poor generalization ability, and cannot effectively learn the reasoning logic and decision-making path of the teacher model.
By aggregating and reducing the intermediate layer features of the teacher model, aligning them with the intermediate layer features of the student model, constructing a feature alignment loss function, and combining supervised fine-tuning and reinforcement learning, the student model is trained so that it not only imitates the output results of the teacher model, but also learns its reasoning logic and decision-making path.
It improves the accuracy of the student model in natural language processing, reduces resource consumption, solves the problem of insufficient generalization ability of the student model, and significantly improves performance in resource-constrained scenarios.
Smart Images

Figure CN120781049A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of knowledge distillation, and in particular to a natural language processing method, device and equipment for large language model knowledge distillation. BACKGROUND
[0002] Large language models (such as GPT-4, DeepSeek-R1, etc.) exhibit breakthrough performance in open-domain question answering, code generation and complex reasoning tasks through the complex architecture of hundreds of billions of parameters. However, the large size of such models results in the consumption of hundreds of GB of video memory and tens of thousands of floating-point operations per inference, which seriously restricts their application in resource-constrained scenarios such as mobile terminals and industrial Internet of Things devices. Knowledge distillation technology can transfer the knowledge of a teacher model to a lightweight student model to balance the contradiction between performance and resource consumption.
[0003] The existing stage-based knowledge distillation method includes a supervised fine-tuning (SFT) stage and a reinforcement learning (RL) stage. The supervised fine-tuning stage uses the teacher model to generate high-quality data (such as mathematical problem solving steps and code completion examples) to perform end-to-end supervised training of the student model. The reinforcement learning stage optimizes the student strategy distribution based on an artificially designed reward function (such as an answer correctness score and a code execution pass rate). The existing method uses the output layer logits or the final answer of the teacher model as a supervision signal, which results in the student model only being able to imitate the "result" of the teacher, poor generalization ability, and low accuracy of natural language processing. SUMMARY
[0004] The embodiments of the present application provide a natural language processing method, device and equipment for large language model knowledge distillation to solve the problem of low accuracy of natural language processing of the student model obtained by the existing stage-based knowledge distillation method of the large language model.
[0005] In a first aspect, the embodiments of the present application provide a natural language processing method for large language model knowledge distillation, comprising: obtaining an initial student model, wherein the number of layers and the dimension of the student model are smaller than those of a teacher model; and the teacher model is a large language model;
[0006] The intermediate layer features of the teacher model are aggregated and reduced in dimension to align with the number of layers and the dimension of the intermediate layers of the student model, to obtain aligned intermediate layer features of the teacher model;
[0007] Based on the difference between the aligned intermediate layer features of the teacher model and the corresponding intermediate layer features of the student model, a feature alignment loss is obtained to construct a first loss function;
[0008] The initial student model is iteratively trained based on the constructed first loss function to obtain a supervised fine-tuned student model;
[0009] Perform reinforcement learning training on the supervised fine-tuned student model to obtain the student model after knowledge distillation;
[0010] Based on the student model after knowledge distillation, the natural language text data to be analyzed is processed.
[0011] In one possible implementation, the intermediate layer features of the teacher model are aggregated and dimensionally reduced to align with the number of layers and dimensions of the intermediate layers of the student model. The aligned intermediate layer features of the teacher model are obtained, including:
[0012] Divide the intermediate layers of the teacher model into L S Group, where L S is the number of intermediate layers of the student model;
[0013] For any set of intermediate layer features of the teacher model, gated attention aggregation is used to obtain the aggregated intermediate layer features;
[0014] For any aggregated intermediate layer feature, a low-rank adapter is used to map the high-dimensional aggregated intermediate layer to the student model dimension, and the intermediate layer feature that is aligned with the number of layers and dimensions of the student model intermediate layer is obtained as the aligned intermediate layer feature of the teacher model.
[0015] In a possible implementation, the intermediate layers of the teacher model are divided into L S Groups include:
[0016] Using the sliding window grouping method, the middle layers of the teacher model are divided into L S groups; each group contains a continuous number of layers Among them, L T represents the number of intermediate layers of the teacher model, w represents the number of layers in each group; sliding step
[0017] In one possible implementation, for any set of intermediate layer features of the teacher model, gated attention aggregation is used to obtain aggregated intermediate layer features including:
[0018] The following formula is used to obtain the characteristics of the intermediate layer after polymerization:
[0019]
[0020] in, Represents the characteristics of the middle layer after the aggregation of each layer in group g; α g,i represents the attention weight of the i-th layer of the g-th group; represents the intermediate layer features of the i-th layer of the g-th group; Represents the trainable parameters of the student model; represents average pooling.
[0021] In a possible implementation, for any aggregated intermediate layer feature, a low-rank adapter is used to map the high-dimensional aggregated intermediate layer to the student model dimension, to obtain an intermediate layer feature aligned with the number of layers and the dimension of the student model intermediate layer as the aligned intermediate layer feature of the teacher model, including:
[0022] The aligned intermediate layer feature of the teacher model is obtained based on the following formula:
[0023]
[0024] wherein, represents the aligned intermediate layer feature; A g and B g represent low-rank matrices; r = L S r represents the number of layers of the student model.
[0025] In a possible implementation, the feature alignment loss is obtained based on the difference between the aligned intermediate layer feature of the teacher model and the corresponding student model intermediate layer feature, including:
[0026] The feature alignment loss is obtained based on the following formula:
[0027]
[0028] wherein, represents the feature alignment loss; S g represents the feature of the gth layer of the student model; γ g represents the weight of the gth layer, wherein the deeper the layer, the greater the weight; d S represents the dimension of the student model.
[0029] In a possible implementation, the reinforcement learning training of the supervised fine-tuned student model includes:
[0030] The KL divergence loss is obtained based on the KL divergence between the output probability distribution of the teacher model and the output probability distribution of the student model;
[0031] The second loss function is constructed based on the KL divergence loss and the feature alignment loss;
[0032] The reinforcement learning training of the supervised fine-tuned student model is performed based on the second loss function and the reward of the preset rule.
[0033] In a possible implementation, the second loss function further includes a grouping relative policy optimization loss and a Logit distribution alignment loss;
[0034] wherein, the second loss function is obtained based on the following formula:
[0035] LRL =λ1·L GRPO +λ2·L align +λ3·L KL +λ4·L logit
[0036] Among them, L RL Represents the second loss function; L GRPO represents the group relative strategy optimization loss; L align represents feature alignment loss; L KL represents KL divergence loss; L logit represents the Logit distribution alignment loss; λ1, λ2, λ3, and λ4 represent the weights of each loss term.
[0037] In a second aspect, an embodiment of the present invention provides a natural language processing device for large language model knowledge distillation, comprising: an acquisition module for acquiring an initial student model, wherein the student model has fewer layers and dimensions than a teacher model; the teacher model is a large language model;
[0038] The alignment module is used to aggregate and reduce the dimensionality of the intermediate layer features of the teacher model, align the number of layers and dimensions with the intermediate layers of the student model, and obtain the aligned intermediate layer features of the teacher model;
[0039] A loss function construction module is used to obtain a feature alignment loss based on the difference between the intermediate layer features of the teacher model after alignment and the intermediate layer features of the corresponding student model, thereby constructing a first loss function;
[0040] A first training module is configured to perform supervised fine-tuning iterative training on the initial student model based on the constructed first loss function to obtain a supervised fine-tuned student model;
[0041] The second training module is used to perform reinforcement learning training on the student model after supervised fine-tuning to obtain the student model after knowledge distillation;
[0042] The processing module is used to process the natural language text data to be analyzed based on the student model after knowledge distillation.
[0043] In a third aspect, an embodiment of the present invention provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the method in the first aspect or any possible implementation of the first aspect is implemented.
[0044] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the method in the first aspect or any possible implementation of the first aspect.
[0045] In a fifth aspect, an embodiment of the present invention provides a computer program product, including a computer program, which, when executed by a processor, implements the method in the first aspect or any possible implementation of the first aspect.
[0046] The embodiment of the present invention aggregates and reduces the intermediate layer features of the teacher model during the supervised fine-tuning stage, aligns them with the intermediate layer features of the student model, and dynamically maps the intermediate layer features of the teacher model to the intermediate layers of the student model. After the intermediate layer features are aligned, a loss function is constructed based on the difference in intermediate layer features between the teacher model and the student model to train the student model. As a result, the trained student model can not only imitate the output results of the teacher model, but also learn deep features such as the teacher model's reasoning logic and decision-making path. The student model can learn feature information at different levels of the teacher model, so as to better understand and imitate the teacher model's reasoning process, thereby improving the accuracy of the student model's natural language processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 A diagram illustrating an application scenario of the natural language processing method for large language model knowledge distillation provided by an embodiment of the present invention;
[0048] Figure 2 This is a flow chart for implementing a natural language processing method for large language model knowledge distillation provided by an embodiment of the present invention;
[0049] Figure 3 is a schematic diagram of a dynamic adaptation process provided by an embodiment of the present invention;
[0050] Figure 4 This is a diagram of the system architecture for the supervision and fine-tuning phase provided by an embodiment of the present invention;
[0051] Figure 5 Schematic diagram of the supervised fine-tuning training process provided by an embodiment of the present invention;
[0052] Figure 6 This is a diagram of the system architecture of the reinforcement learning stage provided by an embodiment of the present invention;
[0053] Figure 7 This is a schematic diagram of the reinforcement learning training process provided by an embodiment of the present invention;
[0054] Figure 8 This is a schematic diagram of the structure of a natural language processing device for large language model knowledge distillation provided by an embodiment of the present invention;
[0055] Figure 9 is a schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0056] The embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0057] Figure 1 The application scenario diagram of the natural language processing method of large language model knowledge distillation provided by the embodiments of the present application is shown as follows. Figure 1 As shown in the figure, knowledge distillation is a technology that transfers knowledge from a larger and more complex model (such as a teacher model) to a smaller and simpler model (such as a student model), aiming to reduce the size and computational cost of the model while maintaining the performance of the model, so that the model is easier to deploy and apply.
[0058] In knowledge distillation, the teacher model is a well-trained model with high performance but may be complex and large, usually having more parameters, deeper network structure or stronger representation ability, which can accurately predict and understand the input data. As the provider of knowledge, the output of the teacher model is used as the target for the student model to learn. The student model is the target model of knowledge distillation, which is a relatively smaller and simpler model, usually having fewer parameters and simpler structure than the teacher model.
[0059] The mainstream large language model knowledge distillation method includes a two-stage method: first, supervised fine-tuning, and then reinforcement learning. Although the two-stage method has achieved certain results, experimental analysis and theoretical verification still have the following key defects: the limitations of black-box knowledge transfer.
[0060] 1. Surface knowledge transfer: existing methods only use the output layer logits or final answer of the teacher model as a supervision signal, ignoring the reasoning logic and decision path contained in the intermediate layer features (such as attention weights, hidden states) of the teacher model, resulting in the student model only being able to imitate the “result” of the teacher, but not being able to learn its “thinking process”;
[0061] 2. Feature expression ability gap: when the high-dimensional feature space (such as 1024-dimensional hidden layer) of the teacher model is directly aligned with the low-dimensional representation (such as 512-dimensional) of the student model, due to the difference in dimension and semantic complexity, feature mismatch and signal distortion occur.
[0062] Therefore, the existing method uses the output layer logits or final answer of the teacher model as a supervision signal, resulting in the student model only being able to imitate the “result” of the teacher, having poor generalization ability and low accuracy.
[0063] During the supervised fine-tuning phase, this embodiment of the present invention uses a multi-level feature mapping mechanism, known as a many-to-one hierarchical mapping, to aggregate the deep features of the teacher model and align them with the shallow features of the student model. Furthermore, the teacher's high-dimensional features are reduced to a low-dimensional space understandable to the student, alleviating the feature expression gap. As a result, the trained student model not only mimics the teacher model's output but also learns its deeper features, such as its reasoning logic and decision paths, improving the student model's natural language processing accuracy.
[0064] Figure 2 This is a flow chart of a natural language processing method for large language model knowledge distillation provided by an embodiment of the present invention. Figure 2 The embodiment of the present invention provides a natural language processing method for knowledge distillation of a large language model, including:
[0065] Step 201: Obtain an initial student model, wherein the number of layers and dimensions of the student model are smaller than those of the teacher model; the teacher model is a large language model;
[0066] For example, the number of layers refers to the number of neural network layers in the model. For example, a neural network model typically includes an input layer, multiple hidden layers, and an output layer. A smaller number of layers indicates a relatively simple model structure.
[0067] For example, dimension refers to the number of neurons in a model's layers. For example, the number of neurons in a single hidden layer is the hidden layer dimension. Smaller dimensions mean a model has a weaker ability to process information, but they also mean it requires fewer computing resources and may run faster.
[0068] Step 202: Aggregate and reduce the dimension of each intermediate layer feature of the teacher model, align the number of layers and dimensions with the intermediate layer of the student model, and obtain the aligned intermediate layer features of the teacher model;
[0069] It's important to note that the teacher model, as a large language model, contains multiple intermediate layers within its neural network architecture. Each intermediate layer extracts and represents feature information at a different level when processing input data. For example, when processing text data, lower intermediate layers may extract features at the word or phrase level, while higher intermediate layers may extract more abstract features at the sentence or semantic level.
[0070] For example, the features of each intermediate layer of the teacher model are aggregated: the large amount of feature data in each intermediate layer is merged or aggregated according to certain rules. For example, operations such as summation, averaging, maximum pooling, and minimum pooling can be used. This results in a more representative comprehensive value, thereby reducing the number and complexity of features. Furthermore, the number of layers in the aggregated teacher model is the same as that of the student model, and they correspond one-to-one, that is, they are aligned with the number of layers in the intermediate layers of the student model.
[0071] For example, dimensionality reduction further reduces the dimensionality of features based on aggregation. Since the teacher model typically has a high dimensionality, directly aligning its features with the student model may result in excessive computational overhead or be unmanageable for the student model. Dimensionality reduction allows the high-dimensional features of the teacher model's intermediate layers to be mapped into a lower-dimensional space that matches the dimensions of the student model's intermediate layers, while preserving as much important information as possible from the original features.
[0072] Therefore, after aggregation and dimensionality reduction, the intermediate layer features of the teacher model are aligned with those of the student model in terms of number of layers and dimensions. This alignment preserves the key knowledge and feature information of the teacher model while also matching the form of the student model.
[0073] The following examples illustrate the methods of aggregation and dimensionality reduction.
[0074] In one possible implementation, the intermediate layer features of the teacher model are aggregated and dimensionally reduced to align with the number of layers and dimensions of the intermediate layers of the student model. The aligned intermediate layer features of the teacher model are obtained, including:
[0075] Step 2021, divide each middle layer of the teacher model into L S Group, where L S is the number of intermediate layers of the student model.
[0076] In some embodiments, the intermediate layers of the teacher model are divided into L S Group includes: Using sliding window grouping method, the middle layers of the teacher model are divided into L S groups; each group contains a continuous number of layers Among them, L T represents the number of intermediate layers of the teacher model, w represents the number of layers in each group; sliding step in, Indicates the floor symbol.
[0077] For example, the teacher model has 24 middle layers (L T =24), the middle layer of the student model is 8 layers (L S =8), then w = 6, s = 3, generating the group:
[0078] Group 1:Layer 1-6→Student Layer 1;
[0079] Group 2:Layer 4-9→Student Layer 2; ...
[0081] Group 8:Layer 19-24→Student Layer 8.
[0082] In step 2022, for any set of intermediate layer features of the teacher model, gated attention aggregation is used to obtain aggregated intermediate layer features.
[0083] In some embodiments, the aggregating any set of intermediate layer features of the teacher model using gated attention to obtain aggregated intermediate layer features includes: obtaining the aggregated intermediate layer features using the following formula:
[0084]
[0085] in, Represents the characteristics of the middle layer after the aggregation of each layer in group g; α g,i represents the attention weight of the i-th layer of the g-th group; represents the intermediate layer features of the i-th layer of the g-th group; Represents the trainable parameters of the student model; represents average pooling.
[0086] This embodiment of the present invention aggregates each set of teacher features using learnable attention weights. This approach automatically identifies and highlights the features most relevant to the task at hand, allowing the student model to focus on key information while removing interference from irrelevant or minor features. This weighted aggregation integrates the dominant elements of multiple layers of teacher features, resulting in a more precise feature representation.
[0087] In step 2023, for any aggregated intermediate layer feature, a low-rank adapter is used to map the high-dimensional aggregated intermediate layer to the student model dimension, and an intermediate layer feature that is aligned with the number of layers and dimensions of the student model intermediate layer is obtained as the aligned intermediate layer feature of the teacher model.
[0088] It should be noted that the dimensions of the aggregated intermediate-layer features are the same as those of the teacher model, which is higher. The low-rank adapter is used to transform the aggregated intermediate-layer features and perform feature projection. Specifically, the teacher's high-dimensional features (e.g., 1024-dimensional) are mapped to the student space (e.g., 512-dimensional) using the low-rank adapter (LoRA). The aligned intermediate-layer features of the teacher model have the same dimensions as those of the student model.
[0089] The embodiments of the present application reduce the number of parameters by using low-rank adapter conversion. For example, the number of parameters is reduced from 1024*512=524288 to 8*(1024+512)=12288, a reduction of 97%.
[0090] In some embodiments, for any post-aggregation intermediate layer feature, a low-rank adapter is used to map the high-dimensional post-aggregation intermediate layer to the student model dimension, to obtain an intermediate layer feature aligned with the number of layers and dimensions of the student model intermediate layer as the aligned intermediate layer feature of the teacher model, including:
[0091] The aligned intermediate layer feature of the teacher model is obtained based on the following formula:
[0092]
[0093] wherein, represents the aligned intermediate layer feature; A g and B g represent a low-rank matrix; r=L S r represents the number of layers of the student model.
[0094] Figure 3 is a dynamic adaptation processing flowchart provided by the embodiments of the present application; refer to Figure 3 For example, the teacher model has 24 layers and the student model has 8 layers. Each layer of the teacher model is grouped by a sliding window, aggregated by a gate, and then mapped to each layer of the student model by a low-rank adapter.
[0095] The embodiments of the present application use adapter-driven feature transformation, introduce a lightweight adapter to reduce the dimensionality of the teacher high-dimensional feature to a low-dimensional space understandable by the student model, and solve the problem of difference in feature expression ability between the teacher model and the student model.
[0096] In step 203, the feature alignment loss is obtained based on the difference between the aligned intermediate layer feature of the teacher model and the corresponding intermediate layer feature of the student model, and the first loss function is constructed therefrom.
[0097] In some embodiments, the first loss function further includes a Logit alignment loss and a cross-entropy loss.
[0098] For example, wherein, represents the first loss function; represents the feature alignment loss; represents the Logit alignment loss; represents the cross-entropy loss. λ1 to λ3 represent the weights of each loss term. For example, λ1=0.4, λ2=0.4, and λ3=0.2.
[0099] In this embodiment of the present invention, the first loss function uses a hybrid supervision approach that uses a weighted sum of multiple loss terms. The student model, trained based on hybrid supervision, not only mimics the output of the teacher model through logit alignment loss and cross-entropy loss, but also learns deep features of the teacher model, such as its reasoning logic and decision-making path, through feature alignment loss, improving the student model's accuracy.
[0100] Exemplarily, the feature alignment loss obtained based on the difference between the intermediate layer features of the teacher model after alignment and the intermediate layer features of the corresponding student model includes:
[0101] The feature alignment loss is obtained based on the following formula:
[0102]
[0103] in, represents feature alignment loss; S g represents the features of the g-th layer of the student model; γ g represents the weight of the g-th layer; d S Represents the dimension of the student model.
[0104] For example, the Logit alignment loss is obtained based on the following formula:
[0105]
[0106] where the temperature scales the KL divergence (initial τ = 5.0, annealed to τ = 1.0).
[0107] For example, the cross entropy loss is obtained based on the following formula:
[0108]
[0109] Figure 4 This is a diagram of the system architecture of the supervision and fine-tuning phase provided by an embodiment of the present invention. Figure 4 , the teacher model and the student model are fed with the same labeled data. It should be noted that since the teacher model is a pre-trained model, it does not need to be retrained during knowledge distillation, so the parameters of the teacher model are frozen. The student model, on the other hand, is trainable.
[0110] The total loss includes cross entropy loss, Logit alignment loss, and feature alignment loss.
[0111] About cross entropy loss: The output Logits of the student model is output through the Logits distributor to obtain the cross entropy loss.
[0112] About Logit alignment loss: The student model outputs Logits and outputs them through the Logits distributor, and together with the output Logits of the teacher model, we get the Logit alignment loss.
[0113] Regarding the feature alignment loss: On the one hand, the student model outputs intermediate-level features. On the other hand, the teacher model outputs intermediate-level features, which are grouped and aggregated by the sliding window grouper and then transformed by the low-rank adapter to obtain intermediate-level features aligned with the student model. This results in the feature alignment loss. It should be noted that the low-rank adapter is trainable. During training, the specific parameters of the low-rank matrix in the low-rank adapter are updated. The following describes the supervised fine-tuning training steps for the student model.
[0114] Step 204: Perform supervised fine-tuning iterative training on the initial student model based on the constructed first loss function to obtain a supervised fine-tuned student model;
[0115] For example, in each iteration, the model calculates a loss value based on the current parameters. Backpropagation is then used to calculate the gradient of the loss value with respect to the model parameters. The model parameters are then updated based on the gradient. Through continuous iteration, the model parameters are gradually optimized, the loss value is gradually reduced, and the model performance continues to improve.
[0116] Figure 5 This is a schematic diagram of the supervised fine-tuning training process provided by an embodiment of the present invention; Figure 5 , first initialize the student model and adapter. Then, input the same labeled data into the student model and teacher model respectively. The intermediate layer features of the teacher model are grouped, aggregated and processed by the adapter to obtain intermediate features. The intermediate features and the Logits of the output layer are used together as supervision signals to construct the first loss function. The student model continuously iterates and updates parameters based on the first loss function until convergence and the final model is obtained. It should be noted that when the student model updates parameters and re-predicts, its intermediate features and output layer Logits are changing, and thus the value of the first loss function is also continuously updated.
[0117] The above steps obtain the supervised fine-tuned student model, and the following steps perform reinforcement learning training.
[0118] Step 205: Perform reinforcement learning training on the supervised fine-tuned student model to obtain a student model after knowledge distillation.
[0119] Traditional reinforcement learning relies on manually designed reward functions (e.g., binary scoring reward rules for correct answers). For example, rewards are calculated for the student model's predictions based on a pre-set reward rule. The student model is then iteratively trained based on the rewards to maximize them.
[0120] Traditional rule-based rewards are one-sided and fail to capture the teacher model's implicit evaluation criteria for decision-making process quality (e.g., rationality of reasoning steps, maintainability of code), leading to student strategies overfitting to specific scoring rules. For example, traditional reinforcement learning relies on artificial rule-based rewards (e.g., scoring for correct answers), which can easily lead to the student model generating repetitive or redundant content through "reward cheating."
[0121] In addition, the lack of teacher guidance causes the student model to optimize independently under pure reward drive. Without real-time feedback from the teacher model on its exploration direction, it is easy to fall into local optimal solutions, such as generating repeated and redundant content to defraud rewards.
[0122] The following example illustrates improved reinforcement learning training. Exemplarily, prior to reinforcement learning training, dynamic data generation and reward calculation are also included. For example, online sampling involves the student model generating five candidate answers {yi1, yi2, ..., yi5} for each input xi; kernel sampling is used to control diversity (p = 0.9). Reward calculation can utilize manually designed rule-based rewards.
[0123] In one possible implementation, performing reinforcement learning training on the supervised fine-tuned student model includes:
[0124] Step 2051, obtaining a KL divergence loss based on the KL divergence of the teacher model output probability distribution and the student model output probability distribution;
[0125] For example, through real-time policy alignment, during the reinforcement learning phase, the teacher model continuously outputs the policy distribution (e.g., token generation probability), and the student model moves closer to the teacher policy through the KL divergence constraint to avoid reward overfitting.
[0126] Step 2052: construct a second loss function based on the KL divergence loss and the feature alignment loss;
[0127] For example, the KL divergence loss (i.e., strategy alignment loss) is combined with the feature alignment loss. Through strategy-feature joint optimization, the student strategy is forced to take into account both the underlying feature expression and the high-level decision logic when updating.
[0128] Step 20532: Based on the second loss function and the rewards of the preset rules, the supervised fine-tuned student model is subjected to reinforcement learning training.
[0129] For example, rule rewards are based on manually designed rules. For example, for a mathematical reasoning task, the rule might be: Rrule = I correctness + 0.5 * step score - 0.2 * number of redundant steps. A correct answer gets 1 point; each key reasoning step (such as applying the law of cosines) adds 0.5 points; and each redundant step (such as repeated calculations) deducts 0.2 points.
[0130] The embodiment of the present invention adopts a hybrid constrained optimization mechanism. On the one hand, it forces the student strategy distribution to be closer to the teacher strategy through real-time KL divergence constraint (token generation probability) to avoid deviation from the correct reasoning path; on the other hand, it introduces feature alignment rewards and uses the semantic similarity of the teacher's intermediate layer features as a reward signal to ensure that the student model strategy optimization is consistent with the underlying feature expression of the teacher model.
[0131] In a possible implementation, the second loss function further includes group relative strategy optimization loss and Logit distribution alignment loss;
[0132] Among them, the second loss function is obtained based on the following formula:
[0133] L RL =λ1·L GRPO +λ2·L align +λ3·L KL +λ4·L logit
[0134] Among them, L RL Represents the second loss function; L GRPO represents the group relative strategy optimization loss; L align represents feature alignment loss; L KL represents KL divergence loss; L logit represents the Logit distribution alignment loss; λ1, λ2, λ3, and λ4 represent the weights of each loss term.
[0135] For example, when processing a mathematical reasoning task, the weights λ1 = 0.5, λ2 = 0.2, λ3 = 0.1, and λ4 = 0.2.
[0136] In some embodiments, the group relative policy optimization loss (GRPO loss, also known as policy optimization based on relative advantage within a group) is obtained based on the following formula:
[0137]
[0138] Where, the clipping threshold ∈=0.2. The advantage function A is obtained based on the following formula t :
[0139]
[0140] Among them, R t represents the original reward of the t-th sample, for example, it can be a reward calculated based on a rule; μ group represents the mean value of the current group reward; σ group Represents the standard deviation of the rewards for the current group. Based on the mean (μ_group) and standard deviation (σ_group) of the rewards for samples within the current group, the raw rewards are normalized.
[0141] In some embodiments, the feature alignment loss is derived based on the following formula:
[0142]
[0143] For example, the intermediate layer features of the teacher model can be pre-cached. For example, before reinforcement learning training, the intermediate features can be pre-calculated. It is stored and can be directly read during the reinforcement learning stage, reducing the real-time computing overhead during reinforcement learning.
[0144] In some embodiments, the Logit distribution alignment loss is obtained based on the following formula:
[0145]
[0146] Among them, dynamic temperature adjustment: linear annealing from τ=3.0 to τ=1.0.
[0147]
[0148] For example, the Logit of the teacher model can be pre-cached. For example, before reinforcement learning training, the data pre-processing stage pre-calculates T logit (x), which can be directly read during the reinforcement learning stage.
[0149] The embodiment of the present invention integrates rule rewards, feature alignment rewards and KL divergence constraints in the strategy optimization of the reinforcement learning stage, which can suppress the strategy drift of the student model.
[0150] Figure 6 This is a diagram of the system architecture of the reinforcement learning phase provided by an embodiment of the present invention. Figure 6 ,The system architecture of the reinforcement learning stage includes: candidate generation, data scoring, teacher model processing and model training.
[0151] 1. Candidate generation: The labeled data is fed into the student model to generate five candidate answers. The candidate answers are sampled (p = 0.9) to obtain (y1, y2, y3, y4, y5) and placed into the candidate pool.
[0152] 2. Data scoring: For the candidate data in the candidate pool, use the rule scorer to score, add the scores, and obtain the training data (x, yi, score).
[0153] 3. Teacher model processing: The teacher model processes the training data, obtains the intermediate layer features (H_g_align) through the adapter, and stores the intermediate features. The logits of the teacher model output layer are stored.
[0154] 4. Model Training: The feature alignment loss is calculated based on the intermediate features of the student model and the stored intermediate features of the teacher model. The KL loss is calculated based on the stored teacher model logits and the student model logits. The GRPO loss is calculated based on the policy distribution of the new policy of the student model and the policy distribution of the old policy. The parameters are updated based on the constructed total loss, and the old policy is updated using the policy replacer. The next iteration is then performed based on the labeled data.
[0155] Figure 7 This is a schematic diagram of the reinforcement learning training process provided by an embodiment of the present invention; Figure 7 ,The reinforcement learning training cycle includes: supervision signal synthesis, candidate generation and scoring, and teacher model evaluation.
[0156] 1. The labeled data is input into the student model, which generates candidate answers based on the current model. After kernel sampling, candidate data is obtained. The candidate data is evaluated according to the rules and the rule scores are obtained.
[0157] 2. Teacher model evaluation: The candidate data is input into the teacher model to obtain intermediate features and output layer Logits.
[0158] 3. Supervision signal synthesis: Based on the rule score, the intermediate features of the teacher model, and the teacher model output layer Logits, a comprehensive supervision signal is obtained.
[0159] 4. The student model updates its strategy based on the comprehensive supervision signal. If converged, the optimized strategy model is obtained. If not, the labeled data is input into the student model for the next iteration.
[0160] Step 206: Process the natural language text data to be analyzed based on the student model after knowledge distillation.
[0161] For example, the natural language text data to be analyzed may be a mathematical reasoning task or a code generation task.
[0162] The embodiment of the present invention aggregates and reduces the intermediate layer features of the teacher model during the supervised fine-tuning stage, aligns them with the intermediate layer features of the student model, and dynamically maps the intermediate layer features of the teacher model to the intermediate layers of the student model. After the intermediate layer features are aligned, a loss function is constructed based on the difference in intermediate layer features between the teacher model and the student model to train the student model. As a result, the trained student model can not only imitate the output results of the teacher model, but also learn deep features such as the teacher model's reasoning logic and decision-making path. The student model can learn feature information at different levels of the teacher model, so as to better understand and imitate the teacher model's reasoning process, thereby improving the accuracy of the student model's natural language processing.
[0163] The embodiment of the present invention breaks through the bottleneck of black box knowledge transfer. Traditional distillation methods only rely on the final output of the teacher model (such as the answer text or the output layer probability), and cannot transfer the complex reasoning logic implied in the intermediate layer (such as the selection of the intermediate lemma of mathematical proof, the API call logic of code generation). The embodiment of the present invention uses dynamic hierarchical feature alignment and low-rank adapter conversion to cross-dimensionally map the deep semantic features of the teacher model (such as the 12th to 24th layers) and the shallow representation of the student model (such as the 4th to 8th layers), so that the student model not only imitates the teacher's "answer results", but also learns its "thinking process". For example, in the geometry proof task, the student model automatically captures the teacher's hidden similar triangle theorem application pattern through feature alignment, rather than mechanically copying the answer text.
[0164] Embodiments of the present invention also eliminate the risk of policy drift. Experiments show that this design reduces the logical error rate of generated content by 63% and reward overfitting by 59%. At the same model scale, this invention improves the student model's accuracy on complex tasks by 30%-50%, while reducing video memory usage by 36.8%, providing a high-performance, low-cost, lightweight language model deployment solution for resource-constrained scenarios.
[0165] In the field of natural language processing and model compression, the present invention proposes a multi-stage teacher-student collaborative distillation method and system for large language models (LLMs). Based on the two-stage collaborative distillation method of supervised fine-tuning (SFT) and reinforcement learning (RL), the present invention realizes efficient knowledge transfer from the teacher model to the student model through intermediate layer feature alignment and hybrid constraint optimization. The core of the solution includes the following modules: supervised fine-tuning stage: through sliding window grouping and low-rank adapter, the teacher's deep features are dynamically mapped to the student space, and the student model is optimized in combination with Logit alignment and answer matching loss; reinforcement learning stage: rule rewards, feature alignment rewards and KL divergence constraints are integrated in policy optimization to suppress policy drift. This solution significantly improves the accuracy and stability of the student model in complex tasks (such as mathematical reasoning and code generation) while reducing resource consumption. The following uses mathematical reasoning tasks as an example to illustrate the technical concept of the present invention.
[0166] 1. Model configuration:
[0167]
[0168] 2. Hardware configuration:
[0169]
[0170] 3. SFT stage dataset construction:
[0171]
[0172]
[0173] Data augmentation: For each problem, the teacher model generates multiple different solution paths, covering the following types:
[0174] Step-by-step decomposition: (e.g., “12×1 / 4=3→12-3=9→…”);
[0175] Comprehensive formula type: (e.g., "12×(3 / 4)×4=36");
[0176] Reverse deduction type: (such as "assuming the final number is x, reverse the deduction x = 36").
[0177] Example of path expansion: Problem: There are 30 red and blue balls in a box. There are 6 more red balls than blue balls. Find the number of red balls. Path 1: Let x = 30 blue balls → x + (x + 6) = 30 → x = 12 → 18 red balls; Path 2: Total + difference method → (30 + 6) / 2 = 18; Path 3: Verify by enumeration (18 red balls, 12 blue balls).
[0178] Below is an example of augmented training data. Each path corresponds to a piece of training data of the form (question, answer, path). In the above example, after augmenting one piece of data, two SFT training data are obtained.
[0179]
[0180]
[0181] 4. RL stage dataset construction
[0182] 4.1. Core Dataset: Main source: All training problems in the SFT phase (GSM8K+MATH+SVAMP), accounting for 80%. New problems: 20% of the problems not included in SFT were extracted from APPS and Ape210K to test generalization.
[0183] 4.2. Candidate Answer and Path Generation: The current student model uses kernel sampling: Top-p = 0.9, filters low-probability tokens, retains reasonable diversity, and generates 5 answers with reasoning paths for each question.
[0184] 4.3. Bonus Score Calculation: Answer Correctness (1 point). Strict Match: Numerical results must be identical to the standard answer (e.g., "18" vs. "18.0" is considered correct). Scientific Notation: Format conversions are permitted (e.g., "1.8e1" must be converted to 18 before verification).
[0185] Core step completeness (maximum 2 points).
[0186]
[0187] Identification of equivalent steps: Mathematical equivalence: For example, "12×(3 / 4)=9" and "12-3=9" are considered equivalent. Logical equivalence: For example, "calculating the number of red balls using the difference method" and "solving the equation" are considered equivalent.
[0188] Redundancy penalty (-0.2 points per step): Redundancy is determined by the following criteria: repeated calculations (e.g., calculating the same value twice); irrelevant derivations (e.g., adding a unit conversion step); and descriptive statements (e.g., "Next we need to do multiplication").
[0189]
[0190]
[0191] 5. Training process design
[0192]
[0193] SFT is divided into three stages of training
[0194]
[0195] RL (Reinforcement Learning) strategy optimization: Temperature annealing, initially τ = 3.0, linearly reduced to τ = 1.0, balancing exploration and exploitation. Early stopping: if the KL divergence for three consecutive batches is greater than 2.0, updates are paused and sampling is restarted.
[0196] 6. Effect example:
[0197] Using traditional distillation technology, the input is: Class A has 40 students, and Class B is 1 / 4 more than Class A. Find the total number of students in the two classes. The incorrect output is: 40 + (40 × 1 / 4) = 50 (the calculation step for Class B is missing).
[0198] The student model of the present invention can fix the above errors:
[0199] Feature alignment test: The cosine similarity between the 14th layer (the third from the bottom) of the student model and the 19-21st layers of the teacher model is 0.38 (threshold 0.6). A deviation in the feature vector for "Class B student count" was detected.
[0200] Strategy constraint intervention: KL divergence constraint forced completion step: 40×5 / 4=50→total number of people=40+50=90.
[0201] Reward feedback: Step score increased from 1.0 to 2.0, total score increased from 2.0 to 3.0.
[0202] The present invention realizes efficient knowledge transfer from the teacher model to the student model through a dynamic interaction mechanism of supervised fine-tuning (SFT) and reinforcement learning (RL) in a collaborative framework of feature layer alignment, policy feedback constraints and cross-stage parameter sharing.
[0203] This method is particularly suitable for the rapid deployment of high-precision lightweight language models in resource-constrained scenarios such as mobile terminals and edge computing nodes. It can improve the reasoning ability of student models under the same training data, or significantly compress the model size to reduce storage and computing costs. Experimental verification shows that this method has significant advantages over traditional solutions in mathematical reasoning (GSM8K) and code generation (HumanEval) tasks. For example, the depth of knowledge transfer is improved: the accuracy of student models in tasks requiring multi-step reasoning is improved by 23%-29%; policy stability is enhanced: the error generation rate caused by policy drift is reduced by 45%-60%; cross-stage collaborative benefits: end-to-end training efficiency (performance gain per unit time) is improved by 2.1 times.
[0204] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0205] The following are device embodiments of the present invention. For details not fully described therein, reference may be made to the corresponding method embodiments described above.
[0206] Figure 8 This is a schematic diagram of the structure of a natural language processing device for large language model knowledge distillation provided by an embodiment of the present invention. For ease of explanation, only the parts related to the embodiment of the present invention are shown, which are detailed as follows: Figure 8 As shown, a natural language processing device 8 for large language model knowledge distillation includes:
[0207] An acquisition module 81 is used to acquire an initial student model, wherein the number of layers and dimensions of the student model are smaller than those of the teacher model; the teacher model is a large language model;
[0208] An alignment module 82 is used to aggregate and reduce the dimensionality of each intermediate layer feature of the teacher model, align the number of layers and dimensions with the intermediate layer of the student model, and obtain the aligned intermediate layer features of the teacher model;
[0209] A loss function construction module 83 is configured to obtain a feature alignment loss based on the difference between the intermediate layer features of the teacher model after alignment and the intermediate layer features of the corresponding student model, thereby constructing a first loss function;
[0210] A first training module 84 is configured to perform supervised fine-tuning iterative training on the initial student model based on the constructed first loss function to obtain a supervised fine-tuned student model;
[0211] The second training module 85 is used to perform reinforcement learning training on the student model after supervised fine-tuning to obtain a student model after knowledge distillation.
[0212] The processing module 86 is used to process the natural language text data to be analyzed based on the student model after knowledge distillation.
[0213] Figure 9 Schematic diagram of an electronic device provided by an embodiment of the present invention. Figure 9 As shown, the electronic device 9 of this embodiment includes a processor 90 and a memory 91. The memory 91 stores a computer program 92. When the processor 90 executes the computer program 92, the steps of the above-described method embodiments are implemented. Alternatively, when the processor 90 executes the computer program 92, the functions of the modules / units in the above-described device embodiments are implemented.
[0214] For example, the computer program 92 can be divided into one or more modules / units, which are stored in the memory 91 and executed by the processor 90 to complete the present application. The one or more modules / units can be a series of computer program instruction segments capable of completing a specific function, which are used to describe the execution process of the computer program 92 in the electronic device 9.
[0215] The electronic device 9 can include, but is not limited to, the processor 90 and the memory 91. Those skilled in the art can understand that the electronic device 9 can further include other components, such as an input / output device, a network access device, a bus, etc. Figure 9 The electronic device 9 is only an example and does not constitute a limitation on the electronic device 9, and can include more or fewer components than the illustration, or combine certain components, or different components, for example, the electronic device 9 can also include an input / output device, a network access device, a bus, etc.
[0216] The processor 90 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0217] The memory 91 can be an internal storage unit of the electronic device 9, such as a hard disk or a memory of the electronic device 9. The memory 91 can also be an external storage device of the electronic device 9, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 9. Further, the memory 91 can include both the internal storage unit and the external storage device of the electronic device 9. The memory 91 is used to store the computer program 92 and other programs and data required by the electronic device 9. The memory 91 can also be used to temporarily store data that has been output or will be output.
[0218] For the convenience and brevity of description, only the above-mentioned division of the functional modules / units is exemplified, and in actual application, the above-mentioned functions can be completed by different functional modules / units according to needs. The above-mentioned modules / units can be realized in the form of hardware, software or a combination of hardware and software.
[0219] An embodiment of the present invention further provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods in the above-mentioned method embodiments.
[0220] An embodiment of the present invention further provides a computer program product, including a computer program, which, when executed by a processor, implements the methods in the above-mentioned method embodiments.
[0221] The computer program includes computer program code, which may be in source code form, object code form, executable file, or some intermediate form. Computer-readable media may include any entity or device capable of carrying computer program code, recording media, USB flash drives, mobile hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunications signals, and software distribution media.
[0222] In the above embodiments, the descriptions of each embodiment have their own focus. For parts not described or recorded in detail in one embodiment, please refer to the relevant descriptions of other embodiments. Unless otherwise specified or there is a logical conflict, the terms and / or descriptions between different embodiments are consistent and can be referenced to each other. The technical features of different embodiments can be combined to form new embodiments based on their inherent logical relationships.
[0223] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A natural language processing method for knowledge distillation of a large language model, characterized by: include: Obtaining an initial student model, wherein the number of layers and dimensions of the student model are smaller than those of the teacher model; the teacher model is a large language model; Aggregate and reduce the dimension of each intermediate layer feature of the teacher model, align the number of layers and dimensions with the intermediate layer of the student model, and obtain the aligned intermediate layer features of the teacher model; Based on the difference between the intermediate layer features of the teacher model after alignment and the intermediate layer features of the corresponding student model, a feature alignment loss is obtained to construct a first loss function; Performing supervised fine-tuning iterative training on the initial student model based on the constructed first loss function to obtain a supervised fine-tuned student model; Perform reinforcement learning training on the supervised fine-tuned student model to obtain the student model after knowledge distillation; Based on the student model after knowledge distillation, the natural language text data to be analyzed is processed.
2. The natural language processing method for large language model knowledge distillation according to claim 1, characterized in that: The intermediate layer features of the teacher model are aggregated and reduced in dimension to align with the number of layers and dimensions of the intermediate layer of the student model. The aligned intermediate layer features of the teacher model include: Divide the intermediate layers of the teacher model into L S Group, where L S is the number of intermediate layers of the student model; For any set of intermediate layer features of the teacher model, gated attention aggregation is used to obtain the aggregated intermediate layer features; For any aggregated intermediate layer feature, a low-rank adapter is used to map the high-dimensional aggregated intermediate layer to the student model dimension, and the intermediate layer feature that is aligned with the number of layers and dimensions of the student model intermediate layer is obtained as the aligned intermediate layer feature of the teacher model.
3. The natural language processing method for large language model knowledge distillation according to claim 2, characterized in that: The intermediate layers of the teacher model are divided into L S Groups include: Using the sliding window grouping method, the middle layers of the teacher model are divided into L S groups; each group contains a continuous number of layers Among them, L T represents the number of intermediate layers of the teacher model, w represents the number of layers in each group; sliding step 4. The natural language processing method for large language model knowledge distillation according to claim 2, characterized in that: For any set of intermediate layer features of the teacher model, gated attention aggregation is used to obtain the aggregated intermediate layer features including: The following formula is used to obtain the characteristics of the intermediate layer after polymerization: in, represents the characteristics of the middle layer after the aggregation of the layers in group g; α g,i represents the attention weight of the i-th layer of the g-th group; represents the intermediate layer features of the i-th layer of the g-th group; Represents the trainable parameters of the student model; represents average pooling.
5. The natural language processing method for large language model knowledge distillation according to claim 2, characterized in that: For any aggregated intermediate layer feature, a low-rank adapter is used to map the high-dimensional aggregated intermediate layer to the student model dimension, and obtain an intermediate layer feature that is aligned with the number of layers and dimensions of the student model intermediate layer. The aligned intermediate layer features of the teacher model include: The aligned intermediate layer features of the teacher model are obtained based on the following formula: r=L S in, Represents the middle layer features after alignment; A g and B g represents a low-rank matrix; r = L S Indicates that r is equal to the number of layers in the student model.
6. The natural language processing method for large language model knowledge distillation according to claim 2, characterized in that: The feature alignment loss obtained based on the difference between the intermediate layer features of the teacher model after alignment and the intermediate layer features of the corresponding student model includes: The feature alignment loss is obtained based on the following formula: in, represents feature alignment loss; S g represents the features of the g-th layer of the student model; γ g represents the weight of the g-th layer; d S Represents the dimension of the student model.
7. The natural language processing method for large language model knowledge distillation according to claim 2, characterized in that: The reinforcement learning training of the supervised fine-tuned student model includes: Based on the KL divergence of the probability distribution of the teacher model output and the probability distribution of the student model output, the KL divergence loss is obtained; Constructing a second loss function based on the KL divergence loss and the feature alignment loss; Based on the second loss function and the rewards according to the preset rules, the supervised fine-tuned student model is subjected to reinforcement learning training.
8. The natural language processing method for large language model knowledge distillation according to claim 7, characterized in that: The second loss function also includes group relative strategy optimization loss and Logit distribution alignment loss; Among them, the second loss function is obtained based on the following formula: L RL =λ1·L GRPO +λ2·L align +λ3·L KL +λ4·L logit Among them, L RL Represents the second loss function; L GRPO represents the group relative strategy optimization loss; L align represents feature alignment loss; L KL represents KL divergence loss; L logit represents the Logit distribution alignment loss; λ1, λ2, λ3, and λ4 represent the weights of each loss term.
9. A natural language processing device for knowledge distillation of a large language model, characterized in that: include: An acquisition module is used to obtain an initial student model, wherein the number of layers and dimensions of the student model are smaller than those of the teacher model; the teacher model is a large language model; The alignment module is used to aggregate and reduce the dimensionality of the intermediate layer features of the teacher model, align the number of layers and dimensions with the intermediate layers of the student model, and obtain the aligned intermediate layer features of the teacher model; A loss function construction module is used to obtain a feature alignment loss based on the difference between the intermediate layer features of the teacher model after alignment and the intermediate layer features of the corresponding student model, thereby constructing a first loss function; A first training module is configured to perform supervised fine-tuning iterative training on the initial student model based on the constructed first loss function to obtain a supervised fine-tuned student model; The second training module is used to perform reinforcement learning training on the student model after supervised fine-tuning to obtain the student model after knowledge distillation; The processing module is used to process the natural language text data to be analyzed based on the student model after knowledge distillation.
10. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the method for natural language processing of large language model knowledge distillation as described in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Reverse channel pruning compression method and device based on multilevel knowledge distillation
CN117744737A
Heterogeneous teacher-student model knowledge distillation method based on middle layer alignment
CN118734934A
Knowledge distillation optimization method based on sparse mixed expert and low-rank adaptation
CN118982072A
Fusing output of artificial intelligence networks
US20200210810A1