Supply chain and e-commerce purchasing field large model compression and online incremental learning method based on three-order combined distillation
The three-stage distillation method for model compression and online learning addresses the inefficiencies of existing LLMs in supply chain and e-commerce by reducing latency and enhancing adaptability, ensuring high accuracy and real-time decision-making.
Patent Information
- Application Number
- CN202510383981.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-03-28
AI Technical Summary
In the supply chain and e-commerce procurement scenarios, existing large language models have problems such as insufficient model lightweighting and poor adaptability of dynamic data, resulting in high inference delay, large hardware resource utilization, and serious accuracy losses, making it difficult to meet the real-time requirements and adaptability of data distribution drift.
The third-order joint distillation method is used for model compression, combined with the hierarchical migration protocol and multimodal optimization, and the model is dynamically updated to adapt to data distribution drift through staged distillation of the encoder, decoder and prediction head, combined with the online incremental learning mechanism.
Significantly reduce inference delay, improve model accuracy, adapt to scenarios such as e-commerce promotions, seasonal demands and supplier quotation fluctuations, realize multi-task optimization, and maintain efficient inference and accuracy.
Smart Images

Figure CN120317899A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and specifically relates to a method for compressing and online incremental learning of large models in the fields of supply chain and e-commerce procurement based on third-order joint distillation. Background Art
[0002] Tasks such as compliance analysis and path reasoning in supply chain management need to process massive amounts of unstructured text data (such as logistics documents, contract terms) and e-commerce procurement transaction information (such as real-time inventory, order execution, and supplier quotes). Existing technologies generally rely on large language models (LLMs) for semantic understanding and decision-making generation, but such models still have the following technical limitations in supply chain and e-commerce procurement scenarios:
[0003] (1) Insufficient model lightweight
[0004] The number of parameters of general large language models usually exceeds tens of billions (for example, 175 billion parameters of GPT-3), resulting in high inference latency and large hardware resource occupation. According to literature reports, the single inference time of traditional supply chain compliance analysis models generally exceeds several seconds, which is difficult to meet the real-time requirements in industrial production. Although knowledge distillation (KD) technology can compress the model through a teacher-student architecture, in the supply chain multi-task scenario, the distillation process is usually accompanied by significant accuracy loss. For example, the accuracy of the compressed model in the entity resolution task generally drops by more than 10%. In the field of e-commerce procurement, when quickly splitting and allocating order demands during e-commerce promotions, the latency and accuracy loss of the model will further exacerbate the actual application obstacles.
[0005] (2) Poor adaptability to dynamic data
[0006] The dynamic characteristics such as logistics path changes and cost fluctuations in the supply chain scenario lead to frequent drift of data distribution. Most existing large language models adopt a static training mode and need to be retrained in full to adapt to new data, resulting in high time and resource overhead. Incremental learning technology theoretically supports online updates, but its ability to adapt to data drift is limited, and it is difficult to achieve a balance between real-time performance and model stability. E-commerce procurement also has situations such as seasonal promotions, frequent changes in market conditions and supplier quotes, making it difficult for traditional models to make accurate and rapid order placement decisions and demand predictions.
[0007] Current technological improvements mainly focus on a single optimization direction. For example, global pruning methods can reduce the number of model parameters, but in the task of supply chain entity relationship extraction, the performance loss of the sparsified model generally exceeds 15%. Similar problems also exist in the automatic merging and splitting algorithms for e-commerce purchase orders. After deep pruning, the order anomaly detection and dynamic scheduling capabilities are likely to decline. In addition, some studies have attempted to optimize the dynamic update mechanism of the model, but the collaborative optimization of lightweight and dynamic adaptability still faces challenges, restricting the large-scale application of large language models in the fields of supply chain and e-commerce procurement. Summary of the Invention
[0008] In view of the deficiencies in the prior art, the present invention provides a method for large model compression and online incremental learning in the fields of supply chain and e-commerce procurement based on third-order joint distillation. Knowledge distillation is achieved through a hierarchical migration protocol and multimodal joint optimization, and an online incremental learning mechanism is combined to address data distribution drift problems such as e-commerce promotions, seasonal demands, and supplier quote fluctuations.
[0009] To achieve the above objectives, the present invention adopts the following technical solutions:
[0010] A method for large model compression and online incremental learning in the fields of supply chain and e-commerce procurement based on third-order joint distillation, comprising the following steps:
[0011] Collect supply chain compliance documents and structured data of purchase orders, and preprocess the supply chain compliance documents and structured data of purchase orders to obtain a dataset;
[0012] Respectively construct a teacher model and a student model based on a large language model. Both the teacher model and the student model include an encoder, a decoder, and a prediction head;
[0013] Perform three-stage knowledge distillation on the teacher model and the student model using the dataset; the three-stage knowledge distillation specifically performs hierarchical knowledge migration in the order of the encoder distillation stage, the decoder distillation stage, and the prediction head distillation stage;
[0014] Collect new data, and determine whether the new data triggers the condition for incremental learning. If so, perform incremental learning using the teacher model and the student model obtained through three-stage knowledge distillation to obtain a trained teacher model and a trained student model;
[0015] Perform collaborative inference using the trained teacher model and the trained student model. If the prediction confidence of the student model is less than the threshold, weight and fuse the outputs of the teacher model and the student model to obtain the results of compliance clause classification, procurement cost distribution prediction, and supplier priority ranking; otherwise, directly use the prediction result of the student model as the final output to obtain the results of compliance clause classification, procurement cost distribution prediction, and supplier priority ranking.
[0016] To optimize the above technical solutions, the specific measures also include:
[0017] Further, the preprocessing of the structured data of the supply chain compliance documents and purchase orders specifically includes:
[0018] Labeling clause classification tags, paragraph segmentation and regularization processing, BPE-based word segmentation, unit unification, field alignment and missing value imputation.
[0019] Further, the encoder includes a self-attention layer and a feed-forward sublayer;
[0020] The specific process of the encoder distillation stage is as follows:
[0021] Divide the dataset into a training set and a validation set, and configure hyperparameters, which include learning rate, batch size, maximum number of iterations, and optimizer;
[0022] After the encoder is trained on the training set, use the validation set for validation and calculate the F1 score. If the fluctuation range of the F1 score is less than the set range for 5 consecutive times, then part of the encoder distillation stage converges, and freeze the encoder parameters, including freezing the parameters of all self-attention layers and feed-forward sublayers. If the loss value of the validation set does not decrease after the set number of iterations, then reduce the learning rate and enable gradient clipping;
[0023] During the process of training the encoder, cache the intermediate representations of the teacher model locally.
[0024] Further, position encoding or segment vectors are added to the input end of the decoder to distinguish different data modalities, and the data modalities include text, numerical values, and timestamps; or an embedding layer and a multi-layer perceptron connected in sequence are added to the input end of the decoder for projecting numerical fields;
[0025] The specific process of the decoder distillation stage specifically includes:
[0026] When processing multi-modal data at the decoder end, map different modalities to an alignable vector space for multi-modal feature fusion;
[0027] Configure hyperparameters, including learning rate and batch size;
[0028] The goals of the decoder distillation stage include attention matrix alignment and semantic distribution alignment; the specific attention matrix alignment is as follows: the attention matrices generated by each layer of the student model are similar to the attention matrices of the corresponding layers of the teacher model, and the similarity between the two is measured by the second norm or mean square error; the specific semantic distribution alignment is as follows: the output of the student model is aligned with the output distribution of the teacher model on the same input, and the difference between the two is measured by KL divergence or mean square error;
[0029] When the path inference accuracy or F1 score reaches the threshold, freeze the decoder parameters; otherwise, increase the number of iterations for fine-tuning the decoder.
[0030] By projecting the student model's attention to the teacher model's space, where ψ(·) represents the projection function from the student model's attention representation space to the teacher model's attention representation space, and W ψ is a trainable projection matrix used to map the dimension or distribution of the student attention matrix to a space equivalent or comparable to that of the teacher attention. is the attention matrix generated by the student model at the l ′ -th layer.
[0031] Update the projection matrix W ψ , and add an L2 regularization term when updating the projection matrix, and update the parameters of the projection matrix during each backpropagation.
[0032] Furthermore, the prediction head distillation stage is specifically as follows:
[0033] Update the parameters of the low-rank adaptation layer W adapt , train the prediction head, randomly sample historical data from the encoder distillation stage and the decoder distillation stage and mix them into the current training to prevent the model from forgetting previous tasks in the final distillation stage; during the training process, evaluate the F1 score of the compliance clause classification task and the Wasserstein distance of the procurement cost distribution prediction task every time a preset number of iterations is reached. If the evaluation metrics do not improve for 3 consecutive times, terminate the prediction head distillation.
[0034] Furthermore, the conditions for incremental learning are specifically the data distribution difference condition or the business metric monitoring condition:
[0035] The data distribution difference condition is: if the class labels of more than a preset percentage of samples in the new data exceed the distribution range of the old data, or the KL divergence difference is greater than the threshold, then trigger incremental learning;
[0036] The business metric monitoring condition is: the domain entity recall rate is less than the set value for 3 consecutive batches, or the MSE of the trading volume prediction exceeds 1.5 times the set threshold, or the supplier ranking NDCG drops by more than a preset percentage.
[0037] Furthermore, the specific process of performing incremental learning using the teacher model and the student model obtained by three-stage knowledge distillation is as follows:
[0038] Adopt the strategy of the prediction head distillation stage, only update the low-rank adaptation layer, and keep the encoder and decoder frozen; if partial feature extraction layers need to be fine-tuned, unfreeze the last 1-2 layers of the encoder on the basis of updating the low-rank adaptation layer, and set the corresponding learning rate to 10% of the initial learning rate.
[0039] New data is collected weekly for incremental learning, and the Mini-Epoch strategy is also adopted during the incremental learning phase.
[0040] Furthermore, the loss function of the three-stage knowledge distillation is as follows:
[0041]
[0042] In the formula, represents the loss value of the three-stage knowledge distillation, L SD represents the semantic distribution distillation loss, and α is a weighting coefficient used to control the proportion of the topological perception attention distillation loss L TA in the total loss, and β is a weighting coefficient used to control the proportion of the domain reinforcement task loss L DT in the total loss;
[0043] The expression of the semantic distribution distillation loss is as follows:
[0044]
[0045] In the formula, L SD is the semantic distribution distillation loss, τ is the temperature coefficient, KL(·) represents the Kullback-Leibler divergence, σ(·) represents the Softmax function, z T and z S represent the log-odds vectors of the teacher model and the student model respectively;
[0046] The expression of the topological perception attention distillation loss is as follows:
[0047]
[0048] In the formula, l represents the layer index in the teacher model, L a is the set of layers selected for distillation in the teacher model, is the mapping or processing operation applied to the attention matrix of the l-th layer of the teacher model, is the mapping applied to the attention matrix ′ of the corresponding l-th layer of the student model;
[0049] The expression of the domain reinforcement task loss is as follows:
[0050]
[0051] In the formula, λ i represents the loss weight of the i-th sub-task, L task,iRepresents the loss function of the $i$-th sub-task, where the sub-tasks include compliance clause classification task, procurement cost distribution prediction task, and supplier priority ranking task;
[0052] After each training round, the loss weights of each sub-task are automatically updated through the gradient information on the validation set, and the formula is as follows:
[0053]
[0054] In the formula, Represents the validation set loss $L$ val The partial derivative of $\lambda$ i is used to measure the error sensitivity of the sub-task in the current model.
[0055] Furthermore, the loss function of the incremental learning is as follows:
[0056]
[0057] In the formula, $L$ 2 total Represents the loss value of incremental learning, $L$ SD Represents the semantic distribution distillation loss, $\alpha$ is a weighting coefficient used to control the proportion of the topological perception attention distillation loss $L$ TA in the total loss, and $\beta$ is a weighting coefficient used to control the proportion of the domain reinforcement task loss $L$ DT in the total loss; $\mu$ represents the smoothness coefficient, and $\theta$ s represents the parameter vector of the student model at the current training stage, represents the parameter vector saved by the student model at the previous moment or the previous incremental stage;
[0058] The expression of the semantic distribution distillation loss is as follows:
[0059]
[0060] In the formula, $L$ SD is the semantic distribution distillation loss, $\tau$ is the temperature coefficient, $KL(\cdot)$ represents the Kullback-Leibler divergence, $\sigma(\cdot)$ represents the Softmax function, and $z$ T and $z$ S represent the logit vectors of the teacher model and the student model respectively;
[0061] The expression of the topological perception attention distillation loss is as follows:
[0062]
[0063] In the formula, $l$ represents the layer index in the teacher model, and $L$ a is the set of layers selected for distillation in the teacher model, is a mapping or processing operation applied to the attention matrix of the l-th layer of the teacher model and is a mapping applied to the attention matrix of the corresponding l-th layer of the student model ′ ; The expression of the domain reinforcement task loss is as follows:
[0064] In the formula, λ
[0065]
[0066] represents the loss weight of the i-th sub-task, and L i represents the loss function of the i-th sub-task. The sub-tasks include the compliance clause classification task, the procurement cost distribution prediction task, and the supplier priority ranking task; task,i After each training round, the loss weight of each sub-task is automatically updated through the gradient information on the validation set. The formula is as follows:
[0067] In the formula,
[0068]
[0069] represents the partial derivative of the validation set loss L with respect to λ val and is used to measure the error sensitivity of the sub-task in the current model. i ;
[0070] Furthermore, the loss of the compliance clause classification task is measured by the cross-entropy loss. The loss function of the compliance clause classification task is as follows:
[0071]
[0072] In the formula, L task,1 represents the loss of the compliance clause classification task, and L CE represents the cross-entropy loss, and k represents the compliance category index
[0073] The loss of the procurement cost distribution prediction task is measured by the Wasserstein distance. The loss function of the procurement cost distribution prediction task is as follows:
[0074]
[0075] In the formula, L task,2 represents the loss of the procurement cost distribution prediction task, and L EMD represents the cumulative distance between the predicted distribution and the true distribution. i is the cost interval index, and CDF pred (i) represents the value of the predicted cumulative distribution function at the interval index i, and CDF true(i) represents the value of the true cumulative distribution function at the interval index i;
[0076] The loss of the supplier priority ranking task is measured by the ranking loss, and the loss function of the supplier priority ranking task is as follows:
[0077]
[0078] In the formula, L task,3 is the loss of the supplier priority ranking task, L Rank represents the ranking loss, s i and s j respectively represent the ranking scores of supplier i and supplier j predicted by the model, y (i,j) represents the preference label of supplier i and j. If i performs better than j, then y (i,j) = +1; if j is better than i, then y (i,j) = -1, and pair(i,j) refers to all comparable supplier pairs in the training set.
[0079] The beneficial effects of the present invention are:
[0080] By gradually distilling and freezing the core components of the model in three stages, the inference latency can be significantly reduced, and refined distillation and adaptation are achieved at the levels of the encoder, decoder, and prediction head.
[0081] By combining the semantic distribution distillation loss L SD with the topology-aware attention distillation loss L TA , knowledge of the teacher model can be respectively learned in terms of output layer distribution alignment and intermediate attention structure alignment, so as to obtain a more comprehensive distillation effect in the multi-modal fusion scenario; at the same time, to balance the different requirements of each sub-task, the domain reinforcement task loss L DT dynamically balances the training weights of different tasks, enabling the student model to maintain sensitivity to business goals while inheriting the knowledge of the teacher; finally, by introducing in the online stage to constrain the drastic update of the student model parameters, it can not only adapt to the drift of the new data distribution but also slow down the forgetting of the original knowledge. The four together constitute a comprehensive optimization objective. Through the joint optimization of semantic distribution, topology-aware attention, and domain reinforcement tasks, the model can balance the knowledge alignment and target task requirements under multi-modal input, and achieve higher inference accuracy and lower model complexity in supply chain and e-commerce procurement scenarios.
[0082] Through the progressive distillation and incremental learning mechanism, while maintaining efficient inference, it can dynamically adapt to scenarios such as e-commerce promotions, seasonal demands, and supplier quote fluctuations, and achieve continuous optimization of multiple tasks (compliance clause classification, entity resolution, path prediction, purchase order optimization, etc.). Description of the Drawings
[0083] Figure 1 This is the overall flowchart of the large model compression and online incremental learning method in the supply chain and e-commerce procurement fields based on third-order joint distillation proposed by the present invention.
[0084] Figure 2 This is the flowchart of the decoder distillation stage. Specific implementation manners
[0085] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0086] Embodiment 1
[0087] The present invention proposes a large model compression and online incremental learning method in the supply chain and e-commerce procurement fields based on third-order joint distillation. The process of this method is as Figure 1 shown and includes the following steps:
[0088] Collect supply chain compliance documents and procurement order structured data, and preprocess the supply chain compliance documents and procurement order structured data to obtain a dataset.
[0089] Supply chain compliance documents: In this embodiment, it is assumed that PDF files and text data after OCR conversion (such as import and export agreements) are used, and clause classification labels (such as "CE certification") are marked. The average length of the text is about 2000-3000 words. After paragraph segmentation and regularization processing (duplicate removal, noise removal, removal of special characters), tokenization based on BPE (Byte Pair Encoding) is performed.
[0090] Procurement order structured data: Order or contract information in JSON format (including fields such as supplier name, goods specifications, etc.). Example fields may include key attributes such as {"supplier_id", "product_type", "order_quantity", "unit_price", "delivery_date"}, and some contain nested levels (such as sub-fields under "product_details"). Standardize the data (unify units, align fields) and impute missing values.
[0091] Mixed-Modal Record: A comparison table of text paragraphs and quotation data in the tender document. By referring to the supplier quotations provided by an external archiving system, an index mapping of "text information - quotation fields" is established so that the model can read both natural language descriptions and numerical fields simultaneously during multi-modal input. It mainly involves quotation ranges, volume discounts, delivery requirements, etc.
[0092] A teacher model and a student model are respectively constructed based on large language models. Both the teacher model and the student model include an encoder, a decoder, and a prediction head.
[0093] In this embodiment, both the teacher model and the student model can be constructed based on large pre-trained language models (LLMs), such as pre-trained language models like BERT and GPT. However, there are differences in scales such as the number of layers, hidden dimensions, and the number of attention heads between the two to achieve a balance between performance and inference efficiency before and after distillation.
[0094] Teacher Model Structure: (1) Overall Framework: It can adopt a Transformer encoder-decoder structure or only use an encoder (for pure text encoding) combined with a downstream prediction head design. If the teacher model is used in a multi-modal input scenario, a corresponding numerical / timestamp Embedding layer is added at the input end of its encoder or decoder, or an MLP is connected after the original text Embedding to map structured fields. (2) Embedding Layer and Decoder: Embedding Layer: Maps text Tokens, numerical fields, etc. to fixed-dimensional vectors (e.g., d = 1024), which may include positional encoding (PositionalEncoding) or segment vectors (Segment Embedding). Decoder: Interacts with the encoder output and can only predict partial sequences in multi-task situations; its multi-head self-attention (Self-Attention) and cross-attention (Cross-Attention) matrix sizes are usually large, making it suitable to guide the student model to align the attention distribution during distillation. (3) Prediction Head: The teacher model usually uses a complete linear layer or a multi-layer perceptron (MLP) as the final output layer, and the dimension directly corresponds to the task requirements. For example, the Softmax layer for classification tasks, the linear layer or multi-layer projection layer for regression tasks. These head parameters will provide a reference distribution with high precision (such as classification probabilities, regression values, etc.) during distillation for the student model to align.
[0095] Student model structure: (1) Overall scale reduction: The student model can inherit the main structure of the teacher model, but reduce the number of Transformer layers (e.g., from 12 layers to 6 layers), reduce the hidden vector dimension d (e.g., from 1024 to 512), or reduce the number of heads H in the multi-head attention (e.g., from 16 to 8). Some parameters can be significantly compressed in scale by using low-rank factorization (e.g., W_adapt = U·V^T) or factorization methods. (2) Decoder and Embedding layer: To ensure the alignment of the input format, the student model still needs to retain the same type of Embedding (text Tokens, numerical fields, positional encodings, etc.), but with a smaller dimension or by merging some redundant vectors. The decoder structure will align with the hierarchical order of the teacher, but can be sparsified in terms of the number of layers or attention heads to reduce the computational cost. (3) Prediction head structure and distillation: The prediction head of the student model can be a scaled-down linear layer or a low-rank adaptation layer, using similar activation functions and output interfaces as the teacher to ensure output alignment.
[0096] Perform three-stage knowledge distillation on the teacher model and the student model using the dataset; the three-stage knowledge distillation is specifically to perform hierarchical knowledge transfer in the order of the encoder distillation stage, the decoder distillation stage, and the prediction head distillation stage; the encoder distillation stage is specifically: The encoder includes a self-attention layer and a feed-forward sub-layer; divide the dataset into a training set and a validation set, and configure hyperparameters, where the hyperparameters include the learning rate, batch size, maximum number of iterations, and optimizer; Exemplary hyperparameter configuration: learning rate 1×10 -4 , Batch Size = 32, iterate about 100,000 steps; The optimizer can be AdamW, β1 = 0.9, β2 = 0.999.
[0097] After the encoder is trained using the training set, use the validation set for validation, calculate the F1 score. If the fluctuation range of the F1 score is less than the set range for 5 consecutive times, then the encoder distillation stage is partially converged, and the encoder parameters are frozen, including freezing the parameters of all self-attention layers and feed-forward sub-layers. If the loss value of the validation set does not decrease after the set number of iterations, then reduce the learning rate and enable gradient clipping (e.g., clipping threshold = 1.0);
[0098] During the process of training the encoder, cache the intermediate representations of the teacher model locally to reduce repeated forward computations.
[0099] Add positional encoding or segment vectors to the input end of the decoder to distinguish different data modalities, where the data modalities include text, numerical values, and timestamps; or add a sequentially connected embedding layer and a multi-layer perceptron to the input end of the decoder for projecting numerical fields;
[0100] The flowchart of the decoder distillation stage is as Figure 2 shown, and the decoder distillation stage specifically includes:
[0101] When processing multi-modal data at the decoder end, different modalities are uniformly mapped to an alignable vector space for multi-modal feature fusion;
[0102] Configure hyperparameters, including the learning rate and batch size;
[0103] The objectives of the decoder distillation stage include attention matrix alignment and semantic distribution alignment; specifically, the attention matrix alignment is as follows: the attention matrices generated by each layer of the student model are similar to those of the corresponding layers of the teacher model, and the similarity between the two is measured by the second norm or mean squared error, so that the attention distribution pattern of the student converges to that of the teacher model. Specifically, the semantic distribution alignment is as follows: the output of the student model is aligned with the output distribution of the teacher model on the same input, and the difference between the two is measured by KL divergence or mean squared error (MSE); this can ensure that when the student makes the final decoding or prediction, it tries to maintain the same "semantic selection" or decision-making tendency as the teacher.
[0104] When the path reasoning accuracy (PRA) or F1 score reaches the threshold, freeze the decoder parameters; otherwise, increase the number of iterations for fine-tuning the decoder.
[0105] By projecting the attention of the student model into the teacher model space, where ψ(·) represents the projection function from the attention representation space of the student model to the attention representation space of the teacher model, and W ψ is a trainable projection matrix, generally with a dimension of (or depending on the reshape / flatten method corresponding to the number of attention heads and sequence length) for mapping the dimension or distribution of the student attention matrix to a space equivalent or comparable to that of the teacher attention, is the attention matrix generated by the l'-th layer of the student model; generally with a shape of
[0106] Update the projection matrix W ψ , and add an L2 regularization term when updating the projection matrix to prevent overfitting, and update the parameters of the projection matrix during each backpropagation.
[0107] Multi-modal fusion (introducing text and numerical embeddings at the decoder input) and attention mapping (comparing the attention distributions of the teacher and student) are carried out in the forward calculation and distillation loss calculation stages respectively.
[0108] Specifically, in the prediction head distillation stage: update the parameters of the low-rank adaptation layer W adapt = U·V T , whose dimension is usually less than 10% - 30% of the original fully connected layer, which can significantly reduce the model size while maintaining high accuracy.
[0109] Among them, U: The size is usually r << d. Responsible for mapping the input features to a more compact space after dimensionality reduction. V: The size is usually Used to re-project the reduced-dimensional representation back to an output space that is the same as or compatible with the original dimension. Train the prediction head, exemplary training: Conduct 20,000 steps, and either Batch Size = 16 or 32 is acceptable. Here, the training rounds can be shortened according to experience. For example, if the metrics have converged on a single task, the training can be stopped in advance. To stabilize the distillation effect, randomly sample the historical data (such as compliance terms, purchase orders) in the encoder distillation stage and the decoder distillation stage and mix them into the current training to prevent the model from forgetting the previous tasks in the final distillation stage; during the training process, every time a preset number of iterations is reached, evaluate the F1 score of the compliance term classification task and the Wasserstein distance of the purchase cost distribution prediction task. If the evaluation metrics do not improve for 3 consecutive times, terminate the prediction head distillation.
[0110] The loss function of the three-stage knowledge distillation is as follows:
[0111]
[0112] In the formula, represents the loss value of the three-stage knowledge distillation, L SD represents the semantic distribution distillation loss, α is a weighting coefficient used to control the proportion of the topological awareness attention distillation loss L TA in the total loss, and β is a weighting coefficient used to control the proportion of the domain reinforcement task loss L DT in the total loss;
[0113] The expression of the semantic distribution distillation loss is as follows:
[0114]
[0115] In the formula, L SD is the semantic distribution distillation loss, τ is the temperature coefficient, KL(·) represents the Kullback-Leibler divergence, σ(·) represents the Softmax function, zx T and z S represent the log-odds vectors of the teacher model and the student model respectively;
[0116] The expression of the topological awareness attention distillation loss is as follows:
[0117]
[0118] In the formula, l represents the layer index in the teacher model, L a is the set of layers selected for distillation in the teacher model, It is a mapping or processing operation applied to the attention matrix of the l-th layer of the teacher model and is a mapping applied to the attention matrix of the corresponding l-th layer of the student model ′ ; The expression of the domain reinforcement task loss is as follows:
[0119] In the formula, λ
[0120]
[0121] represents the loss weight of the i-th sub-task, and L i represents the loss function of the i-th sub-task. The sub-tasks include compliance clause classification tasks, procurement cost distribution prediction tasks, and supplier priority ranking tasks; task,i ;
[0122] After each training round, the loss weight of each sub-task is automatically updated through the gradient information on the validation set. The formula is as follows:
[0123]
[0124] In the formula, represents the partial derivative of the validation set loss L val with respect to λ i and is used to measure the error sensitivity of the sub-task in the current model.
[0125] The loss of the compliance clause classification task is measured by cross-entropy loss. The loss function of the compliance clause classification task is as follows:
[0126]
[0127] In the formula, L task,1 represents the loss of the compliance clause classification task, L CE represents the cross-entropy loss, k represents the compliance category index,
[0128] The loss of the procurement cost distribution prediction task is measured by the Wasserstein distance. The loss function of the procurement cost distribution prediction task is as follows:
[0129]
[0130] In the formula, L task,2 represents the loss of the procurement cost distribution prediction task, L EMD represents the cumulative distance between the predicted distribution and the true distribution, i is the cost interval index, and CDF pred (i) represents the value of the predicted cumulative distribution function at the interval index i, and CDF true (i) represents the value of the true cumulative distribution function at the interval index i;
[0131] The loss of the supplier priority ranking task is measured by the ranking loss, and the loss function of the supplier priority ranking task is as follows:
[0132]
[0133] In the formula, L task,3 is the loss of the supplier priority ranking task, L Rank represents the ranking loss, s i and s j respectively represent the ranking scores of supplier i and supplier j predicted by the model, y (i,j) represents the preference label of supplier i and j. If i performs better than j, then y (i,j) = +1; if j is better than i, then y (i,j) = -1, and pair(i,j) refers to all comparable supplier pairs in the training set.
[0134] Collect new data and determine whether the new data triggers the condition of incremental learning. If so, use the teacher model and student model obtained by three-stage knowledge distillation for incremental learning to obtain the trained teacher model and student model; the conditions for incremental learning are specifically the data distribution difference condition or the business metric monitoring condition:
[0135] The data distribution difference condition is: if the class labels of more than a preset percentage of samples in the new data exceed the distribution range of the old data, or the KL divergence difference is greater than the threshold, then incremental learning is triggered;
[0136] The business metric monitoring condition is: the domain entity recall rate is less than the set value for 3 consecutive batches, or the MSE of the transaction volume prediction exceeds 1.5 times the set threshold, or the NDCG of the supplier ranking drops by more than the preset percentage.
[0137] The specific process of using the teacher model and student model obtained by three-stage knowledge distillation for incremental learning is as follows:
[0138] Follow the strategy of the prediction head distillation stage, only update the low-rank adaptation layer, and keep the encoder and decoder frozen; if some feature extraction layers need to be fine-tuned, unfreeze the last 1-2 layers of the encoder on the basis of updating the low-rank adaptation layer, and set the corresponding learning rate to 10% of the initial learning rate;
[0139] Collect new data for incremental learning every week, and the Mini-Epoch strategy is also adopted in the incremental learning stage.
[0140] The loss function of incremental learning is as follows:
[0141]
[0142] In the formula, L2 total Represents the loss value of incremental learning, L SD Represents the semantic distribution distillation loss, and α is a weighting coefficient used to control the proportion of the topological perception attention distillation loss L TA in the total loss, and β is a weighting coefficient used to control the proportion of the domain reinforcement task loss L DT in the total loss; μ represents the smoothing coefficient, and θ S represents the parameter vector of the student model at the current training stage, represents the parameter vector saved by the student model at the previous moment or the previous incremental stage;
[0143] The expression of the semantic distribution distillation loss is as follows:
[0144]
[0145] In the formula, L SD is the semantic distribution distillation loss, τ is the temperature coefficient, KL(·) represents the Kullback-Leibler divergence, σ(·) represents the Softmax function, z T and z S respectively represent the log-odds vectors of the teacher model and the student model;
[0146] The expression of the topological perception attention distillation loss is as follows:
[0147]
[0148] In the formula, l represents the layer index in the teacher model, and L a is the set of layers selected for distillation in the teacher model, is the mapping or processing operation applied to the attention matrix of the l-th layer of the teacher model, is the mapping applied to the attention matrix ′ of the corresponding l-th layer of the student model;
[0149] The expression of the domain reinforcement task loss is as follows:
[0150]
[0151] In the formula, λ i represents the loss weight of the i-th sub-task, and L task,i represents the loss function of the i-th sub-task, and the sub-tasks include the compliance clause classification task, the procurement cost distribution prediction task, and the supplier priority ranking task;
[0152] After each training round, the loss weights of each subtask are automatically updated based on the gradient information on the validation set, as shown in the following formula:
[0153]
[0154] In the formula, represents the validation set loss L val The partial derivative of λ i is used to measure the error sensitivity of the subtask in the current model.
[0155] Through the joint optimization of semantic distribution, topological attention, and domain reinforcement tasks, the model can balance knowledge alignment under multi-modal input and the requirements of the target task, and achieve higher inference accuracy and lower model complexity in supply chain and e-commerce procurement scenarios.
[0156] Using the trained teacher model and student model for collaborative reasoning, if the prediction confidence of the student model is less than the threshold, the outputs of the teacher model and the student model are weighted and fused to obtain the results of compliance clause classification, procurement cost distribution prediction, and supplier priority ranking; otherwise, the prediction result of the student model is directly used as the final output to obtain the results of compliance clause classification, procurement cost distribution prediction, and supplier priority ranking.
[0157] If the prediction confidence of the student model (Softmax maximum value) < 0.7, the outputs of the teacher and student models are weighted and fused:
[0158] y final = 0.7·y S + 0.3·y T
[0159] In the formula, y final represents the final prediction result, y S is the prediction result of the student model, and y T is the prediction result of the teacher model.
[0160] In the multi-classification scenario, the entropy threshold (e.g., entropy > 1.5) can also be used to determine whether the uncertainty is higher than the acceptable range.
[0161] For highly controversial samples (such as compliance clause conflicts), the dual-model voting mechanism is activated to preferentially adopt the results of the teacher model. If the outputs of the teacher and student models are significantly different (e.g., different label predictions and both confidence levels > 0.6), the sample can be registered as a "high-priority feedback item" and focused on learning during the next batch of online training or manually reviewed by experts.
[0162] When the maximum confidence of the student model (Softmax output) ≥ 0.7, the prediction result y S of the student model is usually directly used as the final output y final, that is
[0163] yfi na l = yS
[0164] This indicates that the student model already has sufficient "confidence" in the current sample, and there is no need to perform teacher-student fusion or fallback to the teacher model prediction.
[0165] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in the present application can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0166] The above is only the preferred implementation manner of the present invention. The protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the concept of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the technical field, several improvements and refinements made without departing from the principle of the present invention should be regarded as within the protection scope of the present invention.
Claims
1. A method for compressing and online incremental learning of large models in the supply chain and e-commerce procurement fields based on third-order joint distillation, characterized in that, It includes the following steps: Collect structured data of supply chain compliance documents and purchase orders, and preprocess the structured data of supply chain compliance documents and purchase orders to obtain a dataset; Construct a teacher model and a student model based on a large language model respectively. Both the teacher model and the student model include an encoder, a decoder, and a prediction head; Perform three-stage knowledge distillation on the teacher model and the student model using the dataset; the specific process of the three-stage knowledge distillation is to perform hierarchical knowledge transfer in the order of the encoder distillation stage, the decoder distillation stage, and the prediction head distillation stage; Collect new data, and determine whether the new data triggers the condition for incremental learning. If so, perform incremental learning using the teacher model and the student model obtained by three-stage knowledge distillation to obtain a trained teacher model and a trained student model; Perform collaborative inference using the trained teacher model and student model. If the prediction confidence of the student model is less than the threshold, weight and fuse the outputs of the teacher model and the student model to obtain the results of compliance clause classification, procurement cost distribution prediction, and supplier priority ranking; otherwise, directly use the prediction result of the student model as the final output to obtain the results of compliance clause classification, procurement cost distribution prediction, and supplier priority ranking.
2. The method for large model compression and online incremental learning in the supply chain and e-commerce procurement fields based on third-order joint distillation according to claim 1, wherein, The preprocessing of the structured data of supply chain compliance documents and purchase orders specifically includes: Labeling clause classification labels, paragraph segmentation and regularization processing, byte pair encoding (BPE)-based word segmentation, unit unification, field alignment, and missing value imputation.
3. The method for large model compression and online incremental learning in the supply chain and e-commerce procurement fields based on third-order joint distillation as claimed in claim 1, wherein The encoder includes a self-attention layer and a feed-forward sub-layer; The specific process of the encoder distillation stage is as follows: Divide the dataset into a training set and a validation set, and configure hyperparameters, including the learning rate, batch size, maximum number of iterations, and optimizer; After the encoder is trained on the training set, use the validation set for validation and calculate the F1 score. If the fluctuation range of the F1 score is less than the set range for 5 consecutive times, the encoder distillation stage is partially convergent, and the encoder parameters are frozen, including freezing the parameters of all self-attention layers and feed-forward sub-layers. If the loss value of the validation set does not decrease after the set number of iterations, reduce the learning rate and enable gradient clipping; During the process of training the encoder, cache the intermediate representations of the teacher model locally.
4. The method for large model compression and online incremental learning in the supply chain and e-commerce procurement fields based on third-order joint distillation according to claim 1, characterized in that Append positional encoding or segment vectors to the input end of the decoder to distinguish different data modalities, where the data modalities include text, numerical values, and timestamps; or add an embedding layer and a multi-layer perceptron connected in sequence to the input end of the decoder to project numerical fields; The specific process of the decoder distillation stage includes: When processing multi-modal data at the decoder end, map different modalities to an alignable vector space for multi-modal feature fusion; Configure hyperparameters, including the learning rate and batch size; The goals of the decoder distillation stage include attention matrix alignment and semantic distribution alignment; the specific process of the attention matrix alignment is as follows: the attention matrices generated by each layer of the student model are similar to the attention matrices of the corresponding layers of the teacher model, and the similarity between the two is measured by the second norm or mean squared error; the specific process of the semantic distribution alignment is as follows: the output of the student model is aligned with the output distribution of the teacher model on the same input, and the difference between the two is measured by KL divergence or mean squared error; When the path inference accuracy or F1 score reaches the threshold, freeze the decoder parameters; otherwise, increase the number of iterations for fine-tuning the decoder. By projecting the student model attention to the teacher model space, where ψ(·) represents a projection function from the student model attention representation space to the teacher model attention representation space, and W ψ is a trainable projection matrix used to map the dimension or distribution of the student attention matrix to a space equivalent or comparable to that of the teacher attention, is the attention matrix generated by the student model at the l ′ -th layer; Update the projection matrix W ψ , add an L2 regularization term when updating the projection matrix, and update the parameters of the projection matrix during each backpropagation.
5. The method for compressing and online incremental learning of large models in the supply chain and e-commerce procurement fields based on third-order joint distillation as claimed in claim 1, wherein The specific process of the prediction head distillation stage is as follows: Update the parameters W of the low-rank adaptation layer adapt , train the prediction head, randomly extract the historical data from the encoder distillation stage and the decoder distillation stage and mix them into the current training to prevent the model from forgetting previous tasks in the final distillation stage; during the training process, every time a preset number of iterations is reached, evaluate the F1 score of the compliance clause classification task and the Wasserstein distance of the procurement cost distribution prediction task. If the evaluation metrics do not improve for 3 consecutive times, terminate the prediction head distillation.
6. The method for compressing and online incremental learning of large models in the supply chain and e-commerce procurement fields based on third-order joint distillation according to claim 1, wherein The conditions for incremental learning are specifically the data distribution difference condition or the business metric monitoring condition: The data distribution difference condition is: If the class labels of more than a preset percentage of samples in the new data exceed the distribution range of the old data, or the KL divergence difference is greater than the threshold, then incremental learning is triggered. The business metric monitoring condition is: The domain entity recall rate is less than the set value for 3 consecutive batches, or the MSE of transaction volume prediction exceeds 1.5 times the set threshold, or the NDCG of supplier ranking drops by more than a preset percentage.
7. The method for compressing and online incremental learning of large models in the supply chain and e-commerce procurement fields based on third-order joint distillation as claimed in claim 1, wherein, The specific process of using the teacher model and student model obtained through three-stage knowledge distillation for incremental learning is as follows: Adopt the strategy of the prediction head distillation stage, only update the low-rank adaptation layer, and keep the encoder and decoder frozen; if some feature extraction layers need to be fine-tuned, unfreeze the last 1-2 layers of the encoder on the basis of updating the low-rank adaptation layer, and set the corresponding learning rate to 10% of the initial learning rate. Collect new data every week for incremental learning, and the Mini-Epoch strategy is also adopted in the incremental learning stage.
8. The method for large model compression and online incremental learning in the supply chain and e-commerce procurement fields based on third-order joint distillation as claimed in claim 1, wherein, The loss function of the three-stage knowledge distillation is as follows: In the formula, represents the loss value of three-stage knowledge distillation, and L SD represents the semantic distribution distillation loss, α is a weighting coefficient used to control the proportion of the topological perception attention distillation loss L TA in the total loss, and β is a weighting coefficient used to control the proportion of the domain reinforcement task loss L DT in the total loss; The expression of the semantic distribution distillation loss is as follows: where L SD is the semantic distribution distillation loss, τ is the temperature coefficient, KL(·) represents the Kullback-Leibler divergence, σ(·) represents the Softmax function, z T and z S represent the logit vectors of the teacher model and the student model, respectively; The expression of the topology-aware attention distillation loss is as follows: where \(l\) represents the layer index in the teacher model, and \(L\) a is the set of layers selected for distillation in the teacher model, is the mapping or processing operation applied to the attention matrix of the \(l\)-th layer of the teacher model, and ′ is the mapping applied to the attention matrix of the corresponding \(l\)-th layer of the student model. The expression of the domain reinforcement task loss is as follows: and where λ i represents the loss weight of the i-th sub-task, and L task,i represents the loss function of the i-th sub-task. The sub-tasks include compliance clause classification tasks, procurement cost distribution prediction tasks, and supplier priority ranking tasks; After each training epoch ends, automatically update the loss weights of each subtask through the gradient information on the validation set. The formula is as follows: In the formula, represents the validation set loss L val is the partial derivative with respect to λ i and is used to measure the error sensitivity of the subtask in the current model.
9. The method for large model compression and online incremental learning in the supply chain and e-commerce procurement fields based on third-order joint distillation according to claim 1, wherein, The loss function of the incremental learning is as follows: Wherein, L 2 total represents the loss value of incremental learning, L SD represents the semantic distribution distillation loss, and α is a weighting coefficient used to control the proportion of the topological perception attention distillation loss L TA in the total loss, and β is a weighting coefficient used to control the proportion of the domain reinforcement task loss L DT in the total loss; μ represents the smoothing coefficient, and θ S represents the parameter vector of the student model at the current training stage, and represents the parameter vector saved by the student model at the previous moment or the previous incremental stage; The expression of the semantic distribution distillation loss is as follows: where L SD is the semantic distribution distillation loss, τ is the temperature coefficient, KL(·) represents the Kullback-Leibler divergence, σ(·) represents the Softmax function, z T and z S represent the logit vectors of the teacher model and the student model, respectively; The expression of the topology-aware attention distillation loss is as follows: where \(l\) represents the layer index in the teacher model, and \(L\) a is the set of layers selected for distillation in the teacher model, is the mapping or processing operation applied to the attention matrix of the \(l\)-th layer of the teacher model, and ′ is the mapping applied to the attention matrix of the corresponding \(l\)-th layer of the student model; The expression of the domain reinforcement task loss is as follows: and where λ i represents the loss weight of the i-th sub-task, and L task,i represents the loss function of the i-th sub-task. The sub-tasks include compliance clause classification tasks, procurement cost distribution prediction tasks, and supplier priority ranking tasks; After each training epoch ends, automatically update the loss weights of each subtask through the gradient information on the validation set. The formula is as follows: In the formula, represents the validation set loss L val partial derivative of i with respect to λ, which is used to measure the error sensitivity of the subtask in the current model.
10. The method for compressing and online incremental learning of large models in the supply chain and e-commerce procurement fields based on third-order joint distillation according to claim 8, wherein, The loss of the compliance clause classification task is measured by the cross-entropy loss. The loss function of the compliance clause classification task is as follows: where L task,1 represents the loss of the compliance clause classification task, and L CE represents the cross-entropy loss, and k represents the compliance category index The loss of the procurement cost distribution prediction task is measured by the Wasserstein distance. The loss function of the procurement cost distribution prediction task is as follows: where L task,2 represents the loss of the procurement cost distribution prediction task, and L EMD represents the cumulative distance between the predicted distribution and the true distribution. i is the cost interval index, and CDF pred (i) represents the value of the predicted cumulative distribution function at the interval index i, and CDF true (i) represents the value of the true cumulative distribution function at the interval index i; The loss of the supplier priority ranking task is measured by the ranking loss. The loss function of the supplier priority ranking task is as follows: Where, L task,3 is the loss of the supplier priority ranking task, and L Rank represents the ranking loss. s i and s j respectively represent the ranking scores of supplier i and supplier j predicted by the model. y (i,j) represents the preference label of supplier i and j. If i performs better than j, then y (i,j) = +1; If j is better than i, then y (i,j) = -1, and pair(i, j) refers to all comparable supplier pairs in the training set.
Citation Information
Patent Citations
Pre-trained language model compression method and platform based on Knowledge distillation
CN111767711A
Knowledge distillation-based end-to-end speech recognition incremental learning method and system
CN115064155A
Track target point prediction method based on knowledge distillation
CN116579423A
Incremental learning method based on knowledge distillation and parameter isolation
CN116883783A
Knowledge distillation method and system based on multi-teacher multi-modal model
CN117669693A
Cited By
Hierarchical identification method, device and system for electronic data trust intensity authentication and medium
CN121071434A
Electronic data trust strength authentication hierarchical identification method, device, system and medium
CN121071434B