Lightweight Natural Language Processing Large Model Training Method

Through dynamic sparse activation and mixed precision quantization optimization of natural language processing large model, the static strategy rigidity and quantitative accuracy loss in the lightweight process are solved, and the model is efficiently adaptable and accurate in complex semantic tasks are achieved.

CN119862925BActive Publication Date: 2025-07-11GUIZHOU NORMAL UNIVERSITY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510355279.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-11
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

When the existing technology lightens the natural language processing large models, there are problems such as rigid static strategies, loss of quantization accuracy and inefficient knowledge transfer, resulting in poor performance of the model in complex semantic tasks.

Method used

Through dynamic sparse activation, hybrid precision quantization and collaborative optimization, the parameter participation degree is adaptively adjusted by semantic complexity, combined with sparse activation mask and quantization bit width, cross-feedback adjustment of the hybrid precision quantization strategy is achieved, and the model training process is optimized.

Benefits of technology

The calculation efficiency and accuracy of the model are improved, and the problems of rigid static strategies, loss of quantitative accuracy and inefficient knowledge transfer are solved, so as to achieve efficient adaptability of the model in complex semantic tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119862925B_ABST
    Figure CN119862925B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for training a lightweight natural language processing large model; it includes the following steps: obtaining processed language data; obtaining an enhanced labeled data set; calculating an activation mask by dynamically activating the sparsification mechanism of the sub-network through semantic complexity; generating a quantization bit width based on the parameter sensitivity of the activation mask; performing cross-feedback adjustment on the mixed-precision quantization strategy; and evaluating the trained student model. Through dynamic sparse activation, mixed-precision quantization, and collaborative optimization, this application solves the core problems in large model lightweighting such as static strategy rigidity, quantization accuracy loss, and inefficient knowledge transfer; dynamic sparse activation replaces traditional static pruning to reduce semantic loss; in order to achieve optimized feature extraction for the enhanced data set, fusion-optimized features are adopted; and mixed-precision quantization effectively reduces the computational complexity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of natural language models, and more specifically, particularly relates to a training method for a lightweight natural language processing large model. Background Art

[0002] In recent years, significant progress has been made in the field of natural language processing (NLP) through large-scale pre-trained models (such as BERT, GPT-3). However, with the exponential growth of the number of model parameters (from hundreds of millions to trillions), training and deploying these large models face huge computational resource consumption and storage costs. The current mainstream technical solutions have the following key problems during the lightweight process:

[0003] Traditional model compression methods (such as pruning, parameter freezing) rely on static strategies, resulting in limited feature expression capabilities; fixed pruning rate problem: typical pruning methods compress by removing low-magnitude weights, but the model cannot dynamically recover the removed parameters after pruning. Poor task adaptability: the performance of the model after static pruning deteriorates significantly in complex tasks (such as multilingual translation, long text generation).

[0004] Existing quantization techniques use the same bit width for all parameters, ignoring the differences in parameter sensitivity; performance collapse of sensitive layers: key parameters in the attention mechanism (such as the Query-Key projection matrix) are sensitive to quantization errors. Insufficient dynamic range adaptation: uniform quantization cannot adapt to the heterogeneity of parameter distributions.

[0005] Traditional distillation methods fail when the ability gap between the teacher and student models is large; capacity mismatch problem: when the gap in the number of parameters between the teacher model and the student model exceeds 100 times, the student model cannot effectively imitate the behavior of the teacher, resulting in a 23% increase in the perplexity of text generation after distillation. Information loss bottleneck: single-stage distillation is difficult to balance knowledge transfer at different granularities.

[0006] Deficiencies of the existing technical solutions:

[0007] How to maintain adaptability to complex semantic tasks while compressing the model scale. Existing methods lack a dynamic perception mechanism and cross-module collaborative optimization, resulting in the following problems: static compression strategies cannot respond to the dynamic changes of input semantics; coarse-grained quantization causes performance degradation of key modules; knowledge residue occurs during the distillation process due to the ability gap between the teacher and student; therefore, we propose a training method for a lightweight natural language processing large model. Summary of the Invention

[0008] Aiming at the problems existing in the prior art, the purpose of the present invention is to provide a lightweight natural language processing large model training method, which adaptively adjusts the parameter participation degree through the input semantic complexity, allocates appropriate quantization bit widths based on the parameter sensitivity, and realizes the joint training of sparse activation, mixed quantization and knowledge distillation, thereby solving the problems of static strategy rigidity, quantization accuracy loss, and low knowledge transfer efficiency in the lightweight of large models, and effectively improving the computing efficiency and accuracy of the model.

[0009] To achieve the above object, the present invention provides the following technical solutions: A lightweight natural language processing large model training method, including the following steps:

[0010] Collect a large amount of natural language data, preprocess the natural language data to obtain processed language data;

[0011] Select a mature large natural language model as the teacher model, and train and process the processed language data through the teacher model to obtain an annotated enhanced data set;

[0012] Take the lightweight natural language processing large model as the student model, then input the enhanced data set into the student model for training, calculate the semantic complexity of the enhanced data set, and dynamically activate the sparsification mechanism of the sub-network through the semantic complexity to calculate the activation mask;

[0013] Input the calculated activation mask into a preset mixed-precision quantization strategy, calculate the parameter sensitivity of the activation mask, and generate a quantization bit width through the parameter sensitivity of the activation mask;

[0014] Then combine the activation mask and the quantization bit width to jointly complete the calculation, and realize the cross-feedback adjustment of the mixed-precision quantization strategy;

[0015] Input the natural language data into the student model, the student model outputs the training result, evaluate the trained student model based on the training result, and output the current student model, which is the natural language processing model for natural language processing, when the evaluation result meets the predetermined conditions.

[0016] Further preferably, the acquisition of the processed language data is to preprocess the natural language data, and the preprocessing steps include text cleaning, spelling correction, removal of duplicate data, and text cutting and splitting;

[0017] The text cleaning is used to remove special symbols, line breaks and spaces from the natural language data;

[0018] The spelling correction is used to correct errors in the natural language data and unify the case format of the text in the natural language data;

[0019] The duplicate data removal is used to remove texts with duplicates or samples with the same annotations in natural language data, which helps to avoid model overfitting;

[0020] The text cutting and splitting is used to cut or split long texts in natural language data. Especially for models with a maximum length limit, by splitting long texts into multiple paragraphs or sentences, each part is convenient for model processing.

[0021] Further preferably, the labeled enhanced dataset is obtained by the following method:

[0022] The teacher model generates pseudo-labels: Use the teacher model to predict or annotate the processed language data to generate pseudo-label data;

[0023] The screening of pseudo-labels: Screen the generated pseudo-label data, and filter out low-confidence pseudo-label data by setting a confidence level;

[0024] Enhanced dataset: Combine the pseudo-label data and the processed language data to generate a new enhanced dataset.

[0025] Further preferably, the calculation formula for setting the confidence level is as follows:

[0026] The confidence level setting uses a dynamic threshold as the threshold, which is set according to the average confidence level of all natural language data output by the teacher model. The formula is as follows:

[0027] ;

[0028] Among them, is the average confidence level, is the quantity of the processed language data, is the th predicted confidence level of the processed language data.

[0029] Further preferably, the steps of the sparsification mechanism of the sub-network are as follows:

[0030] Extract features from the input labeled enhanced dataset, and fuse the extracted features to generate fused features, and then generate multi-dimensional semantic representations, and use the multi-dimensional semantic representations as the decision basis for sparse activation;

[0031] Dynamically select the activated sub-network according to the input fused features, and calculate the activation mask according to the semantic complexity;

[0032] And the steps for feature extraction of the enhanced dataset are as follows:

[0033] The augmented dataset is segmented, and feature extraction is performed on the segmented augmented dataset. Then, the extracted features are fused, and further feature extraction is performed on the fused features;

[0034] First, feature extraction is performed on each character in the augmented dataset, then the extracted features of each one are fused, and further feature extraction is performed to achieve optimized features for all characters;

[0035] And feature extraction is performed on each word in the augmented dataset, then the extracted features of each one are fused, and further feature extraction is performed to achieve optimized features for all words;

[0036] Furthermore, the augmented dataset is segmented. It is evenly segmented by character numbers, and then the augmented dataset is split into two segments. Then, feature extraction is performed on each segment of the augmented dataset and the complete augmented dataset, and then the extracted features are fused, and further feature extraction is performed to achieve optimized features for context features;

[0037] Finally, the optimized character features, optimized word features, and optimized context features are fused to obtain fused features.

[0038] Further preferably, the fused features include character-level features, word-level features, and context features. The augmented dataset uses a lightweight encoder to extract character-level features, word-level features, and context features;

[0039] The calculation formula for the fused features is as follows:

[0040] ;

[0041] Where, represents the fused features, represents the character-level features, represents the word-level features, represents the context features.

[0042] Further preferably, the semantic complexity uses a two-channel feature extractor to perform context dimension analysis processing on the augmented dataset and complete a real-time entropy value calculation engine;

[0043] The analysis calculation of the context dimension is as follows:

[0044] ;

[0045] Where, represents the context dimension features of the augmented dataset, represents the hidden state output by the augmented dataset through the basic Transformer encoder; Document input represented as an enhanced dataset, represented as multi-head self-attention features, represented as hierarchical LSTM features, represented as a document topic vector;

[0046] Dual-channel fusion strategy:

[0047] Dynamic gating fusion algorithm:

[0048] ;

[0049] Among them, represented as multi-dimensional semantic representations, represented as a learnable weight matrix, represented as the Sigmoid activation function, represented as text dimension representations, represented as context dimension representations;

[0050] The calculation of the real-time entropy value calculation engine is as follows:

[0051] Semantic entropy definition and calculation:

[0052] Multi-granularity entropy value fusion formula:

[0053] ;

[0054] Among them, represented as multi-granularity entropy values, represented as Token-level entropy, represented as sentence-level entropy, represented as document-level entropy, respectively represent the weight coefficients of Token-level entropy, the weight coefficients of sentence-level entropy, and the weight coefficients of document-level entropy;

[0055] Among them, the entropy calculation of each level:

[0056] Token-level entropy:

[0057]

[0058] Among them, represented as predefined semantic categories, represented as the number of Tokens in the enhanced dataset, represented as a given Token , which belongs to the category probability, represented as the th Token;

[0059] Sentence-level entropy:

[0060] ;

[0061] Among them, represents the depth distribution of the enhanced dataset, represents the uniform distribution, represents the number of sentences being currently processed, represents the KL divergence, measuring the difference between the actual distribution and the uniform distribution ;

[0062] Document-level entropy:

[0063] ;

[0064] Among them, represents the topic distribution probability of the enhanced dataset, represents the number of topics.

[0065] Further preferably, the parameter sensitivity of the activation mask generates a quantization bit width, and the calculation of the parameter sensitivity of the activation mask is as follows:

[0066] ;

[0067] Among them, the activation frequency of the activation mask is the number of times activated in the most recent steps, = 0.9 is the decay coefficient, represents the exponentially weighted activation frequency of the previous time step , represents the indicator function, indicating whether the parameter is activated at the time step , represents the weight coefficient of the current activation state, complementary to ;

[0068] Parameter sensitivity calculation:

[0069] ;

[0070] Among them, represents the parameter sensitivity of the th activation mask, and represent the weight coefficients, used to balance the contributions of the activation frequency and the parameter sensitivity, and = 0.6, = 0.4, represents the normalized activation frequency, is expressed as the normalized gradient sensitivity;

[0071] The calculation of the quantization bit width is as follows:

[0072] ;

[0073] where is expressed as the value of the quantization bit width, respectively represent the lowest dynamic threshold and the highest dynamic threshold;

[0074] , where is expressed as the 30th percentile of the parameter sensitivity, is expressed as the 70th percentile of the parameter sensitivity.

[0075] Further preferably, the steps of the cross-feedback regulation are as follows:

[0076] Forward propagation: Generate an activation mask according to the input, filter the activated sub-networks, load the quantized weights of the activation parameters, and perform calculations;

[0077] Quantization-aware training: In backpropagation, calculate the pseudo-gradient of the quantization parameters and update the activation frequency and the gradient sensitivity ;

[0078] Dynamic adjustment: Recalculate the sensitivity percentile every step, update the bit width allocation; Use a sliding window to count the activation patterns and optimize the mask generation strategy.

[0079] Further preferably, the calculation for evaluating the student model is as follows:

[0080] ,

[0081] where is the posterior probability, that is, the similarity between the output of the student model and the augmented dataset. N is the total amount of training results for comparison and determination, is expressed as the training result, is expressed as the position sequence of the training result, is the likelihood probability of the similarity between the training result and the augmented dataset, is the prior probability of the similarity between the training result and the augmented dataset, is the likelihood probability of the non-similarity between the training result and the augmented dataset, is the prior probability of the non-similarity between the training result and the augmented dataset.

[0082] The technical effects and advantages of the present invention:

[0083] This application solves the core problems such as static policy rigidity, quantization precision loss, and inefficient knowledge transfer in the lightweighting of large models through dynamic sparse activation, mixed-precision quantization, and collaborative optimization;

[0084] Dynamic sparse activation extracts features from natural language text, fuses the extracted features as the decision basis for sparse activation, calculates the activation mask through semantic complexity, and the semantic complexity is based on the semantic entropy value calculation engine of the input text. The semantic complexity performs characterization analysis of the text dimension and context dimension through a dual-channel feature extractor, and then performs fusion processing to achieve dynamic routing based on real-time semantic analysis, replacing traditional static pruning and reducing semantic loss;

[0085] And in order to achieve optimized feature extraction for the augmented dataset, fused optimized features are adopted, that is, character, vocabulary, and context features are extracted from the augmented dataset, and then the extracted features are fused. Feature extraction is performed on the fused features again to achieve feature optimization, reduce the complexity of analysis and processing of the augmented dataset in large natural language processing models, and improve the efficiency of natural language analysis and processing;

[0086] Mixed-precision quantization calculates the sensitivity of computational parameter gradients through the activation mask and obtains the quantization bit width, that is, calculates the sensitivity of computational parameter gradients through the activation mask, divides the parameters into three levels: critical, medium, and non-critical, effectively reduces the complexity of the calculation, and supports a quantization framework with dynamic bit width switching to improve the accuracy of the model.

[0087] Through dynamic sparsity, mixed quantization, and bidirectional cross-feedback adjustment, through joint optimization processing, the optimal balance between model efficiency and accuracy is achieved, and the core problems such as static policy rigidity, quantization precision loss, and inefficient knowledge transfer in the lightweighting of large models are solved.

[0088] Through the following detailed description of the exemplary embodiments of the present invention with reference to the accompanying drawings, other features and advantages of the present invention will become clear. Brief Description of the Drawings

[0089] Figure 1 is a schematic flowchart of the steps provided by the present invention. Detailed Embodiments

[0090] In order to make the purpose, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0091] As Figure 1 shown, the lightweight natural language processing large model training method provided by the embodiments of the present invention includes the following steps:

[0092] Collect a large amount of natural language data, preprocess the natural language data to obtain processed language data;

[0093] Select a mature large natural language model as the teacher model, and train and process the processed language data through the teacher model to obtain an annotated enhanced data set;

[0094] Take the lightweight natural language processing large model as the student model, then input the enhanced data set into the student model for training, calculate the semantic complexity of the enhanced data set, and dynamically activate the sparsification mechanism of the sub-network through the semantic complexity to calculate the activation mask;

[0095] Input the calculated activation mask into a preset mixed-precision quantization strategy, calculate the parameter sensitivity of the activation mask, and generate a quantization bit width through the parameter sensitivity of the activation mask;

[0096] Then combine the activation mask and the quantization bit width to jointly complete the calculation, and realize cross-feedback adjustment of the mixed-precision quantization strategy;

[0097] Input the natural language data into the student model, the student model outputs the training result, evaluate the trained student model based on the training result, and output the current student model, which is the natural language processing model, for natural language processing when the evaluation result meets the predetermined conditions.

[0098] In this embodiment, preferably, the acquisition of the processed language data is to preprocess the natural language data, and the preprocessing steps include text cleaning, spelling correction, removing duplicate data, and text cutting and splitting;

[0099] The text cleaning is used to remove special symbols, line breaks, and spaces from the natural language data;

[0100] The spelling correction is used to correct errors in the natural language data and unify the case format of the text in the natural language data;

[0101] The removal of duplicate data is used to remove texts with duplicates or samples with the same annotations in the natural language data, which helps to avoid model overfitting;

[0102] The text cutting and splitting is used to cut or split long texts in the natural language data. Especially for models with a maximum length limit, by splitting the long text into multiple paragraphs or sentences, each part is convenient for the model to process;

[0103] It should be noted that by preprocessing natural language data, the format uniformity of natural language can be improved, which is convenient for inputting natural language into a large natural language processing model for analysis and processing. Moreover, the preprocessed natural language can improve the processing efficiency of the large natural language model, reduce errors, and improve accuracy.

[0104] In this embodiment, preferably, the labeled enhanced dataset is obtained by the following method:

[0105] The teacher model generates pseudo-labels: Use the teacher model to predict or label the processed language data to generate pseudo-label data;

[0106] Screening of pseudo-labels: Screen the generated pseudo-label data, and filter out the low-confidence pseudo-label data by setting the confidence level;

[0107] Enhanced dataset: Combine the pseudo-label data and the processed language data to generate a new enhanced dataset;

[0108] It should be noted that by using the teacher model to process the processed language data to obtain an enhanced dataset, and through the training of the enhanced dataset in the student model, the student model can obtain better generalization ability and performance by imitating the teacher model, especially in the case of small sample data or resource-constrained environments. And by setting the confidence level to filter out low-confidence samples, the accuracy of the enhanced dataset can be effectively improved, which is convenient for the student model to perform rapid training and use.

[0109] In this embodiment, preferably, the formula for setting the confidence level is as follows:

[0110] The set confidence level uses a dynamic threshold as the threshold, which is set according to the average confidence level of all natural language data output by the teacher model. The formula is as follows:

[0111] ;

[0112] Where is the average confidence level, is the number of processed language data, is the th predicted confidence level of the processed language data;

[0113] It should be noted that the confidence level is set using a dynamic threshold, which can avoid transmitting unreliable information, focus the training of the student model on the more reliable predictions of the teacher model, improve the efficiency of distillation, and enable the student model to better adapt to different types of samples by dynamically adjusting the intensity of information transmission, thereby improving its generalization ability. The model can automatically adjust the distillation process according to different data distributions to adapt to more complex training scenarios. Using a dynamic threshold as the threshold for the confidence level of the teacher model is a technique that can enhance the knowledge distillation process. By dynamically adjusting the threshold according to the characteristics of different input samples or the prediction confidence level of the teacher model, the student model can selectively learn the knowledge of the teacher model and improve the effect of distillation.

[0114] In this embodiment, preferably, the steps of the sparsification mechanism of the sub-network are as follows:

[0115] Extract features from the input labeled augmented dataset, and fuse the extracted features to generate fused features, and then generate multi-dimensional semantic representations, and use the multi-dimensional semantic representations as the decision basis for sparse activation;

[0116] Dynamically select the activated sub-network according to the input fused features, and calculate the activation mask according to the semantic complexity;

[0117] And the steps of feature extraction for the augmented dataset are as follows:

[0118] Segment the augmented dataset, extract features from the segmented augmented dataset, then fuse the extracted features, and then perform feature extraction again on the fused features;

[0119] First, extract features from each character in the augmented dataset, then fuse each extracted feature, and perform feature extraction again to optimize the features of all characters;

[0120] And extract features from each word in the augmented dataset, then fuse each extracted feature, and perform feature extraction again to optimize the features of all words;

[0121] And then segment the augmented dataset, segment it evenly by character numbers, then divide the augmented dataset into two segments, then extract features from each segment of the augmented dataset and the complete augmented dataset, then fuse the extracted features, and perform feature extraction again to optimize the context features;

[0122] Finally, fuse the optimized character features, optimized word features, and optimized context features to obtain fused features;

[0123] It should be noted that by means of the sparse activation mechanism to improve the computational efficiency and performance of the network, through the strategies of feature extraction, feature fusion, multi-dimensional semantic representation generation, and dynamic selection of sub-network modules, the model can intelligently select which modules are the most important under a given input, thereby reducing unnecessary computations. With the support of semantic complexity, the model can more precisely control the activated sub-networks and improve the efficiency of training and inference, especially when dealing with complex or diverse data.

[0124] In this embodiment, preferably, the fused features include character-level features, word-level features, and context features, and the enhanced dataset uses a lightweight encoder to extract character-level features, word-level features, and context features;

[0125] The calculation formula for the fused features is as follows:

[0126] ;

[0127] Wherein, represents the fused feature, represents the character-level feature, represents the word-level feature, represents the context feature;

[0128] It should be noted that by analyzing the character-level features, word-level features, and context features of the enhanced dataset, the fine-grained understanding of the enhanced dataset can be improved, and the feature fusion of the enhanced household dataset can be realized to obtain the feature information of the enhanced dataset, which is convenient for subsequent computational processing.

[0129] In this embodiment, preferably, the semantic complexity uses a dual-channel feature extractor to perform context dimension analysis and processing on the enhanced dataset, and complete a real-time entropy value calculation engine;

[0130] The analysis and calculation of the context dimension are as follows:

[0131] ;

[0132] Wherein, represents the context dimension feature of the enhanced dataset, represents the hidden state output by the enhanced dataset through the basic Transformer encoder; represents the document input of the enhanced dataset, represents the multi-head self-attention feature, represents the hierarchical LSTM feature, represents the document topic vector;

[0133] Dual-channel fusion strategy:

[0134] Dynamic Gating Fusion Algorithm:

[0135] ;

[0136] Among them, is represented as a multi-dimensional semantic representation, is represented as a learnable weight matrix, is represented as the Sigmoid activation function, is represented as the text dimension representation, is represented as the context dimension representation;

[0137] The calculation of the real-time entropy value calculation engine is as follows:

[0138] Definition and calculation of semantic entropy:

[0139] Multi-granularity entropy fusion formula:

[0140] ;

[0141] Among them, is represented as an activation mask, is represented as the Token-level entropy, is represented as the sentence-level entropy, is represented as the document-level entropy, respectively represent the weight coefficients of the Token-level entropy, the weight coefficients of the sentence-level entropy, and the weight coefficients of the document-level entropy;

[0142] Among them, the entropy calculation of each level:

[0143] Token-level entropy:

[0144]

[0145] Among them, is represented as a predefined semantic category, is represented as the number of Tokens in the enhanced dataset, is represented as a given Token , which belongs to the category probability, is represented as the th Token;

[0146] Sentence-level entropy:

[0147] ;

[0148] Among them, is represented as the depth distribution of the enhanced dataset, is represented as the uniform distribution, is represented as the number of sentences in the current processing, Expressed as the KL divergence, it measures the actual distribution and the uniform distribution difference;

[0149] Document-level entropy:

[0150] ;

[0151] where, represents the topic distribution probability of the enhanced dataset, represents the number of topics;

[0152] Dynamic threshold control strategy:

[0153] ;

[0154] where, = 0.9 (momentum coefficient), EMA is the exponential moving average, is the entropy value sequence of the last 10 steps, represents the dynamic threshold;

[0155] It should be noted that through the dual-channel fusion strategy, feature fusion processing of the text up and down and context dimensions is achieved to obtain multi-dimensional semantic representations, and through the dual-channel fusion strategy, fusion processing is achieved through multi-dimensional semantic representations and multi-level entropy, and the activation mask is calculated to effectively achieve pruning processing of the enhanced dataset and reduce the loss of semantic information.

[0156] In this embodiment, preferably, the parameter sensitivity of the activation mask generates a quantization bit width, and the calculation of the parameter sensitivity of the activation mask is as follows:

[0157] ;

[0158] where, the activation frequency of the activation mask is the number of times activated in the last steps, = 0.9 is the decay coefficient, represents the exponentially weighted activation frequency of the previous time step , represents the indicator function, indicating whether the parameter is activated at the time step , represents the weight coefficient of the current activation state, complementary to ;

[0159] Parameter sensitivity calculation:

[0160] ;

[0161] where, Denoted as the parameter sensitivity of the activation mask, and denoted as the weight coefficient, used to balance the contributions of the activation frequency and the parameter sensitivity, and = 0.6, = 0.4, denoted as the normalized activation frequency, denoted as the normalized gradient sensitivity;

[0162] The calculation of the quantization bit width is as follows:

[0163] ;

[0164] where denoted as the value of the quantization bit width, respectively denoted as the lowest dynamic threshold and the highest dynamic threshold;

[0165] , where denoted as the 30th percentile of the parameter sensitivity, denoted as the 70th percentile of the parameter sensitivity;

[0166] It should be noted that by calculating the parameter sensitivity of the activation mask, calculating the parameter sensitivity through the gradient sensitivity, and calculating and processing the quantization bit width through the parameter sensitivity, setting the dynamic threshold, the key parameter accuracy can be retained for mixed quantization, the performance degradation of uniform quantization can be avoided, and the video memory occupancy can be reduced.

[0167] In this embodiment, preferably, the steps of the cross-feedback regulation are as follows:

[0168] Forward propagation: Generate an activation mask according to the input, filter the activated sub-network, load the quantized weights of the activation parameters, and perform calculations;

[0169] Quantization-aware training: In the backpropagation, calculate the pseudo-gradient of the quantization parameters and update the activation frequency and the gradient sensitivity ;

[0170] Dynamic adjustment: Recalculate the sensitivity percentile every step, update the bit width allocation; Use a sliding window to count the activation patterns and optimize the mask generation strategy;

[0171] It should be noted that through the dynamic sensitivity scoring, mask-aware quantization allocation, and two-way feedback regulation, the optimal balance between the model efficiency and the accuracy is achieved.

[0172] In this embodiment, preferably, the calculation for evaluating the student model is as follows:

[0173] ,

[0174] where, is the posterior probability, that is, the similarity between the output of the student model and the augmented dataset. N is the total amount of training results for comparison and determination. represents the training result. represents the position sequence of the training result. is the likelihood probability of the similarity between the training result and the augmented dataset. is the prior probability of the similarity between the training result and the augmented dataset. is the likelihood probability of the non - similarity between the training result and the augmented dataset. is the prior probability of the non - similarity between the training result and the augmented dataset;

[0175] It should be noted that through the calculation and analysis of the evaluation, the output of the student model can be matched with the result of the teacher model, which is convenient for obtaining the accuracy of the student model. When it is determined that the matching probability between the student model and the teacher model reaches 80%, it is determined that the student model can be used, and the use of the student model output is realized.

[0176] Finally, it should be noted that the above - mentioned are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for training a lightweight natural language processing large model, characterized in that It includes the following steps: Collect a large amount of natural language data, preprocess the natural language data to obtain processed language data; Select a mature large natural language model as the teacher model, and train and process the processed language data through the teacher model to obtain an annotated enhanced data set; Use the lightweight natural language processing large model as the student model, then input the enhanced data set into the student model for training, calculate the semantic complexity of the enhanced data set, and dynamically activate the sparsification mechanism of the sub-network through the semantic complexity to calculate the activation mask; The steps of the sparsification mechanism of the sub-network are as follows: Extract features from the input annotated enhanced data set, and fuse the extracted features to generate fused features, and then generate multi-dimensional semantic representations, and use the multi-dimensional semantic representations as the decision basis for sparse activation; Dynamically select the activated sub-network according to the input fused features, and calculate the activation mask according to the semantic complexity; And the steps of feature extraction for the enhanced data set are as follows: Segment the enhanced data set, extract features from the segmented enhanced data set, then fuse the extracted features, and perform feature extraction again on the fused features; First, extract features from each character in the enhanced data set, then fuse each extracted feature, and perform feature extraction again to optimize the features of all characters; And extract features from each word in the enhanced data set, then fuse each extracted feature, and perform feature extraction again to optimize the features of all words; Then segment the enhanced data set again, perform average segmentation by character numbers, then split the enhanced data set into two segments, then extract features from each segment of the enhanced data set and the complete enhanced data set, then fuse the extracted features, and perform feature extraction again to optimize the context features; Finally, fuse the optimized character features, optimized word features, and optimized context features to obtain fused features; Input the calculated activation mask into a preset mixed-precision quantization strategy, calculate the parameter sensitivity of the activation mask, and generate a quantization bit width through the parameter sensitivity of the activation mask; Then combine the activation mask and the quantization bit width to jointly complete the calculation, and realize cross-feedback adjustment of the mixed-precision quantization strategy; Input the natural language data into the student model, the student model outputs the training result, evaluate the trained student model based on the training result, and output the current student model, which is the natural language processing model, for natural language processing when the evaluation result meets the predetermined conditions.

2. The lightweight natural language processing large model training method according to claim 1, characterized in that: The acquisition of the processed language data is to preprocess the natural language data, and the steps of the preprocessing include text cleaning, spelling correction, removing duplicate data, and text cutting and splitting; The text cleaning is used to remove special symbols, line breaks, and spaces from the natural language data; The spelling correction is used to correct errors in the natural language data and unify the case format of the text in the natural language data; The duplicate data removal is used to remove texts with duplicates or samples with the same annotations in natural language data, which helps to avoid model overfitting; The text cutting and splitting is used to cut or split long texts in natural language data. Especially for models with a maximum length limit, by splitting long texts into multiple paragraphs or sentences, each part is convenient for model processing.

3. The lightweight natural language processing large model training method according to claim 1, characterized in that: The labeled enhanced dataset is obtained through the following method: The teacher model generates pseudo-labels: Use the teacher model to predict or annotate the processed language data to generate pseudo-label data; The screening of pseudo-labels: Screen the generated pseudo-label data, and filter out low-confidence pseudo-label data by setting the confidence level; Enhanced dataset: Combine the pseudo-label data and the processed language data to generate a new enhanced dataset.

4. The lightweight natural language processing large model training method according to claim 3, wherein: The calculation formula for setting the confidence level is as follows: The confidence level setting uses a dynamic threshold as the threshold, which is set according to the average confidence level of all natural language data output by the teacher model. The formula is as follows: ; Among them, is the average confidence level, is the quantity of processed language data, is the predicted confidence level of the nth processed language data.

5. The lightweight natural language processing large model training method according to claim 1, characterized in that: The fused features include character-level features, word-level features, and context features. The enhanced dataset uses a lightweight encoder to extract character-level features, word-level features, and context features; The calculation formula for the fused features is as follows: ; Among them, is represented as a fused feature, is represented as a character-level feature, is represented as a word-level feature, is represented as a context feature.

6. The lightweight natural language processing large model training method according to claim 1, wherein: The semantic complexity uses a two-channel feature extractor to perform context dimension analysis and processing on the enhanced dataset, and complete a real-time entropy calculation engine; The analysis and calculation of the context dimension are as follows: ; Among them, represents enhancing the context dimension features of the dataset, represents the hidden states output by the base Transformer encoder for the enhanced dataset; represents the document input of the enhanced dataset, represents the multi-head self-attention features, represents the hierarchical LSTM features, represents the document topic vector; Two-channel fusion strategy: Dynamic gating fusion algorithm: ; Among them, is represented as a multi-dimensional semantic representation, is represented as a learnable weight matrix, is represented as a Sigmoid activation function, is represented as a text dimension representation, is represented as a context dimension representation; The calculation of the real-time entropy calculation engine is as follows: Semantic entropy definition and calculation: Multi-granularity entropy fusion formula: ; Among them, is expressed as multi-granularity entropy value, is expressed as Token-level entropy, is expressed as sentence-level entropy, is expressed as document-level entropy, respectively represent the weight coefficients of Token-level entropy, the weight coefficients of sentence-level entropy, and the weight coefficients of document-level entropy; Among them, the entropy calculation at each level: Token-level entropy: Among them, is represented as a predefined semantic category, is represented as the number of Tokens in the enhanced dataset, is represented as a given Token , which belongs to the category with a probability of is represented as the th Token; Sentence-level entropy: ; Among them, represents the depth distribution of the enhanced dataset, represents the uniform distribution, represents the number of sentences being currently processed, represents the KL divergence, measuring the actual distribution and the uniform distribution difference; Document-level entropy: ; Among them, represents the topic distribution probability of the enhanced dataset, represents the number of topics.

7. The lightweight natural language processing large model training method according to claim 6, characterized in that: The parameter sensitivity of the activation mask generates a quantization bit width, and the calculation of the parameter sensitivity of the activation mask is as follows: ; among them, the activation frequency of the activation mask is the number of times activated in the most recent steps, = 0.9 is the attenuation coefficient, which represents the exponentially weighted activation frequency of the previous time step ; represents the indicator function, indicating whether the parameter is activated at time step ; represents the weight coefficient of the current activation state, which is complementary to ; Parameter sensitivity calculation: ; wherein, is expressed as the parameter sensitivity of the th activation mask, and is expressed as a weight coefficient for balancing the contributions of the activation frequency and the parameter sensitivity, and = 0.6, = 0.4, is expressed as the normalized activation frequency, is expressed as the normalized gradient sensitivity; The calculation of the quantization bit width is as follows: ; Among them, represents a value of quantization bit width, respectively represent the lowest dynamic threshold and the highest dynamic threshold; , where is represented as the 30th percentile of the parameter sensitivity, is represented as the 70th percentile of the parameter sensitivity.

8. The lightweight natural language processing large model training method according to claim 7, wherein: The steps of the cross-feedback regulation are as follows: Forward propagation: Generate an activation mask according to the input, screen the activated sub-networks, load the quantized weights of the activation parameters, and perform calculations; Quantization-aware training: In backpropagation, calculate the pseudo-gradient of the quantization parameters and update the activation frequency and gradient sensitivity ; Dynamic adjustment: Every step, recalculate the sensitivity quantiles and update the bit-width allocation; Use a sliding window to count the activation patterns and optimize the mask generation strategy.

9. The lightweight natural language processing large model training method according to claim 1, wherein: The calculation for evaluating the student model is as follows: , Among them, is the posterior probability, that is, the similarity between the output of the student model and the enhanced dataset. N is the total amount of training results for comparison and determination. represents the training result. represents the position sequence of the training result. is the likelihood probability of the similarity between the training result and the enhanced dataset. is the prior probability of the similarity between the training result and the enhanced dataset. is the likelihood probability of the non - similarity between the training result and the enhanced dataset. is the prior probability of the non - similarity between the training result and the enhanced dataset.

Citation Information

Patent Citations

  • Named entity recognition method and device, equipment and storage medium

    CN115713082A

  • Large model increment training method and system based on dynamic sparsification

    CN119669714A