Quantification and adaptive model deployment method for large model field of citrus intelligent planting management
By employing large-scale model domain quantization and adaptive model deployment methods, the problems of insufficient model deployment efficiency and poor domain adaptability in citrus planting management were solved, enabling efficient and accurate execution of key tasks in citrus planting and improving the inference accuracy and resource utilization of agricultural edge devices.
Patent Information
- Application Number
- CN202511296987.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-11
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-09-11
AI Technical Summary
Existing technologies in citrus planting and management suffer from problems such as insufficient model deployment efficiency, poor domain adaptability, and rigid multi-task support, resulting in limited computing power and tight memory resources for agricultural field equipment. Furthermore, it is difficult to balance accuracy and efficiency in citrus planting scenarios with general-purpose large-scale visual models.
We adopt a large-scale model domain quantization and adaptive model deployment method, and generate fully quantized, mixed-precision, and full-precision models through scenario quantization fine-tuning and task-aware hybrid deployment. This includes building a calibration dataset, quantizing the Transformer model, performing activation matrix rotation and smoothing factor calculation, fine-tuning instruction set construction, and task complexity classification.
It significantly improves the inference accuracy and stability of the model on agricultural edge devices, enhances its specific adaptability and task robustness to citrus data, and optimizes resource utilization and overall system efficiency.
Smart Images

Figure CN120764594B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of computer artificial intelligence, in particular to the field of large model quantization and adaptive model deployment for citrus intelligent planting management. BACKGROUND
[0002] With the in-depth application of artificial intelligence technology in the field of agriculture, agricultural large models based on multi-modal data fusion have gradually become a key technology for the development of smart agriculture. In traditional agricultural production, citrus planting management mainly relies on manual experience for judgment, which has problems such as low accuracy of disease and pest identification, untimely growth state monitoring, and low efficiency of fruit grading. In recent years, although multi-modal models based on the fusion of visual, semantic and environmental data have shown certain advantages in citrus disease and pest control, growth monitoring and other scenes, the existing technology still has the following outstanding defects: (1) Insufficient model deployment efficiency: agricultural field devices are generally characterized by limited computing power and tight memory resources, while existing large models usually use high-precision floating-point calculations (such as FP32), making it difficult to achieve real-time inference on edge computing devices. (2) Poor domain adaptability: when general visual large models are directly applied to citrus planting scenarios, due to the significant regional specificity of agricultural data (such as differences in disease and pest morphology, varying light conditions, etc.), existing quantization techniques (such as INT8 uniform quantization) can lead to a decrease in key feature extraction capability; (3) Multi-task support is rigid: current deployment schemes mostly use fixed quantization strategies, which cannot adapt to the differentiated task requirements in citrus planting. For example, fruit maturity detection requires high-precision feature retention, while large-area disease and pest screening focuses more on processing speed, and a single quantization mode cannot balance precision and efficiency.
[0003] Therefore, it is urgent to develop an agricultural large model optimization method combining domain quantization fine-tuning and task-aware adaptive deployment to break through the bottlenecks of existing technology in the application of citrus smart planting. SUMMARY
[0004] The purpose of the present application is to provide a large model domain quantization and adaptive model deployment method for citrus intelligent planting management, to achieve efficient and accurate execution of key tasks in citrus planting through scene quantization fine-tuning and task-aware hybrid deployment on edge devices.
[0005] To solve the above technical problems, the technical solution adopted by the present application is a large model domain quantization and adaptive model deployment method for citrus intelligent planting management, comprising the following steps:
[0006] S1: Construct a large model calibration dataset and perform preprocessing, and input the preprocessed large model calibration dataset into a Transformer model to perform forward inference;
[0007] S2: respectively left-rotating the weight matrix of the to-be-quantized layer in the Transformer model, right-rotating the activation matrix, performing GPTQ quantization on the rotated model, and calculating the dynamic adaptive smoothing factor of each channel of the activation matrix;
[0008] S3: grouping the right-rotated activation matrix according to the hidden dimension, performing normalization and symmetric quantization on each group, and generating a quantized Transformer model;
[0009] The quantized Transformer model is composed of multiple stacked Transformer encoders;
[0010] S4: obtaining citrus industry related text data, cleaning and preprocessing the data, and constructing a fine-tuning instruction set for the quantized large model based on the preprocessed data;
[0011] S5: using the S4 to construct the citrus quantized large model fine-tuning instruction and the preprocessed citrus industry related text to supervise the training of the quantized Transformer model, and generating a full-precision model, a full-precision model, and a mixed precision model;
[0012] S6: constructing and training a task complexity classifier, inputting a to-be-classified task into the trained classifier, and outputting a complexity classification result;
[0013] S7: according to the level of the task complexity classification result, determining the full-precision model, the mixed precision model, and the full-precision model as the target deployment model.
[0014] Further, the specific steps of S1 are:
[0015] S1.1: select a general text dataset, tokenize the general text dataset to generate input_ids and attention_mask fields, truncate and key information desensitization process the fields to obtain a calibration dataset;
[0016] S1.2: input the calibration dataset into the Transformer model and perform forward inference in the Transformer model to obtain the activation value of the to-be-quantized layer;
[0017] The Transformer model includes an attention module, a feedforward network module, and an output head module;
[0018] The to-be-quantized layer includes a linear projection layer, a fully connected layer in the feedforward module, a Query matrix, a Key matrix, a Value matrix in the attention module, and an activation function layer;
[0019] S1.3: Calculate the maximum, mean, and standard deviation of the activation values and perform statistical analysis;
[0020] S1.4: Calculate the quantization scaling factor based on the statistical results of the activation values.
[0021] Furthermore, the dynamic adaptive smoothing factor for each channel of the activation matrix in S2 is calculated as follows:
[0022] in: For sequence position, As weight, For the activation matrix, the first One channel, For channel weighting coefficients, For the first The initial smoothing factor of the channel. The maximum number of channels, The number of adjacent channels. For batch size, For sequence length, The number of adjacent channels. For weighted The Dynamic adaptive smoothing factor for each channel.
[0023] Furthermore, the specific steps of S3 are as follows:
[0024] S3.1: Group the activation matrix after right multiplication and rotation according to the hidden dimension to obtain multiple activation sub-matrix groups;
[0025] Grouping the activation matrix after right multiplication and rotation according to the hidden dimension satisfies the following condition:
[0026] Each activation submatrix group contains either continuous or non-continuous hidden dimension indices;
[0027] The group size is dynamically adjusted based on the number of parallel computing units in the hardware.
[0028] S3.2: Normalize each activation matrix. The calculation method is as follows:
[0029] in: For the activation matrix, the first The activation value is obtained by normalizing each channel. For the activation matrix, the first Each channel is assigned an activation value. For weighted The Dynamic adaptive smoothing factor for each channel;
[0030] S3.3: Symmetric quantization is performed on the normalized activation matrix group;
[0031] S3.4: Output recovery is performed on the symmetrically quantized activation matrix, and the calculation method is: .
[0032] Wherein: is the quantized activation value of the i-th channel, is the quantized weight matrix of the i-th channel, is the activation matrix after output recovery, is the dynamic adaptive smoothing factor of the i-th channel with weight .
[0033] Further, in S4, citrus industry related text data is obtained, cleaned and preprocessed, and the cleaning specifically includes: removing duplicate data, removing invalid characters, and removing redundant webpage tags; the preprocessing includes: converting the cleaned text data into question and answer pairs and knowledge graph triples.
[0034] Further, in S5, the S4 constructed citrus quantization large model fine-tuning instruction and the preprocessed citrus industry related text are used to supervise the training of the quantized Transformer model, specifically: the S5 citrus quantization large model fine-tuning instruction and the preprocessed citrus industry related text are used to supervise the training of the quantized Transformer model generated by S3 using LoRA, and the perplexity, BLEU, and ROUGE indicators are used to monitor the training process, and the full quantization model, full precision model, and mixed precision model are generated.
[0035] Further, the two-stage fine-tuning of the quantized Transformer model generated by S3 using LoRA is specifically as follows:
[0036] S5.1: Inject a trainable LoRA module into the quantized Transformer model generated by S3, and keep the backbone network parameters frozen;
[0037] S5.2: Add a classifier to the top layer of the quantized Transformer model;
[0038] S5.3: In the first stage, the top classifier of the quantized Transformer model and the LoRA module in the Transformer are trained, and the remaining parameters are kept frozen;
[0039] S5.4: In the second stage, all trainable LoRA module parameters in the Transformer model are unfrozen, the backbone network parameters are kept in the frozen state, and a hierarchical learning rate strategy is adopted to optimize the LoRA module.
[0040] The beneficial effects of the present application are: the present application effectively suppresses quantization error and eliminates abnormal fluctuations in activation values by combining low-bit quantization with a non-training activation value smoothing technique, significantly improving the inference accuracy and stability of low-bit models on agricultural edge devices while greatly reducing model memory occupancy and computational complexity; the present application effectively overcomes the precision loss problem of directly quantizing general large models in agricultural scenarios by fine-tuning the quantized model using agricultural citrus field data, significantly improving the specific adaptation ability and task robustness of the model to citrus field data distribution; the present application adopts a task-driven mixed precision deployment method, which can dynamically select the optimal quantization strategy according to the precision, speed and resource requirements of different agricultural sub-tasks, maximizing the overall efficiency and resource utilization of the system while ensuring the accuracy of key tasks. BRIEF DESCRIPTION OF DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0042] Figure 1 is a flowchart of the present application;
[0043] Figure 2 is a model quantization process flowchart of the present application;
[0044] Figure 3 is a data processing and model fine-tuning model flowchart in the present application;
[0045] Figure 4 is a self-adaptive task deployment flowchart of the present application. DETAILED DESCRIPTION
[0046] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0047] The citrus intelligent planting management-oriented large model field quantization and adaptive model deployment method will be further described in detail below with reference to the accompanying drawings.
[0048] As shown in Figure 1 The citrus intelligent planting management-oriented large model field quantization and adaptive model deployment method includes the following steps:
[0049] S1: Construct a large model calibration dataset and perform preprocessing, and input the preprocessed large model calibration dataset into a Transformer model to perform forward inference;
[0050] The specific steps of S1 are:
[0051] S1.1: Select a general text dataset, tokenizer encode the general text dataset to generate input_ids and attention_mask fields, and perform truncation and key information desensitization processing on the fields to obtain a calibration dataset;
[0052] S1.2: Input the calibration dataset into the Transformer model and perform forward inference in the Transformer model to obtain the activation value of the layer to be quantized;
[0053] The Transformer model includes an attention module, a feedforward network module, and an output head module.
[0054] The layer to be quantized includes a linear projection layer, a fully connected layer in the feedforward module, a Query matrix, a Key matrix, a Value matrix of the attention module, and an activation function layer.
[0055] S1.3: Calculate the maximum value, mean value, and standard deviation of the activation value and perform statistics;
[0056] S1.4: Calculate the quantization scaling factor based on the statistical results of the activation value.
[0057] In this embodiment, the S1 calibration dataset selects WikiText-2, and then the selected calibration dataset is formatted, the tokenizer is used to generate the standard input field containing input_ids and attention_mask, and the maximum context length (for example, 4096 tokens) of the model is truncated or padded, while the privacy data is desensitized, and the calibration dataset of 128-1024 samples is constructed; then the FP32 in the Transformer model is used for forward inference, the activation values of the linear projection layer, the feedforward network full connection layer, the attention module Q / K / V matrix, and the activation function layer are collected, and the maximum value, mean value and standard deviation statistical indicators are calculated; finally, based on the distribution characteristics of the activation values, the quantization scale parameters of each layer are calculated.
[0058] S2: respectively left-rotating the weight matrix of the layer to be quantized in the Transformer model, right-rotating the activation matrix, performing GPTQ (Gradient Post-Training Quantization) quantization on the rotated model, and calculating the dynamic adaptive smoothing factor of each channel of the activation matrix Figure 2
[0059] Step S2 proposes an offline rotation encoding method for the quantization problem of the self-attention module (Self-Attention) and the feed-forward neural network module (Feed-Forward Network) in the Transformer architecture. This method applies a specific rotation transformation to the linear weight matrix, effectively improving the numerical distribution characteristics of the weight and activation values without changing the original calculation semantics. This method significantly improves the accuracy performance of subsequent low-bit quantization (such as INT4).
[0060] The right-rotating technology principle of the activation matrix of the layer to be quantized in the Transformer model in step S2 is as follows:
[0061] For linear transformation , where is the weight matrix, is the activation matrix, and an orthogonal rotation matrix is introduced, where is the order of the weight matrix, satisfying , where is the identity matrix, without changing the final output, left-rotating the weight matrix , and its calculation method is as follows: , where is the transformed matrix, and at the same time, the input activation matrix Right multiplication transformation is performed, which is calculated as follows wherein, is the transformed activation matrix The following function relationship still holds:
[0062]
[0063] wherein, is the weight matrix, is the activation matrix, is the orthogonal rotation matrix.
[0064] In addition, the orthogonal rotation matrix is selected as an orthogonal matrix obtained by orthogonal triangular decomposition, The decomposition to generate the orthogonal matrix is a core operation in linear algebra, which can decompose any invertible real matrix into the product of an orthogonal matrix and an upper triangular matrix. The column vectors of the orthogonal matrix are not only orthogonal to each other but also have a unit length, and the transpose of the orthogonal matrix is its inverse matrix.
[0065] Right multiplication rotation is performed on the activation matrix of the layer to be quantized in the Transformer model, and GPTQ (Gradient Post-Training Quantization) quantization is performed on the rotated model. In specific implementation, the following strategies are adopted:
[0066] The weights of each linear layer are unilaterally rotated, i.e.,
[0067] For the of the attention module, and the in the feedforward module, right multiplication rotation is also performed, wherein are the Query weight matrix, the Key weight matrix, and the Value weight matrix, respectively, are the weight matrices for upsampling and downsampling of the feedforward neural network, respectively; correspondingly, the activation matrix is also right multiplied.
[0068] After rotation is completed, the new weight matrix is obtained, which is more suitable for low-bit quantization in numerical distribution. GPTQ is then used to compress and quantize the rotated model.
[0069] The calculation principle of the dynamic adaptive smoothing factor of each channel of the activation matrix in S2 is as follows: for the input activation matrix wherein, is the batch size, is the sequence length, To hide the dimension, the maximum value is calculated for each channel as a smoothing factor. To enhance robustness to noise and distribution shift, a dynamic adaptive smoothing factor is proposed, which introduces channel weights and historical statistics. The smoothing factor is dynamically adjusted according to the importance of the channel to avoid excessive compression of sparse channels by a single maximum value.
[0070] The dynamic adaptive smoothing factor for each channel of the activation matrix in S2 is calculated as follows:
[0071] in: For sequence position, As weight, For the activation matrix, the first One channel, For channel weighting coefficients, For the first The initial smoothing factor of the channel. The maximum number of channels, The number of adjacent channels. For batch size, For sequence length, The number of adjacent channels. For weighted The Dynamic adaptive smoothing factor for each channel.
[0072] S3: Group the activation matrix after right multiplication and rotation according to the hidden dimension, perform normalization and symmetric quantization on each group, and generate the quantized Transformer model;
[0073] The specific steps for S3 are as follows:
[0074] S3.1: Group the activation matrix after right multiplication and rotation according to the hidden dimension to obtain multiple activation sub-matrix groups;
[0075] Grouping the activation matrix after right multiplication and rotation according to the hidden dimension satisfies the following condition:
[0076] Each activation submatrix group contains either continuous or non-continuous hidden dimension indices;
[0077] The group size is dynamically adjusted based on the number of parallel computing units in the hardware.
[0078] S3.2: Normalize each activation matrix. The calculation method is as follows:
[0079] in: For the activation matrix, the first The activation value is obtained by normalizing each channel. For the activation matrix, the first Each channel is assigned an activation value. For weighted The Dynamic adaptive smoothing factor for each channel;
[0080] S3.3: Perform symmetric quantization on the normalized activation matrix.
[0081] S3.4: Output recovery is performed on the activation matrix after symmetric quantization. The calculation method is as follows:
[0082] in: For the quantified first The activation value of the channel, For the quantified first The channel weight matrix, This is the activation matrix after output recovery. For weighted The Dynamic adaptive smoothing factor for each channel.
[0083] The quantized Transformer model consists of multiple stacked Transformer encoders;
[0084] S4: Acquire textual data related to the citrus industry, clean and preprocess it, and construct a fine-tuning instruction set for a large-scale quantitative model based on the preprocessed data. Figure 3 );
[0085] The cleaning process specifically includes: removing duplicate data, removing invalid characters, and removing redundant web page tags; preprocessing includes: converting the cleaned text data into question-answer pairs and knowledge graph triples.
[0086] In this implementation method S4, the specific sources of textual data related to the citrus industry are: citrus planting guidelines published by agricultural technology extension stations, expert consultation and Q&A records, agricultural research papers, government-related agricultural data, agricultural forum exchange content, and farmers' actual problems and experience summaries published on social media, while also integrating symptom descriptions and prevention and control plans from agricultural technology manuals.
[0087] In this embodiment, the specific construction method of the fine-tuning instruction set for building a large quantization model based on preprocessed data is as follows: First, the original text needs to be encoded into input_ids and attention_mask using the official or open-source tokenizer. Based on the context length limit of 2048 tokens, the text is windowed and packaged to organize all samples into JSON or dataset objects. Then, the supervised fine-tuning (SFT) method is used to construct question-and-answer fine-tuning instructions.
[0088] For the llama model, such as: "instruction":"Please briefly introduce the transmission mode and control method of citrus Huanglongbing disease."
[0089] "input":"",
[0090] "output":"Citrus Huanglongbing is transmitted by citrus psylla and is a bacterial disease. There is no specific drug at present, and it is recommended to remove diseased plants and control pests to control transmission."
[0091] S5: using the S4 constructed citrus quantification large model fine-tuning instruction and the preprocessed citrus industry related text to supervise the training of the quantified Transformer model, to generate full quantization model, full precision model and mixed precision model;
[0092] S5 is: using the citrus quantification large model fine-tuning instruction in S4 and the preprocessed citrus industry related text, and using LoRA to supervise the training of the quantified Transformer model generated by S3, and using perplexity, BLEU and ROUGE indicators to monitor the training process, to generate full quantization model, full precision model and mixed precision model;
[0093] The two-stage fine-tuning of the quantified Transformer model generated by S3 using LoRA is specifically as follows:
[0094] S5.1: inject a trainable LoRA module into the quantified Transformer model generated by S3, and keep the backbone network parameters frozen;
[0095] S5.2: add a classifier to the top layer of the quantified Transformer model;
[0096] S5.3: In the first stage, the top classifier of the quantified Transformer model and the LoRA module in the Transformer are trained, and the remaining parameters are kept frozen;
[0097] S5.4: In the second stage, all trainable LoRA module parameters in the Transformer model are unfrozen, the backbone network parameters are kept in a frozen state, and a hierarchical learning rate strategy is used to optimize the LoRA module.
[0098] In this application, the LoRA (Low-Rank Adaptation) high-efficiency parameter fine-tuning technology is used, only a low-rank update matrix is added to a small number of trainable layers (such as query / key projection in Attention), to reduce the computing power and memory occupation. The loss function used in training is usually token-level cross-entropy loss.
[0099] To maximize the response accuracy of the model to the citrus instruction problem, during the fine-tuning of the large language model of the application, the model input sequence length is usually set to 1024, which can be flexibly adjusted according to the demand of the downstream task for context modeling capability; according to the tokenizer used by the model, the vocabulary size is set to 32768 to ensure coverage of agricultural terms. In order to avoid the influence of special tokens (such as padding) on loss calculation, the training uses the ignore_index=-100 parameter to shield invalid positions, ensuring effective gradient update. To solve the problem of model overfitting, a label smoothing mechanism can be selectively introduced, set between 0.0 and 0.1, to enhance the model's generalization ability to ambiguous tokens. In the batch loss aggregation mode, the "mean" strategy is used to average the loss values of all valid tokens, improving the stability of loss calculation under different sample lengths. At the same time, the training enables the teacher forcing strategy, that is, the model predicts at each step with the real token as input, which accelerates the convergence speed and improves the semantic coherence of the output. The above control parameters jointly constitute the loss calculation framework in the training stage, ensuring the stability of the model in the agricultural citrus field task, and optimizing the generation accuracy and language understanding ability,
[0100] S6: Construct and train a task complexity classifier, input the task to be classified into the trained classifier, and output the complexity classification result Figure 4 );
[0101] S7: According to the level of the task complexity classification result, determine the full-precision model, mixed-precision model, and full-precision model as the target deployment model.
[0102] The application point of the application: by fusing low-bit quantization, domain fine-tuning and adaptive deployment mechanism, the problem of large language model in agricultural field application is effectively solved. Large consumption of computing resources and low inference efficiency.
[0103] Each embodiment in the specification is described in a related manner, and the same and similar parts between each embodiment can be referred to each other. Each embodiment focuses on the difference from other embodiments. Especially, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and the related parts can be referred to the part of the method embodiment.
[0104] The above is only a preferred embodiment of the application, and is not used to limit the protection scope of the application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the application is included in the protection scope of the application.
Claims
1. A method for model field quantization and adaptive model deployment for large models in citrus intelligent plantation management, characterized in that, The method comprises the following steps: S1: Constructing a large model calibration dataset and preprocessing, inputting the preprocessed large model calibration dataset into a Transformer model to perform forward inference; S2: respectively multiplying the weight matrix of the to-be-quantized layer in the Transformer model on the left, multiplying the activation matrix on the right, performing GPTQ quantization on the rotated model, and calculating the dynamic adaptive smoothing factor of each channel of the activation matrix; S3: grouping the right-rotated activation matrix by hidden dimension, performing normalization and symmetric quantization on each group to generate a quantized Transformer model; The quantized Transformer model is composed of multiple stacked Transformer encoders; S4: obtaining citrus industry related text data, cleaning and preprocessing, and constructing a fine-tuning instruction set for the quantized large model based on the preprocessed data; S5: using S4 to construct a citrus quantized large model fine-tuning instruction and preprocessed citrus industry related text to supervise the training of the quantized Transformer model, generating a full-quantization model and a full-precision full-quantization model, a full-precision model, and a mixed-precision model; S6: constructing and training a task complexity classifier, inputting the to-be-classified task into the trained classifier to output a complexity classification result; S7: according to the level of the task complexity classification result, determining the full-quantization model, the mixed-precision model, and the full-precision model as the target deployment model; The specific steps of using S4 to construct a citrus quantized large model fine-tuning instruction and preprocessed citrus industry related text to supervise the training of the quantized Transformer model in S5 are as follows: Using the citrus quantized large model fine-tuning instruction and preprocessed citrus industry related text in S5 and adopting LoRA to supervise the training of the quantized Transformer model generated in S3, and using perplexity, BLEU, and ROUGE indicators to monitor the training process, generating a full-quantization model, a full-precision model, and a mixed-precision model; The specific steps of using LoRA to perform two-stage fine-tuning on the quantized Transformer model generated in S3 are as follows: S5.1: injecting a trainable LoRA module into the quantized Transformer model generated in S3, and keeping the backbone network parameters frozen; S5.2: adding a classifier to the top layer of the quantized Transformer model; S5.3: In the first stage, the top classifier of the quantized Transformer model and the LoRA module in the Transformer are trained, and the remaining parameters are kept frozen; S5.4: In the second stage, unfreeze all trainable LoRA module parameters in the Transformer model, keep the backbone network parameters in a frozen state, and use a hierarchical learning rate strategy to optimize the LoRA module.
2. The method of claim 1, wherein the method is characterized by, The specific steps of S1 are as follows: S1.1: Select a general text dataset, tokenize the general text dataset, generate input_ids and attention_mask fields, truncate and key information desensitization processing, and obtain a calibration dataset; S1.2: Input the calibration dataset into the Transformer model and perform forward inference in the Transformer model to obtain the activation value of the layer to be quantized; The Transformer model includes an attention module, a feedforward network module, and an output head module; The layer to be quantized includes a linear projection layer, a fully connected layer in the feedforward module, a Query matrix, a Key matrix, and a Value matrix of the attention module, and an activation function layer; S1.3: Calculate the maximum value, mean value, and standard deviation of the activation value and perform statistics; S1.4: Based on the statistical results of the activation value, calculate the quantization scaling factor. 3.The method of claim 1, wherein the method is characterized by, The specific steps of S3 are: S3.1: Group the right-multiplication rotated activation matrix by hidden dimension to obtain a plurality of activation sub-matrix groups; Grouping the right-multiplication rotated activation matrix by hidden dimension satisfies the following conditions: Each activation sub-matrix group contains continuous or non-continuous hidden dimension indexes; The group size is dynamically adjusted according to the number of hardware parallel computing units; S3.2: Normalize each activation matrix, the calculation method is: Wherein: is the activation value of the normalized channel in the activation matrix, is the activation value of the normalized channel in the activation matrix, is the activation value of the normalized channel in the activation matrix, is the activation value of the normalized channel in the activation matrix, is the dynamic adaptive smoothing factor of the channel with weight , and is the channel number. S3.3: Symmetric quantization of normalized activation matrix group; S3.4: Output recovery is performed on the activation matrix after symmetric quantization. The calculation method is as follows: in: For the quantified first The activation value of the channel, For the quantified first The channel weight matrix, This is the activation matrix after output recovery. For weighted The Dynamic adaptive smoothing factor for each channel.
4. The method of claim 1, wherein the method is characterized by, In S4, the text data related to the citrus industry is obtained, cleaned and preprocessed, and the cleaning specifically includes removing duplicate data, removing invalid characters, and removing redundant webpage tags; The pre-processing includes converting the cleaned text data into question and answer pairs and knowledge graph triples.
Citation Information
Patent Citations
System and method for compressing and reconstructing audio files
CA2569536A1
Efficient quantization method for Vision Transform neural network
CN119830971A