Industrial scene intelligent question answering system based on large language model and use method

By performing multi-task, low-order adaptive fine-tuning and quantization of sensitivity of low-sample grouping for large language models, combined with the industrial knowledge base optimization search and generation module, the efficiency and accuracy of the Q&A system in industrial scenarios are solved, and efficient and professional fault diagnosis and maintenance guidance are achieved.

CN120277245APending Publication Date: 2025-07-08SHANGHAI UNIV OF ENG SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510350029.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing question-and-answer system based on large language models has problems such as low model fine-tuning efficiency, high resource consumption, lack of logic and professionalism in generated answers, and insufficient coverage of search results in industrial scenarios, which is difficult to meet the complexity and high requirements of industrial scenarios.

Method used

The low-order adaptive method for multi-task, few-sample grouping sensitivity is used to fine-tune and quantify the large language model, combine industrial data to build a knowledge base, optimize the search and generation module, and deploy it on local servers or edge devices. Through dynamic task weight decomposition, few-sample weight initialization and regularization constraints, group sensitive quantization is realized, and the adaptability and efficiency of the model are enhanced.

Benefits of technology

It realizes efficient and accurate fault diagnosis and maintenance guidance in industrial scenarios, reduces hardware resource requirements, meets privacy protection and real-time response needs, and improves diagnostic accuracy and professionalism of maintenance guidance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277245A_ABST
    Figure CN120277245A_ABST
Patent Text Reader

Abstract

The invention relates to an industrial scene intelligent question-answering system based on a large language model and a use method, and belongs to the technical field of intelligent question-answering systems, and the method comprises the following steps: 1, constructing an industrial scene knowledge base by using industrial data; 2, selecting a large language model as a base; 3, performing fine adjustment and quantification on the base large language model by using a multi-task few-sample grouping sensitivity low-order adaptive method; 4, optimizing a retrieval module and a generation module of the base large language model; and 5, combining the optimized large language model with the industrial scene knowledge base to form an industrial scene intelligent question-answering system, and deploying the industrial scene intelligent question-answering system on a local server or edge equipment. According to the method, hardware resource requirements and delay can be effectively reduced, diagnosis accuracy and professional maintenance guidance are improved, privacy protection, real-time response and diversified task requirements in an industrial scene are met, and accurate response to complex tasks in the industrial scene is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent question - answering systems, and particularly relates to an intelligent question - answering system for industrial scenarios based on large language models and a usage method thereof. Background Art

[0002] In traditional industrial scenarios, equipment maintenance methods usually rely on manual experience and manual queries, with low efficiency, poor accuracy, and being easily affected by the subjective factors of operators. In order to improve maintenance efficiency and accuracy, the use of large language models (such as GPT, BERT, etc.) for natural language processing has been widely applied to tasks such as text understanding, information retrieval, and question - answering generation. Therefore, intelligent question - answering systems based on artificial intelligence have gradually become one of the solutions.

[0003] However, in industrial scenarios, due to characteristics such as high task complexity, scarce data volume, and strong real - time requirements, existing question - answering systems based on large language models still face many technical challenges in aspects such as model fine - tuning efficiency, quantization optimization, and retrieval - generation logic. First of all, most existing systems are mainly optimized for general scenarios and fail to effectively adapt to the special needs of the industrial field. For example, in equipment fault diagnosis and maintenance tasks, existing technologies fail to perform structured classification and highly relevant processing on the knowledge base in combination with specific industrial rules (such as fault codes, maintenance records, etc.), resulting in insufficient coverage of retrieval results and being unable to accurately solve complex industrial problems. Secondly, the generation module of existing technologies lacks logic and professionalism. The generation module usually relies on general - purpose language models and has limited support for specific task logics in industrial scenarios (such as multi - step fault diagnosis and maintenance operations), and the generated answers often lack coherence and professionalism. For example, in multi - step operation scenarios, existing technologies are difficult to provide clear and detailed solutions. In addition, the lack of templated generation and multi - step reasoning strategies results in answers that are not logically rigorous enough to meet the accuracy requirements of the industrial field.

[0004] At the same time, in terms of model fine - tuning, traditional large language models use the full - parameter fine - tuning method, which requires huge computing resources and storage space for training and fine - tuning, with low efficiency. Especially for devices with limited computing power, it is neither efficient nor practical and is difficult to adapt to the multi - task and dynamic requirements in industrial scenarios; in terms of quantization technology, existing methods mostly adopt a unified quantization strategy, treating all weights with the same precision, failing to distinguish according to the sensitivity of weights to tasks, often resulting in insufficient precision for key task weights and causing a decline in model performance, especially being particularly obvious in industrial scenarios with high - precision requirements; and most existing technologies rely on cloud - based deployment solutions, and in industrial scenarios, the requirements for privacy protection and real - time response are difficult to be met.

[0005] Therefore, the existing technologies have obvious deficiencies in multi-task adaptation, model quantization, retrieval generation, and resource and equipment consumption costs, restricting their in-depth application in intelligent question-and-answer systems in industrial scenarios and making it difficult to meet the complexity and high requirements of specific industry needs. Summary of the Invention

[0006] The object of the present invention is to address the problems existing in the prior art and provide an intelligent question-and-answer system for industrial scenarios based on large language models and its usage method, which can effectively reduce the hardware resource requirements and latency, improve the diagnostic accuracy rate and the professionalism of maintenance guidance, and meet the privacy protection, real-time response, and diverse task requirements in industrial scenarios, achieving accurate responses to complex tasks in industrial scenarios.

[0007] The technical solutions are as follows:

[0008] An intelligent question-and-answer system for industrial scenarios based on large language models includes the following steps:

[0009] Step 1: Construct an industrial scenario knowledge base using industrial data, where the sources of industrial data include equipment manuals, fault logs, maintenance records, and industry standards and specifications;

[0010] Step 2: Select a large language model as the base, and the large language model is one of LLaMA3-8B-GSLoRA, ChatGLM2-6B, and LLaMA3;

[0011] Step 3: Use the Multi-Task Few-Shot Grouped Sensitivity Low-Rank Adaptation (MF-GSLoRA) method to fine-tune and quantize the base large language model, group the model weights or embedding vectors according to the sensitivity to performance and tasks as grouped sensitive vector data, and determine the processing methods of weights at different quantization levels;

[0012] Step 4: Optimize the retrieval module and generation module of the base large language model;

[0013] Step 5: Combine the optimized large language model and the industrial scenario knowledge base to form an intelligent question-and-answer system for industrial scenarios and deploy it on a local server or edge device.

[0014] In Step 3, when using the Multi-Task Few-Shot Grouped Sensitivity Low-Rank Adaptation method to fine-tune and quantize the large language model and using it as grouped sensitive vector data to determine the processing methods of weights at different quantization levels, it includes the following steps:

[0015] Step 31: Introduce a dynamic task weight decomposition strategy. Based on the sharing of the basic weight matrix, generate task-specific low-rank matrices for different tasks, and dynamically adjust the low-rank matrix A during the multi-task fine-tuning process through a task scheduler i and B i 's parameter distribution, which is dynamically updated by the task scheduler in combination with the insertion points of the low-rank adapters (LoRA), ensuring multi-task weight sharing while retaining task specificity: W′ = W + α(A shared B shared + A task B task ), where A shared , B shared represent the shared task weights, and A task B task represent the task-specific weights;

[0016] Step 32: Enhance the fine-tuning effect in the few-shot scenario through few-shot weight initialization and regularization constraints. Among them, few-shot weight initialization combines domain knowledge or pre-trained weights to initialize the low-rank matrices A and B. Regularization constraints are to add regularization constraints to the update process of the low-rank matrices to prevent excessive deviation of the model weights during few-shot fine-tuning. And before fine-tuning, generate domain-related embedding representations using the pre-trained model as the initial values of the low-rank matrices to reduce weight drift in the few-shot scenario, and then use regularization to limit the change range of the weights to ensure fine-tuning stability, and insert additional regularization terms in the low-rank adaptive low-rank adapter: Loss reg = λ(||A|| 2 + ||B|| 2 );

[0017] Step 33: Conduct sensitivity analysis on the matrices A and B inserted by the low-rank adapter. For each weight or vector, calculate its gradient contribution to a specific task where L(W) is the task loss function. A large gradient contribution means high sensitivity, indicating a significant impact on the model performance. The formula for measuring the sensitivity of the weights is: where S(W) represents the sensitivity of the weight ω, and according to the sensitivity analysis results of the weights, the weights are divided into two groups, and a quantization level q(i) is defined for each group to control the precision of the grouped weights: where q(i) represents the quantization level of the i-th group; S(ω i) represents the sensitivity of the i-th weight; τ represents the grouping sensitivity threshold. For the high-sensitivity weight group, 8-bit Normal-Float (NF8) quantization is adopted to maintain high precision and reduce the impact of quantization on performance. For the low-sensitivity weight group, 4-bit Normal-Float (NF4) quantization is adopted to further compress the weights, reduce the computational cost and storage requirements, combine the quantized weights with the low-rank matrix, optimize the inner product calculation of the matrix during inference, and perform the product calculation of the low-rank matrix in each group of weights; combining the grouped weight tags and quantization levels, the weight calculation is finally integrated into the following formula: where W′ represents the final weight of the model, W represents the weight matrix, α represents the low-order adapter scaling factor to control the impact of the inserted weights; N represents the number of weight groups, q(i) represents the quantization level (4-bit or 8-bit) of the i-th group of weights, A i B i represents the core part of the low-rank matrix for weight adjustment;

[0018] Step 34: Insert task-related low-rank matrices in the attention layer or feed-forward layer of the base large language model for multi-task scenarios, and insert additional context processing weights between layers to enhance the cross-task context modeling ability.

[0019] The optimization of the retrieval module and generation module of the base large language model in Step 4 includes the following steps:

[0020] Step 41: Convert the user query and document content into high-dimensional vectors through the fine-tuned and vectorized large language model in Step 3, use the vector database for similarity calculation, screen relevant content, and calculate the similarity between the user query q and d based on the dot product cosine similarity for vector retrieval: where q is the user query vector, d is the document vector, and n is the dimension of the vector; then through similarity sorting, select the top k relevant documents: R = TopK(Similarity(q, d1), Similarity(q, d2),..., Similarity(q, d N )); To improve the efficiency and accuracy of retrieval, a two-stage retrieval mechanism of rough retrieval and fine retrieval is adopted. In the rough retrieval stage, perform preliminary similarity calculation on the vectorized user question q and the document vectors d1, d2.d3…d n in the knowledge base, screen out a batch of highly relevant documents, and obtain the preliminary relevant document set R coarse , and use the similarity threshold τ to screen the preliminary results: R coarse = {d|Similarity(q, d) > τ}, and then filter the results based on the rule R(d), and perform secondary screening in the rough retrieval results by combining domain rules to obtain the optimized high-quality document set R refined={d|d∈R coarse , where R(d) = 1}, and finally output the refined document set: R final = R refined , then concatenate the user question q with the refined highly relevant document set R refined to form the context input c and pass it to the generation module;

[0021] Step 42: Input the context input c into the specific template rule T to generate the answer A: A = T(C), where

[0022] the generation module generates the answer step by step based on the retrieved content C = {R1, R2,...R n}: A t = f(A t-1 , C), where A t represents the answer generated at the t-th step; A t-1 represents the answer generated in the previous round; f represents the generation function, and the final answer is the set of answers for all steps:

[0023] Step 43: The retrieval module first concatenates the relevant documents into the context input: where Q is the user query and R i is the retrieval result; then the generation module generates a structured answer based on the context input and achieves efficient collaboration through context fusion.

[0024] A method for using an industrial scenario intelligent question answering system based on a large language model. For the above industrial scenario intelligent question answering system, the following steps are executed:

[0025] Step 1: The user inputs a question to the industrial scenario intelligent question answering system;

[0026] Step 2: The industrial scenario intelligent question answering system tokenizes the question and performs semantic analysis;

[0027] Step 3: The industrial scenario intelligent question answering system performs vector retrieval on the question;

[0028] Step 4: The industrial scenario intelligent question answering system retrieves the documents related to the question and performs screening and sorting;

[0029] Step 5: The industrial scenario intelligent question answering system calls the optimized large language model for input inference;

[0030] Step 6: The industrial scenario intelligent question answering system optimizes the answer;

[0031] Step 7: The industrial scenario intelligent question answering system displays the optimized returned answer to the user.

[0032] Beneficial effects:

[0033] 1) The present invention adopts grouped sensitive quantization, multi-task dynamic weight decomposition, and few-shot weight initialization to construct an intelligent industrial defect and maintenance Q&A system based on the MF-GSLoRA fine-tuning quantization optimization method and retrieval-generation collaborative optimization, realizing the intelligent and efficient application of industrial equipment fault diagnosis and maintenance guidance.

[0034] 2) It can effectively reduce the hardware resource requirements and latency, improve the diagnostic accuracy and the professionalism of maintenance guidance, and meet the privacy protection, real-time response, and diverse task requirements in industrial scenarios, achieving precise response to complex tasks in industrial scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 is the process framework diagram of an intelligent Q&A system for industrial scenarios based on a large language model of the present invention;

[0036] Figure 2 is the schematic flow diagram of fine-tuning and quantization of the base large language model by the multi-task few-shot grouped sensitivity low-order adaptive method;

[0037] Figure 3 is the schematic flow diagram of the retrieval-generation collaborative optimization method;

[0038] Figure 4 is the flow chart of the usage method of an intelligent Q&A system for industrial scenarios based on a large language model;

[0039] Figure 5 is the change graph of accuracy and recall rate under different model versions in the embodiment;

[0040] Figure 6 is the change process of the loss smoothing curve before and after fine-tuning in the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] In order to make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments:

[0042] As Figure 1 shown, an intelligent Q&A system for industrial scenarios based on a large language model includes the following steps:

[0043] Step 1: Construct an industrial scenario knowledge base using industrial data;

[0044] Step 2: Select a large language model as the base;

[0045] Step 3: Use the multi-task few-shot grouped sensitivity low-order adaptive method to fine-tune and quantize the base large language model, group the model weights or embedding vectors according to their sensitivity to performance and tasks, and adopt different processing methods for the weights at different quantization levels;

[0046] Step 4: Optimize the retrieval module and generation module of the base large language model;

[0047] Step 5: Combine the optimized large language model and the industrial scenario knowledge base to form an industrial scenario intelligent question-answering system, and deploy it on a local server or edge device.

[0048] Further, the industrial data sources in Step 1 include device manuals, fault logs, maintenance records, and industry standards and specifications.

[0049] Further, the large language model in Step 2 is one of LLaMA3-8B-GSLoRA, ChatGLM2-6B, and LLaMA3.

[0050] Further, the multi-task few-shot grouped sensitivity low-order adaptive method for fine-tuning and quantizing the large language model in Step 3 includes the following steps:

[0051] Step 31: Introduce a dynamic task weight decomposition strategy, generate task-specific low-rank matrices for different tasks based on the sharing of the base weight matrix, and dynamically adjust the low-rank matrices A i and B i 's parameter distribution, which is dynamically updated by the task scheduler in combination with the low-order adapter (LoRA) insertion point to ensure multi-task weight sharing while retaining task specificity: W′ = W + α(A shared B shared +A task B task ), where A shared , B shared represents the shared task weights, and A task B task represents the task-specific weights;

[0052] Step 32: Enhance the fine-tuning effect in few-shot scenarios through few-shot weight initialization and regularization constraints. Among them, few-shot weight initialization combines domain knowledge or pre-trained weights to initialize the low-rank matrices A and B. Regularization constraint adds regularization constraints to the update process of the low-rank matrices to prevent excessive deviation of model weights during few-shot fine-tuning. And before fine-tuning, generate domain-related embedding representations using a pre-trained model as the initial values of the low-rank matrices to reduce weight drift in few-shot scenarios. Then, limit the change amplitude of weights through regularization to ensure fine-tuning stability, and insert an additional regularization term in the low-order adaptive low-order adapter: Loss reg = λ(||A|| 2 + ||B|| 2 )

[0053] Step 33: Conduct sensitivity analysis on the inserted matrices A and B in the low-order adapter. For each weight or vector, calculate its gradient contribution to a specific task where L(W) is the task loss function. A large gradient contribution indicates high sensitivity, suggesting a significant impact on model performance. The formula for measuring the sensitivity of weights is: where S(W) represents the sensitivity of weight ω, and according to the sensitivity analysis results of weights, divide the weights into two groups. Define the quantization level q(i) for each group to control the precision of the grouped weights: where q(i) represents the quantization level of the i-th group; S(ω i ) represents the sensitivity of the i-th weight; τ represents the grouped sensitivity threshold. For the high-sensitivity weight group, use 8-bit Normal-Float (NF8) quantization to maintain high precision and reduce the impact of quantization on performance. For the low-sensitivity weight group, use 4-bit Normal-Float (NF4) quantization to further compress the weights, reduce computational costs and storage requirements. Combine the quantized weights with the low-rank matrices to optimize the inner product calculation of the matrices during inference. In each group of weights, perform the product calculation of the low-rank matrices; combine the grouped weight labels and quantization levels, and finally integrate the weight calculation into the following formula: where W′ represents the final weight of the model, W represents the weight matrix, α represents the low-order adapter scaling factor to control the impact of the inserted weights; N represents the number of weight groups, q(i) represents the quantization level (4-bit or 8-bit) of the i-th group of weights, A i B i represents the core part of the weight adjustment by the low-rank matrices

[0054] Step 34: Insert task-related low-rank matrices in the attention layer or feed-forward layer of the base large language model for multi-task scenarios, and insert additional context processing weights between layers to enhance cross-task context modeling capabilities

[0055] Furthermore, the optimization of the retrieval module and the generation module of the base large language model in step 4 includes the following steps:

[0056] Step 41: Convert the user query and the document content into high-dimensional vectors through the fine-tuned and vectorized large language model in step 3, calculate the similarity using a vector database, and filter relevant content. The vector retrieval calculates the similarity between the user query q and d based on the dot product cosine similarity: where q is the user query vector, d is the document vector, and n is the dimension of the vector; then, through similarity ranking, select the top k relevant documents: R = TopK(Similarity(q, d1), Similarity(q, d2),..., Similarity(q, d N )); To improve the efficiency and accuracy of retrieval, a two-stage retrieval mechanism of rough retrieval and fine retrieval is adopted. In the rough retrieval stage, the vectorized user question q is initially calculated for similarity with the document vectors d1, d2, d3,... d n in the knowledge base, and a batch of highly relevant documents are filtered out to obtain the initial relevant document set R coarse , and the similarity threshold τ is used to screen the initial results: R coarse = {d|Similarity(q, d) > τ}, and then the results are filtered based on the rule R(d). In the rough retrieval results, secondary screening is performed in combination with domain rules to obtain the optimized high-quality document set R refined = {d|d ∈ R coarse , R(d) = 1}, and finally the document set after fine retrieval is output: R final = R refined , and then the user question q is concatenated with the highly relevant document set R refined after fine retrieval to form the context input C, which is passed to the generation module;

[0057] Step 42: Input the context input C into a specific template rule T to generate the answer A: A = T(C), and the

[0058] generation module generates the answer step by step based on the retrieved content C = {R1, R2,... R n}: A t = f(A t-1 , C), where A t represents the answer generated at the t-th step; A t-1 represents the answer generated in the previous round; f represents the generation function, and the final answer is the set of answers for all steps:

[0059] Step 43: First, the retrieval module concatenates the relevant documents into the context input: Where Q is the user query, and R i is the retrieval result; then the generation module generates a structured answer based on the context input through the task module, and realizes efficient collaboration through context fusion.

[0060] A method for using an industrial scenario intelligent question answering system based on a large language model. For the above industrial scenario intelligent question answering system, the following steps are executed:

[0061] Step 1: The user inputs a question to the industrial scenario intelligent question answering system;

[0062] Step 2: The industrial scenario intelligent question answering system tokenizes the question and performs semantic analysis;

[0063] Step 3: The industrial scenario intelligent question answering system performs vector retrieval on the question;

[0064] Step 4: The industrial scenario intelligent question answering system retrieves documents related to the question and performs screening and sorting;

[0065] Step 5: The industrial scenario intelligent question answering system calls the optimized large language model for input inference;

[0066] Step 6: The industrial scenario intelligent question answering system optimizes the answer;

[0067] Step 7: The industrial scenario intelligent question answering system displays the optimized returned answer to the user.

[0068] Example: Obtain industrial equipment-related text data from multiple data sources (such as equipment operation manuals, troubleshooting guides, maintenance logs, industry standards and specifications, etc.), and extract the core content related to industrial defect analysis and equipment maintenance to form a structured text data set. Use this data set to perform MF-GSLoRA fine-tuning and model merging quantization on the LLaMA3 model.

[0069] Convert the processed data into structured information and store it in an industrial equipment maintenance knowledge base in a unified format. The knowledge base content covers various technical reference information such as equipment function descriptions, fault diagnosis, maintenance steps, and industry standards. At the same time, use the fine-tuned large language model to perform semantic vectorization processing on each piece of text in the knowledge base, encode the text content into a high-dimensional vector representation, and retain its core semantic features. The vector data is stored in an efficient vector database and its index management is performed. When the user asks a question, the system matches the most relevant text data through vector retrieval technology to ensure retrieval accuracy.

[0070] In terms of fine-tuning, a multi-task fine-tuning technique is adopted to design a dynamic task weight decomposition strategy to generate a low-rank matrix for a specific task. At the same time, in combination with the few-shot fine-tuning technique, the low-rank matrix is initialized with domain knowledge and regularization constraints are added to prevent the model from overfitting. In terms of quantization, grouped sensitive vector data is introduced to perform sensitivity analysis on the model weights, which are divided into two groups: high-sensitive weights and low-sensitive weights. 8-bit and 4-bit Normal-Float quantization methods are used respectively to balance performance and resource consumption. The optimized weights are integrated through formulas to further improve the model inference efficiency and task adaptability.

[0071] Optimized based on RAG (Retrieval-Augmented Generation), in the retrieval module, by combining vectorization technology and domain rules, efficient similarity calculation and screening are performed on user queries to quickly locate the most relevant document content. In the generation module, through multi-step task reasoning, it is ensured that the generated answers are logical and operable.

[0072] As shown in Table 1, using the Ollama local deployment framework, the optimized LLaMA3-8B-GSLoRA model is deployed on a local server. The comparison of answers before and after model fine-tuning and quantization is shown in Table 2. The accuracy and recall rates of the original model, fine-tuned model, and quantized model are evaluated respectively as Figure 5 shown. It can be clearly seen the evolution trends of different versions of the model in the two core indicators of accuracy and recall. The performance of the fine-tuned model is significantly improved, and the quantized model is further optimized on this basis. And the change process of the loss smooth curve before and after fine-tuning is as Figure 6 shown, which proves that the model training remains stable. This result fully shows that through the combined optimization of fine-tuning and quantization, the locally deployed model after fine-tuning and quantization can more accurately respond to complex industrial Q&A tasks, and can also balance resource efficiency and actual application requirements.

[0073] Table 1 Experimental environment configuration information

[0074]

[0075]

[0076] Table 2 Comparison of answers before and after fine-tuning and quantization

[0077]

[0078]

[0079] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, and improvements made within the principle and spirit of the present invention shall be included in the protection scope of the present invention.

Claims

1. An intelligent Q&A system for industrial scenarios based on large language models, characterized in that: It includes the following steps: Step 1: Use industrial data to build an industrial scenario knowledge base; Step 2: Select a large language model as the base; Step 3: Use the multi-task few-shot grouped sensitivity low-order adaptive method to fine-tune and quantize the base large language model. Group the model weights or embedding vectors according to their sensitivity to performance and tasks as grouped sensitive vector data to determine the processing method of weights at different quantization levels; Step 4: Optimize the retrieval module and generation module of the base large language model; Step 5: Combine the optimized large language model and the industrial scenario knowledge base to form an industrial scenario intelligent question-answering system and deploy it on a local server or edge device.

2. The intelligent Q&A system for industrial scenarios based on large language models according to claim 1, wherein: In the above-mentioned Step 1, the sources of industrial data include equipment manuals, fault logs, maintenance records, and industry standards and specifications.

3. The intelligent Q&A system for industrial scenarios based on large language models according to claim 1, characterized in that: In the above-mentioned Step 2, the large language model is one of LLaMA3-8B-GSLoRA, ChatGLM2-6B, and LLaMA3.

4. The intelligent Q&A system for industrial scenarios based on large language models according to claim 1, characterized in that: In the above-mentioned Step 3, the multi-task few-shot grouped sensitivity low-order adaptive method is used to fine-tune and quantize the large language model, including the following steps: Step 31: Introduce a dynamic task weight decomposition strategy. Based on the sharing of the basic weight matrix, generate task-specific low-rank matrices for different tasks, and dynamically adjust the low-rank matrices A i and B i 's parameter distributions, which are dynamically updated by the task scheduler in combination with the low-order adapter insertion points to ensure multi-task weight sharing while retaining task specificity: W′ = W + α(A shared B shared + A task B task ), where A shared , B shared represent the shared task weights, and A task B task represent the task-specific weights; Step 32: Enhance the fine-tuning effect in few-shot scenarios through few-shot weight initialization and regularization constraints. Among them, few-shot weight initialization combines domain knowledge or pre-trained weights to initialize the low-rank matrices A and B. Regularization constraints are to add regularization constraints to the update process of the low-rank matrices to prevent excessive deviation of model weights during few-shot fine-tuning. And before fine-tuning, generate domain-related embedding representations using the pre-trained model as the initial values of the low-rank matrices to reduce weight drift in few-shot scenarios, and then limit the change range of weights through regularization to ensure fine-tuning stability. Additionally, insert an additional regularization term in the low-order adaptive low-order adapter: Loss reg = λ(||A|| 2 + ||b|| 2 ); Step 33: Conduct a sensitivity analysis on matrices A and B inserted by the low-order adapter. For each weight or vector, calculate its gradient contribution to a specific task where L(W) is the task loss function. A large gradient contribution indicates high sensitivity, suggesting a significant impact on the model performance. The formula for measuring the sensitivity of weights is as follows: where S(W) represents the sensitivity of weight ω. Based on the results of the sensitivity analysis of weights, the weights are divided into two groups. For each group, a quantization level q(i) is defined to control the precision of the grouped weights: where q(i) represents the quantization level of the i-th group; S(ω i ) represents the sensitivity of the i-th weight; τ represents the grouping sensitivity threshold. For the high-sensitivity weight group, 8-bit Normal-Float quantization is used, and for the low-sensitivity weight group, 4-bit Normal-Float quantization is used. Combine the quantized weights with the low-rank matrix to optimize the inner product calculation of the matrix during inference. In each group of weights, perform the product calculation of the low-rank matrix; Combine the grouped weight labels and quantization levels, and finally integrate the weight calculation into the following formula: where W′ represents the final weight of the model, W represents the weight matrix, α represents the low-order adapter scaling factor to control the impact of the inserted weights; N represents the number of weight groups, q(i) represents the quantization level of the i-th group of weights, A i B i represents the core part of the weight adjustment by the low-rank matrix; Step 34: Insert task-related low-rank matrices in the attention layer or feed-forward layer of the base large language model for multi-task scenarios, and insert additional context processing weights between layers to enhance the cross-task context modeling ability.

5. An industrial scenario intelligent Q&A system based on a large language model according to claim 1, characterized in that: In the above-mentioned Step 4, the optimization of the retrieval module and generation module of the base large language model includes the following steps: Step 41: Convert the user query and document content into high-dimensional vectors using the fine-tuned and vectorized large language model in Step 3. Use a vector database to calculate similarity, screen relevant content, and calculate the similarity between the user query q and d based on dot product cosine similarity for vector retrieval: where q is the user query vector, d is the document vector, and n is the dimension of the vector; then, through similarity ranking, select the top k relevant documents: R = TopK(similarity(q, d1), Similarity(q, d2),..., Similarity(q, d N )); To improve the efficiency and accuracy of retrieval, a two-stage retrieval mechanism of rough retrieval and refined retrieval is adopted. In the rough retrieval stage, the vectorized user question q and the document vectors d1, d2, d3,... d n in the knowledge base are used to perform a preliminary similarity calculation, and a batch of highly relevant documents are screened out to obtain the preliminary relevant document set R coarse . The similarity threshold τ is used to screen the preliminary results: R coarse = {d|Similarity(q, d) > τ}. Then, based on the rule R(d), the results are filtered, and in the rough retrieval results, secondary screening is performed by combining domain rules to obtain the optimized high-quality document set R refined = {d|d ∈ R coarse , R(d) = 1}. Finally, the refined document set is output: R final = R refined . Then, the user question q and the refined highly relevant document set R refined are concatenated into the context input C and passed to the generation module; Step 42: Input the context C into a specific template rule T to generate an answer A: A = T(C). The generation module generates the answer step by step based on the retrieved content C = {R1, R2,...R n}, gradually generating the answer: A t = f(A t-1 , C), where A t represents the answer generated at the t-th step; A t-1 represents the answer generated in the previous round; f represents the generation function, and the final answer is the set of answers for all steps: Step 43: The retrieval module first splices relevant documents into a context input: where Q is the user query and R i is the retrieval result; then the generation module generates a structured answer through the task module based on the context input and achieves efficient collaboration through context fusion.

6. A method for using an intelligent question-answering system in an industrial scenario based on a large language model, characterized in that: For any industrial scenario intelligent question-answering system as described in Claims 1 to 5, perform the following steps: Step 1: The user inputs a question to the industrial scenario intelligent question-answering system; Step 2: The industrial scenario intelligent question-answering system tokenizes the question and performs semantic analysis; Step 3: The industrial scenario intelligent question-answering system performs vector retrieval on the question; Step 4: The industrial scenario intelligent question-answering system retrieves documents related to the question and performs screening and sorting; Step 5: The industrial scenario intelligent question-answering system calls the optimized large language model for input inference; Step 6: The industrial scenario intelligent question-answering system optimizes the answer; Step 7: The industrial scenario intelligent question-answering system displays the optimized returned answer to the user.