Industrial product anomaly detection and intelligent question and answer method and system based on large model
By combining large language models and traditional technologies, an industrial defect question and answer system is built, the problem that existing detection methods cannot deeply analyze the impact of defects is solved, accurate identification and intelligent diagnosis are achieved, and professional defect handling solutions are provided.
Patent Information
- Application Number
- CN202510407181.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-18
AI Technical Summary
Existing industrial defect detection methods cannot deeply analyze the impact of defects on product functions and performance, and lack intelligent interaction capabilities, so they cannot provide professional diagnostic suggestions and solutions.
Combining the cognitive reasoning ability of large language models and traditional defect detection technology, a real industrial scene image question and answer data set is constructed, defect detection is carried out through multimodal large language models, low-rank adaptation fine-tuning and joint training are carried out, and multi-task loss function is designed for model optimization.
It realizes accurate identification and impact assessment of industrial product defects, provides professional and reliable defect diagnosis and treatment solutions, and improves the professionalism and efficiency of the model in industrial applications.
Smart Images

Figure CN120336470A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence applications, and specifically relates to a large-model-based industrial product anomaly detection and intelligent question-answering method and system. Background Art
[0002] In the field of industrial production, product quality inspection and defect identification are key links to ensure product quality. At present, traditional industrial defect detection mainly relies on machine vision and deep learning methods to identify and segment defects. Although certain results have been achieved in appearance defect detection, there are still some problems: for example, the existing defect detection methods mainly focus on the visual feature recognition of defects, and cannot deeply analyze the actual impact of defects on product functions and performance, resulting in the lack of practical application value of the detection results; in addition, in actual industrial applications, operators often need to analyze the causes and make processing decisions for the detected defects, but the existing systems lack intelligent interaction capabilities and cannot provide professional diagnostic suggestions and solutions. Large language models show strong capabilities in natural language understanding and knowledge reasoning, but they are less used in the field of industrial defect detection, especially in the field of intelligent analysis. How to combine the cognitive capabilities of large language models with traditional defect detection technologies to achieve deep understanding and intelligent diagnosis of defects is a technical problem that needs to be solved urgently. Summary of the invention
[0003] In order to solve the problems existing in the prior art, the present invention provides an industrial product anomaly detection and intelligent question-answering method and system based on a large model. By combining the cognitive reasoning ability of a large language model with traditional defect detection technology, accurate identification of industrial product defects and impact assessment can be achieved. By constructing industrial field data and fine-tuning training, the professionalism of the model for defect detection and question-answering tasks is improved, so that the system can better serve actual industrial application scenarios. Based on the multi-round dialogue capability of the large model, professional and reliable defect diagnosis and processing solutions are provided to operators.
[0004] To achieve the above object, the present invention provides the following solutions:
[0005] A method for industrial product anomaly detection and intelligent question answering based on a large model, the method comprising:
[0006] Build a real industrial scene image question answering dataset;
[0007] Based on the constructed dataset, defect detection is performed using the trained multimodal large language model;
[0008] Conduct a comprehensive quantitative assessment of defect detection results.
[0009] Preferably, constructing a real industrial scene image question answering dataset includes:
[0010] Collect real images of industrial products;
[0011] Generate relevant questions based on the characteristics of each type of industrial product and a preset question template;
[0012] Annotate the relevant questions by domain experts in the corresponding field to generate standard answers containing professional technical judgments;
[0013] Form a training set D based on the real images of industrial products and the standard answers containing professional technical judgments train ={(I1,Q1,A1),(I2,Q2,A2),...,(I n ,Q n ,A n )}, where I i represents the preprocessed component image, Q i represents the questions designed based on professional knowledge, and A i represents the standard answer containing the basis for technical judgment.
[0014] Preferably, the multimodal large language model includes: three core modules, namely dual - path encoding, adapter fine - tuning, and large language model generation;
[0015] Among them, the dual - path encoding module is used to extract features from the input defect images and question texts; the adapter module is used to achieve the fusion and optimization of cross - modal features through two key components, multi - head attention and linear perceptron; the large language model is used to perform reasoning based on the fused features to generate defect diagnosis results.
[0016] Preferably, the training of the multimodal large language model includes: two stages, namely low - rank adaptation fine - tuning and joint training;
[0017] Among them, in the fine - tuning stage, the low - rank adaptation LoRA technology is used to perform domain - adaptation fine - tuning on the pre - trained multimodal large language model by introducing two low - rank matrix decompositions into the original model weight matrix :
[0018] ΔW = BA
[0019] Among them, d represents the dimension of the original weight matrix, k represents the dimension of the low - rank matrix, r << min(d,k) and r is the rank of the original matrix;
[0020] The calculation method of the updated weight matrix is:
[0021] W′ = W + αΔW
[0022] Among them, α is the scaling factor, W′ represents the updated weight, and ΔW is the update amount of the weight;
[0023] The fine-tuning process uses a learning rate η that meets the preset requirements, and the weight update formula is:
[0024]
[0025] where L(W t ) represents the loss value of the loss function with respect to the weight W t , represents the gradient, and W t represents the t-th weight matrix in the model;
[0026] After completing the fine-tuning, it enters the joint training stage. Based on the constructed industrial defect Q&A dataset, end-to-end model optimization is carried out. The training adopts a multi-task learning strategy. The samples in the dataset are organized in the form of "image-question-answer" triples, and a correctness label is assigned to each answer. A hybrid loss function is designed to optimize the model performance:
[0027] L total = λ1L qa + λ2L cls
[0028] where the Q&A loss L qa uses cross-entropy loss to guide the model to generate accurate defect descriptions by calculating the difference between the answer generated by the model and the standard answer; the correctness judgment loss L cls then uses binary cross-entropy loss to guide the model to learn to judge the accuracy of the generated answer; L total is the sum of the two losses to obtain the overall loss;
[0029] Through performance evaluation on the validation set and multiple rounds of optimization iterations, a model that can identify defect features and generate diagnostic results is finally obtained.
[0030] Preferably, a comprehensive quantitative evaluation of the defect detection results includes: constructing a hierarchical evaluation system, including an accuracy rate index and a logic scoring index.
[0031] Preferably, at the accuracy rate evaluation level, two measurement indicators, single-sample accuracy rate and grouped accuracy rate, are introduced;
[0032] where the formula for calculating the single-sample accuracy rate is where N correct represents the number of questions correctly answered by the model, and N total represents the total number of questions in the test set;
[0033] The formula for evaluating the grouped accuracy rate is defined as where N group_correct represents the number of groups in which the model achieves all correct answers in the Q&A group of the same image, that is, all questions corresponding to a defect sample are correctly answered, Ngroup_total Represents the total number of defective samples in the test set.
[0034] Preferably, in terms of answer logic scoring, a multi-dimensional evaluation index system is constructed based on answer accuracy;
[0035] Among them, the defect classification score Evaluates the recognition ability of the evaluation model for different types of defects;
[0036] Spatial positioning score Measures the accuracy of the model in positioning the defect location;
[0037] Inference completeness score Evaluates the logical integrity of the model's diagnostic analysis; among them, w1, w2, and w3 correspond to the weight coefficients of different evaluation dimensions, reflecting the importance of each dimension in the comprehensive score, and satisfying the normalization constraint w1 + w2 + w3 = 1; N C_type Represents the number of defect types correctly identified by the model, N C_location Represents the number of defect locations correctly located by the model, N C_reason Represents the number of samples for which the model can correctly perform causal analysis, logical inference, or comprehensive judgment during the inference process;
[0038] The formula for calculating the system comprehensive performance score is S = λ1ST + λ2SL + λ3SR, where λ i (i = 1, 2, 3) are the weight coefficients of each dimension and satisfy Normalization constraint;
[0039] Through this evaluation index system, the performance of the model in the industrial defect diagnosis task is quantitatively analyzed and systematically evaluated.
[0040] The present invention also provides an industrial product anomaly detection and intelligent question-answering system based on a large model, which is used to implement any one of the above methods, and is characterized in that the system includes: a construction module, a detection module, and an evaluation module;
[0041] The construction module is used to construct a real industrial scenario image question-answering data set;
[0042] The detection module is used to perform defect detection using a trained multi-modal large language model based on the constructed data set;
[0043] The evaluation module is used to comprehensively and quantitatively evaluate the defect detection results.
[0044] Compared with the prior art, the beneficial effects of the present invention are:
[0045] By adopting the technical solutions of dual-channel encoding and multi-task training in the industrial defect Q&A system, the present invention can effectively improve the model's understanding and Q&A capabilities for industrial defects; through the low-rank adaptation fine-tuning technology, while ensuring the model performance, the number of trainable parameters is reduced to 0.1% of the original model, significantly improving the model training efficiency; by training with the constructed high-quality industrial defect Q&A dataset, the model achieves excellent performance with a single-sample accuracy of 87% and a grouped accuracy of 47% respectively; evaluated based on a multi-dimensional scoring system, the scores in the three dimensions of defect type recognition, location positioning, and reasoning process completeness all reach above 4.5 points, verifying the excellent performance of the model in the industrial defect Q&A task and providing a reliable technical solution for intelligent quality inspection in the industrial field. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions of the present invention, the accompanying drawings required for use in the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0047] Figure 1 Schematic diagram of the defect Q&A system construction process for the embodiments of the present invention;
[0048] Figure 2 Schematic diagram of the industrial defect Q&A dataset construction process for the embodiments of the present invention;
[0049] Figure 3 Schematic diagram of the multi-modal large language model fine-tuning and adaptation optimization process for the embodiments of the present invention;
[0050] Figure 4 Schematic diagram of the real image of electronic components for the embodiments of the present invention;
[0051] Figure 5 Schematic diagram of a method for industrial product anomaly detection and intelligent Q&A based on a large model for the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0053] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the drawings and specific embodiments.
[0054] Example 1
[0055] As Figure 5 shown, an embodiment of the present invention provides an industrial product anomaly detection and intelligent Q&A method based on a large model, and the method includes:
[0056] Construct a real industrial scenario image Q&A dataset;
[0057] Based on the constructed dataset, use the trained multi-modal large language model for defect detection;
[0058] Conduct a comprehensive quantitative evaluation of the defect detection results.
[0059] In this embodiment, the construction process of the defect Q&A system is as Figure 1 shown:
[0060] This is the design process of an intelligent Q&A system for industrial defect scenarios: Since there is a lack of publicly available Q&A datasets in the industrial defect field, the research first constructed a high-quality defect Q&A sample library based on standard industrial defect datasets such as MVTec through a systematic Q&A generation method combined with manual annotation. Secondly, to improve the model's understanding ability and Q&A performance in the industrial defect field, the research adopted a variety of technical means to adapt and train the large model. In terms of instruction fine-tuning, a set of professional instruction templates for industrial defect scenarios was designed, including multi-dimensional task guidance such as defect feature description, location positioning, and type judgment, to enhance the model's understanding and expression ability of industrial defects. At the same time, the adapter technology was introduced, and by inserting lightweight trainable modules into the pre-trained model, efficient domain adaptation was achieved, which not only maintained the model's general understanding ability but also significantly improved its professional performance in the industrial defect field. Finally, a comprehensive performance evaluation and verification of the trained model were conducted from multiple dimensions such as answer accuracy, defect type recognition, defect location positioning, and details of the reasoning process.
[0061] In this embodiment, the construction of the industrial defect Q&A dataset is as Figure 2 shown:
[0062] The construction of the industrial defect Q&A dataset is based on the MVTec AD dataset. This dataset contains images of 15 categories of industrial products and objects, each category including normal samples and different types of defect samples. For each defect case, the basic information of the object, object description, defect type, and defect description can be systematically recorded using the images in the dataset and the corresponding defect annotation information. Specifically, the basic information of the object comes from the product categories in the dataset, such as bottles, cables, etc.; the object description is based on the feature description of the normal sample images; the defect type uses the specific types marked in the dataset, such as contamination, cutting, holes, etc.; the defect description is explained according to the detailed defect features of the defect sample images and annotations. For example, in the cable category, the normal sample shows the cross-sectional structure of the cable, and we can record its features such as "a cross-sectional image of multiple cables of different colors and sizes neatly arranged". For the defect samples, according to the annotations in the dataset, we can accurately describe specific defect types and manifestations such as "the cable is deformed under external force, resulting in the overall structure being bent" or "the outer insulation protection layer of the cable is damaged, causing the internal wires to be exposed". This description method based on the standard dataset not only ensures the professionalism and accuracy of the data but also provides a reliable information basis for subsequent Q&A generation.
[0063] Based on this basic information, a large language model is used to generate Q&A. By designing reasonable prompt templates, the model is guided to generate relevant questions and answers around multiple dimensions such as the existence, location, type, and impact of defects. In the cable case, the model generated typical questions such as "Does this cable have any visible defects?" and "Where is the defect located in this cable cross-section?", as well as the corresponding accurate answers. Through reasonable prompt engineering design, it is ensured that the generated Q&A pairs cover the key aspects of defect diagnosis and analysis, providing a comprehensive perspective on defect understanding. The model will consider the particularity of the industrial scenario during the generation process, using corresponding professional terms and expressions for different types of defects, making the generated Q&A pairs both professional and easy to understand. At the same time, through multi-round optimized prompt templates, it is ensured that the generated questions are hierarchical and logical, forming a complete Q&A chain from the basic identification of defects to in-depth risk assessment.
[0064] The generated Q&A pairs also go through a manual review process to evaluate and correct the accuracy, professionalism, and expression norms of the Q&A. The review process mainly includes three aspects: First, check the accuracy of the Q&A content to ensure that the Q&A content strictly corresponds to the defect images and annotation information in the MVTec AD dataset; second, standardize the use of professional terms to avoid ambiguous expressions; third, optimize the description logic to make the Q&A structure clear and hierarchical. For different types of industrial defects, detailed review criteria are formulated based on the existing annotation information in the dataset, with a focus on whether the Q&A accurately reflects the key features of the defects. Through this meticulous review process, it is ensured that the finally constructed Q&A dataset can truly reflect the characteristics of industrial defects and provide a reliable training basis for subsequent industrial defect Q&A systems.
[0065] In this embodiment, the fine-tuning and adaptation optimization of the multi-modal large language model are as Figure 3 shown:
[0066] Model architecture: After completing the construction of the industrial defect Q&A dataset, this solution proposes a defect Q&A system based on a multi-modal large language model, which realizes efficient industrial defect diagnosis and analysis by deeply integrating visual and language information. As shown in the figure, the overall system consists of three core modules: dual-path encoding, adapter fine-tuning, and large language model generation. Among them, the dual-path encoding module extracts features from the input defect images and question texts respectively; the adapter module realizes the fusion and optimization of cross-modal features through two key components, multi-head attention and linear perceptron; the large language model then performs reasoning based on the fused features to generate accurate and professional defect diagnosis results. Based on the constructed high-quality Q&A dataset, model training and optimization are carried out for different types of industrial defect scenarios, enabling the system to accurately identify and analyze various complex defect situations. Through systematic training and tuning, the model can fully understand the visual features of defect images and the semantic information of related questions, accurately answer industrial detection questions, and provide intelligent support for quality control and defect diagnosis in industrial production.
[0067] Training process: Model training is divided into two stages: low-rank adaptation fine-tuning and joint training. In the fine-tuning stage, the low-rank adaptation (LoRA) technique is used to perform domain adaptation fine-tuning on the pre-trained model. By adding two low-rank matrices A and B to the original model weight matrix (d represents the dimension of the original weight matrix, and k represents the dimension of the low-rank matrix), and decomposing them:
[0068] ΔW = BA
[0069] where, d represents the dimension of the original weight matrix, k represents the dimension of the low-rank matrix, r << min(d, k) and r is the rank of the original matrix.
[0070] The calculation method of the updated weight matrix is as follows:
[0071] W′ = W + αΔW
[0072] Where α is the scaling factor, W′ represents the updated weight, and ΔW is the update amount of the weight. This fine-tuning method reduces the number of trainable parameters from d×k to r(d + k), significantly improving the training efficiency. A relatively small learning rate η is used in the fine-tuning process, and the weight update formula is:
[0073]
[0074] Where L(W t ) represents the loss value of the loss function with respect to the weight W t , represents the gradient, and W t represents the t-th weight matrix in the model. This fine-tuning method significantly reduces the number of trainable parameters, improving the training efficiency while maintaining the model performance. A relatively small learning rate is used in the fine-tuning process to maintain the original feature extraction ability of the model, and at the same time, the model parameters are gradually adjusted to adapt to the industrial defect scenario.
[0075] After the fine-tuning is completed, it enters the joint training stage, and end-to-end model optimization is performed based on the constructed industrial defect Q&A dataset. The training adopts a multi-task learning strategy, organizing the samples in the dataset into the form of "image - question - answer" triples, and annotating the correctness label for each answer. A hybrid loss function is designed to optimize the model performance:
[0076] L total = λ1L qa + λ2L cls
[0077] Where the Q&A loss L qa adopts the cross-entropy loss, guiding the model to generate accurate defect descriptions by calculating the difference between the answer generated by the model and the standard answer; the correctness judgment loss L cls adopts the binary cross-entropy loss, guiding the model to learn to judge the accuracy of the generated answer, and the sum of the two losses gives the overall loss L total . This multi-stage and multi-task training method not only ensures the model's in-depth understanding of industrial defect features but also improves the professionalism and reliability of the generated content. Through performance evaluation on the validation set and multiple rounds of optimization iterations, a model that can accurately identify defect features and generate reliable diagnostic results is finally obtained.
[0078] In this embodiment, regarding the performance evaluation of the defect Q&A system:
[0079] To comprehensively evaluate the performance of the proposed industrial defect Q&A system, this study constructed a hierarchical evaluation system, including an accuracy index and a logical scoring index. At the accuracy evaluation level, two measurement indicators, single-sample accuracy and grouped accuracy, were introduced. The formula for single-sample accuracy is where N correct represents the number of questions correctly answered by the model, and N total represents the total number of questions in the test set. The evaluation formula for grouped accuracy is defined as where N group_correct represents the number of groups in which the model achieves all correct answers in the Q&A group of the same image (i.e., all questions corresponding to a defective sample are correctly answered), and N group_total represents the total number of defective samples in the test set.
[0080] In terms of answering logic scoring, a multi-dimensional evaluation index system was constructed based on answer accuracy. The defect classification score evaluates the model's ability to identify different types of defects; the spatial positioning score measures the accuracy of the model in locating the defect position; the reasoning completeness score evaluates the logical integrity of the model's diagnostic analysis. Among them, w1, w2, and w3 correspond to the weight coefficients of different evaluation dimensions, reflecting the importance of each dimension in the comprehensive score, and satisfying the normalization constraint w1 + w2 + w3 = 1; N C_type represents the number of defect types correctly identified by the model, N C_location represents the number of defect positions correctly located by the model, N C_reason represents the number of samples in which the model can correctly perform causal analysis, logical inference, or comprehensive judgment during the reasoning process. The formula for the comprehensive performance score of the system is S = λ1ST + λ2SL + λ3SR, where λ i (i = 1, 2, 3) are the weight coefficients of each dimension and satisfy the normalization constraint. Through this evaluation system, the performance of the model in the industrial defect diagnosis task can be quantitatively analyzed and systematically evaluated. The evaluation results show that the system demonstrates good accuracy and generalization in the industrial defect diagnosis task.
[0081] In summary, 1. The present invention provides a method for constructing a multi-modal Q&A dataset based on industrial product defect images. This method performs professional Q&A annotation for typical defects such as electronic components to form standardized training data containing image-question-answer triples;
[0082] 2. The present invention proposes an industrial product defect diagnosis method based on a multi-modal large language model. Through the dual-channel feature extraction of a visual encoder and a language encoder, combined with the cross-modal feature fusion of an adapter module, accurate identification and analysis of defects are achieved;
[0083] 3. The present invention provides a method for implementing a hierarchical evaluation system, which combines the single-sample accuracy rate, the grouped accuracy rate, and the scoring indicators in three dimensions of defect classification, spatial positioning, and reasoning completeness to achieve a comprehensive quantitative evaluation of the defect diagnosis results.
[0084] Example Two
[0085] The present invention provides a defect Q&A system based on industrial scenario images, and the specific implementation steps are as follows:
[0086] Step 1: Construction of a real industrial scenario image Q&A data set
[0087] The present invention first collects real images of electronic components, including resistor types (thin film resistors, wafer resistors), inductor types (block inductors, surface mount power inductors), connector types (Type-C connectors, copper pillars), button types (long-legged buttons, short-legged buttons), LED types (LED lamp beads, LED pads), and fasteners (flat nuts), etc. For the characteristics of each type of component, relevant questions are generated based on a preset question template, and the question template includes: "What is the nominal value of the resistor in the figure?", "How is the pin integrity of the Type-C connector?", "Is the welding quality of the LED pad qualified?", etc. Then, domain experts with an electronic engineering professional background annotate these questions to generate standard answers containing professional technical judgments. Finally, a training set D train ={(I1,Q1,A1),(I2,Q2,A2),...,(I n ,Q n ,A n )} is formed, where I i represents the preprocessed component image, Q i represents the question designed based on professional knowledge, and A i represents the standard answer containing the basis for technical judgment. As Figure 4 shown in Table 1.
[0088] Table 1
[0089]
[0090] Step 2: Industrial defect diagnosis reasoning
[0091] Based on the constructed training data set, the constructed multimodal large language model is used for defect diagnosis reasoning. The trained multimodal large language model is used to diagnose defects in images of electronic components such as block inductors. First, the image of the electronic component to be inspected is input into the visual encoder to extract visual features, and the preset inspection questions such as "Is there any abnormality?" are input into the language encoder to obtain text features. The output features of the visual encoder and the language encoder are fused through the designed adapter module, in which the multi-head attention mechanism automatically aligns the association between the image area and the question keywords to achieve effective integration of cross-modal information. Finally, the fused features are input into the large language model for reasoning, and a standardized diagnostic answer is generated for each question. For example, for the question "Is there any abnormality?", the model output result is "Yes, there are defects in multiple views of the square inductor", which accurately points out the location and type of the defect.
[0092] Step 3: Defect Diagnosis and Evaluation
[0093] In order to verify the reliability of the diagnosis results, a hierarchical evaluation method is adopted. First, the single sample accuracy is calculated based on all test images, that is, the number of correct answers of the statistical model when answering single questions such as "Is there a defect" and "Does the defect affect the conductivity" is obtained to obtain the accuracy; at the same time, the grouping accuracy is calculated, requiring the model to correctly answer all 5 questions of the same image sample to count the correct group number to obtain the grouping accuracy. Secondly, the diagnostic level of the model for each defective sample is evaluated from three dimensions: defect classification, location positioning, and reasoning completeness. For example, for the above-mentioned block inductor sample, its ST score (accurately identifying the defect type), SL score (correctly positioning to the middle and lower right of the 2nd, 3rd and 4th views) and SR score (logical completeness of the analysis result) are calculated respectively. Finally, the scores of the three dimensions are weighted and summed according to the preset weight coefficients to obtain the comprehensive score S of the defective sample. In this way, this embodiment realizes a comprehensive quantitative evaluation of the defect diagnosis results.
[0094] Embodiment 3
[0095] The present invention also provides an industrial product anomaly detection and intelligent question-answering system based on a large model, the system is used to implement any one of the methods described, the system comprises: a construction module, a detection module, and an evaluation module;
[0096] A construction module for building a real industrial scene image question answering dataset;
[0097] The detection module is used to perform defect detection based on the constructed dataset using the trained multimodal large language model;
[0098] Evaluation module, used to conduct comprehensive quantitative evaluation of defect detection results.
[0099] The embodiments described above are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. An industrial product anomaly detection and intelligent question answering method based on a large model, characterized in that, The method includes: Constructing a real industrial scenario image question-answering dataset; Based on the constructed dataset, using a trained multi-modal large language model for defect detection; Conducting a comprehensive quantitative evaluation of the defect detection results.
2. The method according to claim 1, characterized in that, Constructing a real industrial scenario image question-answering dataset includes: Collecting real images of industrial products; Generating relevant questions based on preset question templates according to the characteristics of each type of industrial product; Annotating the relevant questions by domain experts in the corresponding field to generate standard answers containing professional technical judgments; Form a training set D based on the real images of industrial products and the standard answers containing professional technical judgments train ={(I1, Q1, A1), (I2, Q2, A2),..., (I n , Q n , A n )}, where I i represents the preprocessed component image, Q i represents the questions designed based on professional knowledge, and A i represents the standard answer containing the basis for technical judgment.
3. The method according to claim 1, characterized in that, The multi-modal large language model includes: three core modules, namely dual-route encoding, adapter fine-tuning, and large language model generation; Among them, the dual-route encoding module is used to extract features from the input defect images and question texts; the adapter module is used to achieve the fusion and optimization of cross-modal features through two key components, multi-head attention and linear perceptron; the large language model is used to reason based on the fused features and generate defect diagnosis results.
4. The method according to claim 1, wherein The training of the multi-modal large language model includes: two stages, low-rank adaptation fine-tuning and joint training; Among them, in the fine-tuning stage, the Low-Rank Adaptation (LoRA) technique is used to perform domain adaptation fine-tuning on the pre-trained multi-modal large language model by introducing two low-rank matrix decompositions into the original model weight matrix : ΔW = BA Among them, d represents the dimension of the original weight matrix, k represents the dimension of the low-rank matrix, r << min(d, k) and r is the rank of the original matrix; The calculation method of the updated weight matrix is: W′ = W + αΔW Where α is the scaling factor, W′ represents the updated weight, and ΔW is the update amount of the weight; During the fine-tuning process, a learning rate η that meets the preset requirements is adopted, and the weight update formula is: where \(L(W\) t ) represents the loss value of the loss function with respect to the weight \(W\) t , represents the gradient, and \(W\) t represents the \(t\)-th weight matrix in the model; After completing the fine-tuning, enter the joint training stage. Based on the constructed industrial defect question-answering dataset, conduct end-to-end model optimization. The training adopts a multi-task learning strategy, organizes the samples in the dataset into the form of "image-question-answer" triples, and annotates the correctness label for each answer. A mixed loss function is designed to optimize the model performance: L total = λ1L qa + λ2L cls Among them, the Q&A loss L qa adopts the cross-entropy loss and guides the model to generate accurate defect descriptions by calculating the difference between the answers generated by the model and the standard answers; the correctness judgment loss L cls adopts the binary cross-entropy loss to guide the model to learn to judge the accuracy of the generated answers; L total is the sum of the two losses to obtain the overall loss; Through performance evaluation on the validation set and multiple rounds of optimization iterations, finally obtain a model that can identify defect features and generate diagnosis results.
5. The method according to claim 1, characterized in that Conducting a comprehensive quantitative evaluation of the defect detection results includes: constructing a hierarchical evaluation system, including an accuracy rate index and a logical scoring index.
6. The method according to claim 5, characterized in that At the level of accuracy rate evaluation, two measurement indicators, single-sample accuracy rate and grouped accuracy rate, are introduced; Among them, the calculation formula for single-sample accuracy is where N correct represents the number of questions correctly answered by the model, and N total represents the total number of questions in the test set; The formula for evaluating the grouping accuracy is defined as where N group_correct represents the number of groups in which the model achieves all correct answers in the Q&A groups of the same image, that is, all questions corresponding to a defective sample are correctly answered, and N group_total represents the total number of defective samples in the test set.
7. The method according to claim 6, characterized in that In terms of answer logic scoring, a multi-dimensional evaluation index system is constructed based on answer accuracy; Among them, the defect classification score evaluates the recognition ability of the model for different types of defects; Spatial positioning score Measure the accuracy of the model in positioning the defect location; Inference Completeness Score Evaluate the logical integrity of the diagnostic analysis of the model; where w1, w2, w3 are the weight coefficients corresponding to different evaluation dimensions, reflecting the importance of each dimension in the comprehensive score, satisfying the normalization constraint w1 + w2 + w3 = 1; N C_type Indicates the number of defect types correctly identified by the model, N C_location Indicates the number of defect locations correctly located by the model, N C_reason Indicates the number of samples that the model can correctly perform causal analysis, logical inference, or comprehensive judgment during the inference process; The calculation formula for the system comprehensive performance score is S = λ1ST + λ2SL + λ3SR, where λ i (i = 1, 2, 3) are the weight coefficients of each dimension, and satisfy normalization constraint; Through this evaluation index system, conduct quantitative analysis and systematic evaluation of the model's performance in the industrial defect diagnosis task.
8. An industrial product anomaly detection and intelligent Q&A system based on a large model, the system being used to implement the method described in any one of claims 1-7, characterized in that, The system includes: a construction module, a detection module, and an evaluation module; The construction module is used to construct a real industrial scenario image question-answering dataset; The detection module is used to perform defect detection based on the constructed dataset using a trained multi-modal large language model; The evaluation module is used to conduct a comprehensive quantitative evaluation of the defect detection results.
Citation Information
Cited By
Industrial anomaly detection method, system, equipment and medium
CN121350465A
Instruction fine tuning data set generation method, electronic equipment and storage medium
CN121579077A