Industrial defect intelligent detection method and system based on visual enhancement fine tuning large model

Through the multimodal feature fusion strategy based on visual enhancement fine-tuning large model, the problems of low manual detection efficiency and poor adaptability of traditional automated detection are solved, and efficient, accurate and highly adaptable industrial defect detection is achieved, which is suitable for multi-variety and small-batch production scenarios.

CN120411050APending Publication Date: 2025-08-01SHINE OPTICS TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510547852.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

In the prior art, artificial visual inspection has low efficiency and fluctuations in accuracy, making it difficult to meet the real-time inspection needs of high-speed production lines. In addition, traditional automated inspection lacks the comprehensive analysis capabilities for complex industrial scenarios, and it is difficult to adapt to the flexible production needs of multiple varieties and multiple processes.

Method used

An intelligent industrial defect detection method based on visual enhancement fine-tuning large model is adopted. Through feature extraction and alignment of multi-spectral image data, process parameter real-time data and historical detection record data, combined with multimodal feature fusion strategy, the visual enhancement fine-tuning large model is used for defect detection, so as to achieve data correlation across batches and time and highly adaptable detection.

Benefits of technology

It improves the efficiency and accuracy of industrial defect detection, is highly adaptable, can quickly adapt to new scenarios, reduce data acquisition and labeling costs, shorten production line replacement time, and improve production efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120411050A_ABST
    Figure CN120411050A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of industrial defect detection, in particular to an industrial defect intelligent detection method and system based on a visual enhancement fine-tuning large model, and the method comprises the steps: receiving the demand data of a user, and determining a corresponding detection scene and an industrial camera corresponding to the detection based on the demand data; judging whether the current detection scene is a new scene or not; when the judgment result is no, multispectral image data collected by the industrial camera and corresponding process parameter real-time data are collected; historical detection record data are called from a historical database; performing feature extraction and alignment on the multispectral image data, the process parameter real-time data and the historical detection record data; performing feature fusion on the aligned features, and outputting corresponding multi-modal feature data; and determining a visual enhancement fine tuning model, taking the output multi-modal feature data as input data, and outputting corresponding defect detection result data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of industrial defect detection, and particularly to an industrial defect intelligent detection method and system based on a vision-enhanced fine-tuned large model. Background Art

[0002] In the process of industrial production, defect detection is a key link to ensure product quality. The current mainstream detection methods still mainly rely on manual visual inspection, supplemented by a small amount of traditional algorithm detection. The main problems are as follows:

[0003] Manual detection relies on quality inspection personnel to visually judge each piece one by one. Limited by the visual range of the human eye (only perceiving the visible light band) and fatigue characteristics, it is difficult to meet the real-time detection requirements of high-speed production lines. Not only is the detection efficiency low, but also the accuracy fluctuates greatly.

[0004] Secondly, existing automated detections are mostly based on single-scene data (such as visible light images or single process parameters), lacking the comprehensive analysis ability for complex industrial scenes. Manual preset detection rules are required, and it is difficult to adapt to the flexible production requirements of multiple varieties and multiple processes.

[0005] Based on this, there is an urgent need for an industrial defect intelligent detection method and system based on a vision-enhanced fine-tuned large model, which can realize efficient, accurate and highly adaptable industrial defect detection functions, and greatly improve the efficiency, high accuracy and strong adaptability of industrial defect detection. Summary of the Invention

[0006] One of the purposes of the present invention is to provide an industrial defect intelligent detection method and system based on a vision-enhanced fine-tuned large model, which can realize efficient, accurate and highly adaptable industrial defect detection functions, and greatly improve the efficiency, high accuracy and strong adaptability of industrial defect detection.

[0007] To achieve the above purpose, an industrial defect intelligent detection method based on a vision-enhanced fine-tuned large model is provided, including the following steps:

[0008] S1. Receive the demand data corresponding to the user, and determine the corresponding detection scene and the industrial camera corresponding to this detection based on the demand data;

[0009] S2. According to the determined detection scene, judge whether the current detection scene is a new scene;

[0010] S3. When the judgment result is that the current detection scene is not a new scene, collect the multi-spectral image data collected by the industrial camera and the real-time data of the corresponding process parameters;

[0011] S4. Retrieve the corresponding historical detection record data from the historical database;

[0012] S5. Based on a preset feature alignment strategy, extract and align features from the multispectral image data, real-time process parameter data, and historical detection record data;

[0013] S6. After the features extracted from the multispectral image data, real-time process parameter data, and historical detection record data are aligned, based on a preset multimodal feature fusion strategy, fuse the aligned features and output corresponding multimodal feature data;

[0014] S7. Based on the current detection scenario, retrieve the customized AOI scheme corresponding to this detection scenario from the historical database, determine the corresponding vision-enhanced fine-tuning large model, take the output multimodal feature data as input data, input it into the vision-enhanced fine-tuning large model corresponding to this detection scenario, and output corresponding defect detection result data.

[0015] The technical principle and effect of this solution: In this solution, user requirement data is received to determine the detection scenario and industrial camera. It is judged whether the detection scenario is a new scenario. New scenarios need additional processing, while non-new scenarios directly collect data. The multispectral image data and real-time process parameter data collected by the industrial camera are acquired, and at the same time, historical detection record data is retrieved from the historical database. Multispectral images can provide rich visual information; process parameters reflect the state of the production process; historical data is used for comparative analysis and provides a reference for defect judgment.

[0016] Using the preset feature alignment strategy, key features are extracted from multi-source data, and the features of different modal data are aligned in the same scale and semantic space, which is convenient for subsequent fusion. The multimodal feature fusion strategy is adopted to fuse the aligned features to obtain multimodal feature data, integrating the advantages of different types of data and enabling the model to obtain more comprehensive information. According to the detection scenario, the customized AOI scheme and the corresponding vision-enhanced fine-tuning large model are retrieved from the historical database, the multimodal feature data is input into the model, and the model uses the patterns and rules learned from a large amount of data to output defect detection result data.

[0017] Traditional supervised fine-tuning (SFT) models are prone to overfitting due to insufficient data in few-shot scenarios, resulting in a significant decline in detection performance. This solution constructs a richer feature space by fusing multi-source data (multispectral images, process parameters, historical detection records), and through retrieving historical detection record data and combining the preset feature alignment strategy, realizes cross-batch and cross-time data association. Even when the data volume of the current detection scenario is small, the transfer learning of historical data can effectively alleviate the overfitting problem and improve the generalization ability of the model in few-shot scenarios.

[0018] Single-modal visual models only focus on image data and ignore the internal relationship between process parameters and defects, resulting in difficult traceability of detection results. Through a multi-modal feature fusion strategy, this solution realizes the fusion of multi-modal features. The comprehensive feature vector formed after the fusion of multi-modal features can capture complex associations that cannot be represented by a single modality, greatly improving the efficiency and accuracy of process defect detection. That is, it realizes the functions of efficient, accurate, and highly adaptable industrial defect detection, greatly improving the efficiency, high accuracy, and strong adaptability of industrial defect detection.

[0019] Further, the preset feature alignment strategy is as follows:

[0020] Input the multi-spectral image data, real-time process parameter data, and historical detection record data into the visual encoder, process parameter encoder, and historical knowledge encoder respectively. Based on the feature extraction strategies corresponding to each encoder, extract the feature vectors corresponding to the multi-spectral image data, real-time process parameter data, and historical detection record data respectively and align them into a unified feature space.

[0021] Beneficial effects: This solution realizes the standardized processing of heterogeneous data through each data-specific encoder and the mapping mechanism of the unified feature space, greatly activating the correlation potential between cross-modal data, greatly improving the richness and consistency of feature expression, and increasing the robustness of the model.

[0022] Further, the preset multi-modal feature fusion strategy is as follows:

[0023] Concatenate the aligned feature vectors to obtain the corresponding concatenated feature vectors;

[0024] Input the concatenated feature vectors into each preset fusion function in sequence, take the output of the previous fusion function as the input, and input it into the next fusion function until the output of the last fusion function is used as the finally fused multi-modal feature data;

[0025] During the fusion process, according to the current corresponding detection scenario, determine the weight ratios corresponding to the feature vectors of the multi-spectral image data, real-time process parameter data, and historical detection record data respectively in each fusion function.

[0026] Beneficial effects: The complexity of industrial detection scenarios (such as product types, defect forms, and process stage differences) requires the importance of different modality data to be dynamically adjustable. This strategy is realized through a scene-aware weight allocation mechanism. According to the current detection scenario, use the attention mechanism to automatically adjust the weight ratios of the multi-spectral image, process parameters, and historical record features. Through dynamic weight adjustment, complementary features of different modalities can be forced to be activated.

[0027] Further, the visually enhanced fine-tuning large model preset in S7 is a model trained based on a preset training strategy;

[0028] The training strategy is as follows:

[0029] S700. Obtain historical multispectral image data, corresponding historical process parameter real-time data, and historical detection record data from the historical database; obtain corresponding actual defect detection data from the historical database;

[0030] S701. According to the historical multispectral image data, corresponding historical process parameter real-time data, and historical detection record data, execute steps S5 and S6, and output corresponding historical multimodal feature data;

[0031] S702. Construct a visually enhanced fine-tuning large model;

[0032] S703. Use the output historical multimodal feature data as input data and input it into the constructed visually enhanced fine-tuning large model to output corresponding training defect detection result data;

[0033] S704. According to the output training defect detection result data and corresponding actual defect detection data, match corresponding verifiable reward functions from the database, calculate the function values corresponding to the corresponding verifiable reward functions, and adjust the model parameters corresponding to the visually enhanced fine-tuning large model according to the corresponding function values until the corresponding function values meet the preset requirements. At this time, the training of the visually enhanced fine-tuning large model is completed. The verifiable reward functions include visual reward functions and industrial reward functions.

[0034] Beneficial effects: In this solution, by integrating historical multispectral image data (visual features), historical process parameter real-time data (process logic features), and historical detection record data (business rule features), a multimodal feature system including image vision, production processes, and quality standards is constructed. By extracting multimodal features through steps S5 and S6, the deep association between image pixel features and process parameters is realized.

[0035] Introduce visual reward functions (such as image segmentation accuracy, defect classification accuracy) and industrial reward functions (such as detection efficiency, false alarm rate control), and transform the defect detection problem into a reinforcement learning optimization task. Through the comparison between the training defect detection results and actual detection data, dynamically adjust the model parameters (such as neural network weights) to form a self-optimizing closed loop of "data input → model inference → reward feedback → parameter update".

[0036] Further, S3 further includes:

[0037] When the judgment result indicates that the current detection scenario is a new scenario, a small amount of multi-spectral image data and real-time process parameter data collected by the industrial camera corresponding to this detection scenario are acquired, and the corresponding historical detection record data is retrieved;

[0038] The actual defect detection results corresponding to the multi-spectral image data and real-time process parameter data at this time are labeled to form corresponding labeled content, and the labeled content is associated with the corresponding collected data;

[0039] According to the multi-spectral image data with labeled content, real-time process parameter data, and historical detection record data, steps S701 to S704 are executed, and a fine-tuned visual enhancement fine-tuning large model is output;

[0040] After the fine-tuned visual enhancement fine-tuning large model is output, according to the multi-spectral image data, real-time process parameter data, and historical detection record data obtained in real time currently, steps S5 to S7 are executed, and the corresponding defect detection result data is output. At this time, the visual enhancement fine-tuning large model in S7 is the fine-tuned visual enhancement fine-tuning large model.

[0041] Beneficial effects: In this solution, when the detection scenario is a new scenario (such as a new product model, new process parameters), only a small amount of multi-spectral image data and process parameters need to be collected, combined with historical detection record data (such as defect standards for similar products), and a suitable model can be quickly constructed through fine-tuning. Traditional deep learning models require a large amount of labeled data (usually hundreds to thousands of images), while this process compresses the amount of labeled data required for the new scenario to "a small amount" through transfer learning plus reinforcement fine-tuning, greatly reducing the data collection and annotation costs, especially suitable for multi-variety, small-batch production scenarios. Retrieving historical detection record data (such as defect type definitions, severity grading standards) can transfer the prior knowledge of the old scenario to the new scenario, realizing cross-scenario knowledge reuse and reducing the workload of manually redefining standards.

[0042] Manually label a small amount of data in the new scenario (such as marking the defect location and type), and strongly associate and store it with the collected images and process parameters. This "collect and label immediately" mechanism ensures that: the labeled data directly reflects the actual needs of the new scenario, avoiding the distribution deviation between historical data and the new scenario; a real-time feedback link of "new data - new annotation - model fine-tuning" is formed, enabling the model to adapt to scenario changes, rather than relying on periodic large-scale data updates.

[0043] The fine-tuned model can be directly connected to the real-time detection process (executing S5 - S7), realizing full automation from image acquisition, feature extraction, defect inference to result output. There is no need for manual intervention in model deployment, which is suitable for scenarios where product models are quickly switched on the production line (such as flexible manufacturing production lines), shortening the changeover time and improving production efficiency.

[0044] Further, S3 further includes:

[0045] When outputting the fine-tuned vision-enhanced large model, associate the fine-tuned vision-enhanced large model with the corresponding current detection scenario to form a customized AOI solution corresponding to the current detection scenario.

[0046] Beneficial effects: Different detection scenarios vary in terms of product characteristics, process requirements, defect types, etc. The customized AOI solution can be optimized according to the specific characteristics of the current detection scenario to ensure that the detection ability of the model highly matches the scenario requirements.

[0047] The present invention also provides an industrial defect intelligent detection system based on a vision-enhanced large model, using the above-mentioned industrial defect intelligent detection method based on a vision-enhanced large model. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 It is a flowchart of the industrial defect intelligent detection method based on a vision-enhanced large model in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0049] The following is a further detailed description through specific embodiments:

[0050] Embodiment 1

[0051] The industrial defect intelligent detection method based on a vision-enhanced large model is basically as Figure 1 shown, and includes the following steps:

[0052] S1. Receive the demand data corresponding to the user, and determine the corresponding detection scenario and the industrial camera corresponding to this detection based on the demand data;

[0053] S2. According to the determined detection scenario, determine whether the current detection scenario is a new scenario;

[0054] S3. When the judgment result is that the current detection scenario is not a new scenario, collect the multi-spectral image data collected by the industrial camera and the real-time process parameter data;

[0055] The above S3 further includes:

[0056] When the judgment result is that the current detection scenario is a new scenario, obtain a small amount of multi-spectral image data and real-time process parameter data collected by the industrial camera corresponding to this detection scenario, and retrieve the corresponding historical detection record data;

[0057] Annotate the actual defect detection results corresponding to the multi-spectral image data and real-time process parameter data at this time to form corresponding annotation content, and associate the annotation content with the corresponding collected data;

[0058] According to the multi-spectral image data with annotation content, real-time process parameter data, and historical detection record data, execute steps S701 to S704, and output the fine-tuned visual enhancement fine-tuning large model;

[0059] After outputting the fine-tuned visual enhancement fine-tuning large model, according to the currently real-time obtained multi-spectral image data, real-time process parameter data, and historical detection record data, execute steps S5 to S7, and output the corresponding defect detection result data. At this time, the visual enhancement fine-tuning large model in S7 is the fine-tuned visual enhancement fine-tuning large model.

[0060] When outputting the fine-tuned visual enhancement fine-tuning large model, associate the fine-tuned visual enhancement fine-tuning large model with the corresponding current detection scenario to form a customized AOI scheme corresponding to the current detection scenario.

[0061] S4. Retrieve the corresponding historical detection record data from the historical database;

[0062] S5. Based on the preset feature alignment strategy, perform feature extraction and alignment on the multi-spectral image data, real-time process parameter data, and historical detection record data;

[0063] The preset feature alignment strategy is:

[0064] Input the multi-spectral image data, real-time process parameter data, and historical inspection record data into the visual encoder, process parameter encoder, and historical knowledge encoder respectively. Based on the feature extraction strategies corresponding to each encoder, extract the feature vectors corresponding to the multi-spectral image data, real-time process parameter data, and historical inspection record data respectively and align them to a unified feature space. In this embodiment, for visual feature extraction: an improved ViT-GAN architecture (Vision Transformer + Generative Adversarial Network) is adopted, and a cross-spectral attention mechanism is introduced in the image encoding stage to dynamically adjust the weights of different spectral channels through a spectral difference matrix. For example, for semiconductor wafer inspection, a feature correlation matrix is established between the 250-400nm ultraviolet spectrum and the 800-1100nm infrared spectrum. For process parameter embedding: a parameter-defect mapping table is constructed, and real-time process parameters (such as temperature, pressure, exposure time) are converted into high-dimensional vectors through Graph Embedding technology. For example, a GCN (Graph Convolutional Network) is used to model the physical constraint relationships between parameters (such as the correlation between temperature and material expansion coefficient). For historical knowledge fusion: a temporal knowledge graph embedding module is developed to convert historical defect cases (including defect types, occurrence locations, process parameter windows) into dynamic knowledge vectors. For example, the ComplEx model is used to process time series knowledge triples (defect type, occurs in, process stage).

[0065] S6. After the features extracted from the multi-spectral image data, real-time process parameter data, and historical inspection record data are aligned, based on a preset multi-modal feature fusion strategy, perform feature fusion on the aligned features and output the corresponding multi-modal feature data;

[0066] The preset multi-modal feature fusion strategy is as follows:

[0067] Concatenate the aligned feature vectors to obtain the corresponding concatenated feature vectors;

[0068] Input the concatenated feature vectors into each preset fusion function in sequence, take the output of the previous fusion function as the input and input it into the next fusion function until the output of the last fusion function is used as the finally fused multi-modal feature data;

[0069] In this embodiment, the architecture corresponding to the fusion is hierarchical fusion processing. Specifically, the goal of the hierarchical fusion architecture is to fuse visual features, process parameter features, and historical knowledge features, and through multi-level processing, extract more representative comprehensive features. This architecture includes three encoders, namely a visual encoder (visual_encoder), a process parameter encoder (parameter_encoder), and a historical knowledge encoder (knowledge_encoder), and a series of fusion layers (fusion layer processing: the concatenated feature vector fused will sequentially pass through a series of fusion layers. The fusion layers are defined by nn.ModuleList and contain 3 nn.Sequential modules with the same structure. Each nn.Sequential module includes operations such as the following linear transformation, layer normalization, activation function, and linear transformation), and gradually integrates information of different modalities in a hierarchical manner. After being processed by multiple fusion layers, a fused feature vector is finally obtained and used as the output of the hierarchical fusion architecture. This fused feature vector synthesizes information from multiple aspects such as vision, process parameters, and historical knowledge, and can be used for subsequent tasks such as industrial defect detection.

[0070] During the fusion process, according to the current corresponding detection scenario, determine the weight ratios of the respective corresponding feature vectors of the hyperspectral image data, real-time process parameter data, and historical detection record data in each fusion function. In this embodiment, an attention gating mechanism is designed to automatically adjust the contribution degrees of each modality according to the detection scenario. For example, in a high-precision detection scenario, the weight ratio of visual features is increased to 70%, the process parameter ratio is 20%, and the historical knowledge ratio is 10%.

[0071] S7. Based on the current detection scenario, retrieve the customized AOI scheme corresponding to this detection scenario from the historical database, and determine the corresponding visual enhancement fine-tuning large model. Use the output multi-modal feature data as input data and input it into the visual enhancement fine-tuning large model corresponding to this detection scenario to output the corresponding defect detection result data.

[0072] The pre-set visual enhancement fine-tuning large model in S7 is a model trained based on a pre-set training strategy;

[0073] The training strategy is as follows:

[0074] S700. Obtain historical hyperspectral image data from the historical database, as well as the corresponding historical real-time process parameter data and historical detection record data; obtain the corresponding actual defect detection data from the historical database;

[0075] S701. Execute steps S5 and S6 based on historical multi-spectral image data, as well as corresponding real-time historical process parameter data and historical inspection record data, and output corresponding historical multi-modal feature data;

[0076] S702. Construct a vision-enhanced fine-tuning large model; in this embodiment, the model architecture corresponding to the vision-enhanced fine-tuning large model is:

[0077] Vision Encoder: Use the pre-trained Qwen2-VL-7B as the basis to process image inputs and extract relevant features;

[0078] Language Text Encoder: Process text inputs to understand questions or instructions;

[0079] Cross-Attention Mechanism: Facilitate the interaction between visual and language features to achieve effective integration of multi-modal information;

[0080] Decoder: Generate model outputs based on the fused features, including the inference process and the final answer.

[0081] S703. Take the output historical multi-modal feature data as input data and input it into the constructed vision-enhanced fine-tuning large model to output corresponding training defect detection result data;

[0082] S704. According to the output training defect detection result data and the corresponding actual defect detection data, match the corresponding verifiable reward function from the database, calculate the function value corresponding to the corresponding verifiable reward function, and adjust the model parameters corresponding to the vision-enhanced fine-tuning large model according to the corresponding function value until the corresponding function value meets the preset requirements. At this time, the training of the vision-enhanced fine-tuning large model is completed. The verifiable reward function includes a visual reward function and an industrial reward function. In this embodiment, Visual-RFT is based on a reinforcement learning framework and uses a verifiable reward function to evaluate the output of the model. Specific reward functions are designed for different visual tasks. For example, in the object detection task, the IoU (Intersection over Union) reward is used to measure the overlap degree between the predicted box and the ground truth box; in the classification task, the CLS reward is used to judge the accuracy of the classification result. A policy optimization algorithm such as Group Relative Policy Optimization (GRPO) is adopted to update the model parameters by comparing the reward values of a group of candidate responses, rather than relying on a critic model. This method can more directly optimize the behavior strategy of the model and improve the learning efficiency. A hybrid training method combining iterative supervised fine-tuning (Iterative SFT) and GRPO reinforcement learning is adopted, which can dynamically adjust the length of the thought chain.

[0083] This embodiment also discloses an industrial defect intelligent detection system based on visual reinforcement fine-tuning of a large model, which uses the above-mentioned industrial defect intelligent detection method based on visual reinforcement fine-tuning of a large model.

[0084] The above are only embodiments of the present invention. Specific structures and common knowledge such as characteristics well known in the art are not described in detail here. Those of ordinary skill in the art know all the common general technical knowledge in the technical field to which the invention belongs before the filing date or the priority date, can learn all the existing technologies in this field, and have the ability to apply conventional experimental means before this date. Those of ordinary skill in the art can, under the inspiration given in this application, complete and implement this solution in combination with their own abilities. Some typical well-known structures or well-known methods should not become obstacles for those of ordinary skill in the art to implement this application. It should be noted that for those skilled in the art, without departing from the structure of the present invention, several modifications and improvements can be made, which should also be regarded as the protection scope of the present invention, and these will not affect the implementation effect of the present invention and the practicality of the patent. The protection scope required by this application should be based on the content of its claims, and the specific implementation manners described in the specification can be used to interpret the content of the claims.

Claims

1. An intelligent industrial defect detection method based on visually enhanced fine-tuning of large models, characterized in that: It includes the following steps: S1. Receive the demand data corresponding to the user, and determine the corresponding detection scenario and the industrial camera for this detection based on the demand data; S2. According to the determined detection scenario, judge whether the current detection scenario is a new scenario; S3. When the judgment result is that the current detection scenario is not a new scenario, collect the multi-spectral image data collected by the industrial camera and the corresponding real-time process parameter data; S4. Retrieve the corresponding historical detection record data from the historical database; S5. Based on the preset feature alignment strategy, perform feature extraction and alignment on the multi-spectral image data, real-time process parameter data, and historical detection record data; S6. After the features extracted from the multi-spectral image data, real-time process parameter data, and historical detection record data are aligned, based on the preset multi-modal feature fusion strategy, perform feature fusion on the aligned features and output the corresponding multi-modal feature data; S7. Based on the current detection scenario, retrieve the customized AOI scheme corresponding to this detection scenario from the historical database, determine the corresponding vision enhancement fine-tuning large model, take the output multi-modal feature data as input data, input it into the vision enhancement fine-tuning large model corresponding to this detection scenario, and output the corresponding defect detection result data.

2. The industrial defect intelligent detection method based on visually enhanced fine-tuning of a large model according to claim 1, characterized in that: The preset feature alignment strategy is: Input the multi-spectral image data, real-time process parameter data, and historical detection record data into the vision encoder, process parameter encoder, and historical knowledge encoder respectively. Based on the feature extraction strategies corresponding to each encoder, extract the feature vectors corresponding to the multi-spectral image data, real-time process parameter data, and historical detection record data respectively and align them into a unified feature space.

3. The industrial defect intelligent detection method based on visually enhanced fine-tuning of large models according to claim 2, wherein: The preset multi-modal feature fusion strategy is: Concatenate the aligned feature vectors to obtain the corresponding concatenated feature vectors; Input the concatenated feature vectors into each preset fusion function in sequence, take the output of the previous fusion function as input and input it into the next fusion function until the output of the last fusion function is used as the finally fused multi-modal feature data; During the fusion process, according to the current corresponding detection scenario, determine the weight ratios corresponding to the feature vectors of the multi-spectral image data, real-time process parameter data, and historical detection record data in each fusion function.

4. The industrial defect intelligent detection method based on visual reinforcement fine-tuning of a large model according to claim 3, characterized in that: The vision enhancement fine-tuning large model preset in S7 is a model trained based on the preset training strategy; The training strategy is: S700. Obtain historical multi-spectral image data, corresponding historical real-time process parameter data, and historical detection record data from the historical database; obtain the corresponding actual defect detection data from the historical database; S701. According to the historical multi-spectral image data, corresponding historical real-time process parameter data, and historical detection record data, execute steps S5 and S6 and output the corresponding historical multi-modal feature data; S702. Construct a vision enhancement fine-tuning large model; S703. Use the output historical multi-modal feature data as input data, and input it into the constructed visual reinforcement fine-tuning large model to output the corresponding training defect detection result data; S704. According to the output training defect detection result data and the corresponding actual defect detection data, match the corresponding verifiable reward function from the database, calculate the function value corresponding to the corresponding verifiable reward function, and adjust the model parameters corresponding to the visual reinforcement fine-tuning large model according to the corresponding function value until the corresponding function value meets the preset requirements. At this time, the training of the visual reinforcement fine-tuning large model is completed. The verifiable reward function includes a visual reward function and an industrial reward function.

5. The industrial defect intelligent detection method based on visually enhanced fine-tuning of a large model according to claim 4, characterized in that: S3 further includes: When the judgment result is that the current detection scene is a new scene, obtain a small amount of multi-spectral image data and real-time process parameter data collected by the industrial camera corresponding to the detection scene, and retrieve the corresponding historical detection record data; Annotate the actual defect detection results corresponding to the multi-spectral image data and real-time process parameter data at this time to form corresponding annotation content, and associate the annotation content with the corresponding collected data; According to the multi-spectral image data with annotation content, real-time process parameter data, and historical detection record data, execute steps S701 to S704, and output the fine-tuned visual reinforcement fine-tuning large model; After outputting the fine-tuned visual reinforcement fine-tuning large model, according to the multi-spectral image data, real-time process parameter data, and historical detection record data obtained in real time currently, execute steps S5 to S7, and output the corresponding defect detection result data. At this time, the visual reinforcement fine-tuning large model in S7 is the fine-tuned visual reinforcement fine-tuning large model.

6. The industrial defect intelligent detection method based on visually enhanced fine-tuning of a large model according to claim 5, characterized in that: S3 further includes: When outputting the fine-tuned visual reinforcement fine-tuning large model, associate the fine-tuned visual reinforcement fine-tuning large model with the corresponding current detection scene to form a customized AOI solution corresponding to the current detection scene.

7. An intelligent industrial defect detection system based on vision-enhanced fine-tuning of large models, characterized in that: Use the industrial defect intelligent detection method based on the visual reinforcement fine-tuning large model according to any one of claims 1 to 6 above.

Citation Information

Cited By

  • Multi-mode AI collaborative industrial visual defect detection system

    CN121708018A