Large model method and device for AVI detection defect data judgment and medium

By constructing a large model AVIGPT that combines image features and language prompts, the problems of low detection rate, high false alarm rate and poor generalization ability of traditional detection methods in defect detection are solved, and defect detection with high accuracy and low false alarm rate in small sample scenarios are achieved, with good migration and scalability.

CN120510107APending Publication Date: 2025-08-19JIANGSU PROVISION ELECTRONICS CO LTD

Patent Information

Application Number
CN202510579426.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

The existing industrial vision detection methods have low detection rate, high false alarm rate, poor generalization ability in defect detection, and rely on a large number of manual thresholds to be set, making it difficult to migrate to new product lines or defect types, and lack semantic level intelligent reasoning and judgment capabilities.

Method used

AVIGPT, a defect detection big model is constructed, combined with image features and language prompt information, and a multimodal training big model is used to train a large model, including frozen image encoder, text encoder, decoder, prompt word encoder and LLM model. Through multimodal alignment and semantic level determination, accurate re-judgment of AVI detection defect data is achieved.

Benefits of technology

The large-model method gets rid of threshold dependence, has good mobility and scalability, and can achieve high-accurate defect detection in small sample scenarios, significantly improves the detection rate and reduces the false alarm rate, and enhances the perception of complex patterns and tiny anomalies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120510107A_ABST
    Figure CN120510107A_ABST
Patent Text Reader

Abstract

The invention discloses a large model method and device for AVI detection defect data judgment and a medium. The large model method comprises the steps that a defect detection large model is constructed; the method comprises the following steps: acquiring an industrial product photo, generating an image text data pair after text labeling, inputting the image text data pair into a defect detection large model, generating intermediate feature representation by matching an image encoder and a text encoder, and generating a prediction mask by a decoder based on the intermediate feature representation; performing multi-modal alignment on the predicted text mask and the text feature information to obtain modal feature representation; inputting the modal feature representation and the predicted image mask into a cue word encoder to generate a Prompt vector code with the same dimension as the image feature information; and inputting the Prompt vector code into the LLM model, and outputting a semantic level judgment result after language questioning. The large model method has good mobility, expansibility, strong generalization ability and perception ability, gets rid of threshold dependence, and realizes accurate redetermination of AVI detection defect data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision and artificial intelligence technology, and in particular to a large model method, device and medium for AVI detection defect data determination. Background Art

[0002] Currently, industrial visual inspection, especially defect detection in integrated circuit manufacturing, generally uses traditional image processing and machine learning methods. These inspection methods often rely on a large number of manually set threshold parameters, and feature extraction is mainly based on low-level visual information such as edges and textures. This leads to two key problems:

[0003] 1) Low detection rate and high false alarm rate: Due to the complexity and morphology of defects, traditional detection methods have difficulty in dealing with high-dimensional and diverse data, especially when dealing with small or complex defects;

[0004] 2) Poor generalization capability: Existing detection methods are mostly dedicated algorithms that are dependent on specific patterns and defects and are difficult to migrate to new product lines or different defect types.

[0005] Furthermore, while existing detection methods have improved detection accuracy, they still require a large number of labeled samples to train the classifier and have limited generalization capabilities for defects outside the training dataset. Consequently, existing detection methods still suffer from the following technical issues: They cannot set thresholds without extensive manual experience; they are ineffective with small numbers of samples; and they lack semantic-level intelligent reasoning and judgment capabilities.

[0006] In view of this, the present invention is proposed. Summary of the Invention

[0007] In order to overcome the above-mentioned defects, the present invention provides a large-model method, device, and medium for AVI detection defect data judgment. The large-model method combines image features and language prompt information, and introduces multimodal training large-model capabilities. It has good migration, scalability, strong generalization ability, and the ability to perceive complex patterns and minor anomalies, and is completely free from threshold dependence, thereby enabling accurate re-judgment of AVI detection defect data, greatly improving the detection accuracy, and having significant advantages, especially in small-sample defect scenarios.

[0008] The technical solution adopted by the present invention to solve the technical problem is: a large model method for AVI detection defect data determination, comprising:

[0009] Build a large defect detection model AVIGPT, which integrates a frozen image encoder, a frozen text encoder, a decoder, a prompt word encoder, and an LLM model;

[0010] Obtain photos of industrial products, annotate them with text, and generate image-text data pairs;

[0011] The obtained image-text data pair is input into the defect detection large model AVIGPT, the image encoder and the text encoder cooperate to generate an intermediate feature representation containing image feature information and text feature information, and the decoder generates a prediction mask containing a predicted image mask and a predicted text mask based on the intermediate feature representation;

[0012] Perform multimodal alignment on the obtained predicted text mask and the obtained text feature information to obtain modal feature representation;

[0013] Inputting the obtained modal feature representation and the obtained predicted image mask into the prompt word encoder, and generating a prompt vector encoding with the same dimension as the obtained image feature information after convolution and fully connected network mapping processing;

[0014] The obtained Prompt vector code is input into the LLM model, and after a language question is input into the LLM model, the LLM model outputs a semantic level judgment result.

[0015] As a further improvement of the present invention, it also includes: training the large defect detection model AVIGPT.

[0016] As a further improvement of the present invention, the method for training the large defect detection model AVIGPT includes:

[0017] After obtaining several industrial product photos and annotating them with text, a dataset consisting of several image-text data pairs A is generated; at the same time, the dataset is divided into a training set, a test set, and a validation set;

[0018] The dataset is input into the large defect detection model AVIGPT, the image encoder and the text encoder cooperate to generate a plurality of intermediate feature representations A each containing image feature information A and text feature information A, and the decoder generates a prediction mask A containing a predicted image mask A and a predicted text mask A based on the intermediate feature representations A;

[0019] Performing multimodal alignment on the obtained predicted text mask A and the obtained text feature information A to obtain a modal feature representation A; simultaneously calculating a cross entropy loss A between the obtained predicted text mask A and the obtained text feature information A, and correcting the weight distribution of the cross entropy loss A using a FocalLoss loss function to minimize the cross entropy loss A, thereby obtaining the optimized decoder;

[0020] The obtained predicted image mask A is input into the prompt word encoder, and after convolution and fully connected network mapping processing, a Prompt vector encoding A with the same dimension as the obtained image feature information A is generated; the obtained Prompt vector encoding A and the obtained modal feature representation A are input into the LLM model to generate a text response; the cross entropy loss B between the obtained text response and the obtained text feature information A is calculated, and the weight gradient of the cross entropy loss B relative to the prompt word encoder is adjusted using an optimization algorithm to minimize the cross entropy loss B, thereby updating and optimizing the parameters of the prompt word encoder.

[0021] As a further improvement of the present invention, the content of the text annotation of the industrial product photo includes: a defect abnormality at a certain position of the product, or a non-defect abnormality of the product.

[0022] As a further improvement of the present invention, a gradient descent method or a stochastic gradient descent method is used to adjust the weight gradient of the cross entropy loss B relative to the prompt word encoder.

[0023] As a further improvement of the present invention, the formula for calculating the cross entropy loss A and the cross entropy loss B is:

[0024]

[0025] The Focal Loss loss function formula is:

[0026]

[0027] As a further improvement of the present invention, the language questions input into the LLM model include: whether this product drawing has defects or abnormalities;

[0028] The semantic level determination result output by the LLM model includes: a defect abnormality at a certain position of the product, or a non-defect abnormality of the product.

[0029] The present invention also provides a device for determining defective data for AVI detection, which includes a processor capable of implementing the steps of the large model method for determining defective data for AVI detection described in the present invention.

[0030] The present invention also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the large model method for determining AVI detection defect data of the present invention are implemented.

[0031] The beneficial effects of the present invention are as follows: compared with traditional defect detection methods, the large model method provided by the present invention has the following advantages: ① The large model method is based on semantic large model reasoning and does not require threshold setting, that is, it is completely free from dependence on the threshold; ② The large model method relies on the strong generalization ability of the large model, and even if the defect samples are very few, high-accuracy judgment can still be completed; ③ The large model method combines image features and language prompt information through multimodal reasoning, and can accurately output defect anomalies; ④ The large model method has good portability and scalability, and there is no need to redesign the model structure for each new defect or image. It only needs to replace the prompt or a few samples to be applicable to a variety of application scenarios; ⑤ The large model method significantly enhances the model's perception of complex patterns and minor anomalies by fusing images and texts in a multimodal manner, thereby having a higher detection rate and a lower false alarm rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 This is a flow chart of the large model method for determining AVI detection defect data described in Example 1 of the present invention;

[0033] Figure 2 This is a block diagram of the basic architecture of the large defect detection model AVIGPT described in Example 1 of the present invention.

[0034] The following description is made with reference to the accompanying drawings:

[0035] 1. Image encoder; 2. Text encoder; 3. Decoder; 4. Prompt word encoder; 5. LLM model; 6. Multimodal alignment module. DETAILED DESCRIPTION

[0036] The preferred embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0037] Example 1:

[0038] Please see the attached Figure 1 As shown, this embodiment 1 provides a large model method for AVI detection defect data determination, including the following processing steps:

[0039] S1: Build a large defect detection model AVIGPT.

[0040] The defect detection model AVIGPT adopts a multimodal large model architecture that combines the capabilities of natural language processing (NLP) and computer vision (CV) to achieve comprehensive understanding and analysis of multimodal information, thereby enabling a more comprehensive understanding and processing of complex data. Figure 2As shown, the defect detection large model AVIGPT includes a frozen image encoder 1, a frozen text encoder 2, a decoder 3, a prompt word encoder 4 and a frozen LLM model 5 (large language model), wherein the image encoder 1, the text encoder 2 and the LLM model 5 are all frozen to minimize the computational cost; the decoder 3 and the prompt word encoder 4 need to be fine-tuned.

[0041] Further explanation: ① Frozen image encoder 1 and text encoder 2: This means they only perform forward propagation during training to extract image features and / or generate text features, without updating their parameters. This significantly reduces computing resources, shortens training time and costs, and improves model generalization. ② Frozen LLM model 5: This means its pre-trained parameters are not updated during training, and only forward propagation is used to extract text features.

[0042] S2: Obtain photos of industrial products, perform text annotation, and generate image-text data pairs.

[0043] In this application, the industrial products (such as PCBs) whose images are captured have already undergone automated optical inspection (AVI). Therefore, after acquiring photos of the industrial products through a computer vision system, based on the inspection results of the AVI inspection equipment, appropriate text annotations can be made on the photos using software such as Adobe Photoshop or a universal image editor. Specifically, the text annotations may include information about defects or abnormalities in a certain location on the product, or information about the product's normal condition.

[0044] It can be understood that the present application obtains the image data and text data of industrial products by combining a computer vision system with the above-mentioned text annotation software, that is, obtains the image-text data pair.

[0045] S3: The obtained image-text data pair is input into the defect detection large model AVIGPT, and the image encoder 1 and the text encoder 2 cooperate to generate an intermediate feature representation containing image feature information and text feature information, and the decoder 3 generates a prediction mask containing a predicted image mask and a predicted text mask based on the intermediate feature representation.

[0046] Specifically, the image feature information is obtained after the image encoder 1 processes the image data in the image-text data pair; and the text feature information is obtained after the text encoder 2 processes the text data in the image-text data pair.

[0047] The intermediate feature representation is the basis for the decoder 3 to perform prediction. The decoder 3 receives the intermediate feature representation as input and generates a prediction mask with the same size as the input information.

[0048] S4: Perform multimodal alignment on the obtained predicted text mask and the obtained text feature information to obtain a modal feature representation.

[0049] Specifically, the multimodal alignment refers to the process of establishing associations and synchronization between data from different modalities at the temporal, spatial, or semantic levels to achieve unified representation and collaborative reasoning of cross-modal information. In the present application, the multimodal alignment method preferably adopts an alignment method based on deep learning. It can be understood that the defect detection large model AVIGPT described in the present application also integrates a multimodal alignment module 6 for multimodally aligning the obtained predicted text mask with the obtained text feature information.

[0050] S5: The obtained modal feature representation and the obtained predicted image mask are input into the prompt word encoder 4, and after convolution and fully connected network mapping processing, a Prompt vector encoding with the same dimension as the obtained image feature information is generated.

[0051] Specifically, the Prompt vector encoding is a combination of two vectors, where one vector corresponds to the obtained modality feature representation and the other vector corresponds to the obtained predicted image mask.

[0052] S6: Input the obtained Prompt vector code into the LLM model 5, and after inputting a language question into the LLM model 5, such as: input "Is there a defect abnormality in this product drawing?"; the LLM model 5 outputs a semantic-level judgment result based on the obtained Prompt vector code, such as: output "There is a defect abnormality in a certain position of the product, or the product has no defect abnormality."

[0053] From the above, it can be seen that compared with traditional defect detection methods, the large model method provided in this application is an accurate re-determination of AVI detection defect data, which has the following advantages: ① The large model method is based on semantic large model reasoning and does not require threshold setting, that is, it is completely free from dependence on the threshold; ② The large model method relies on the strong generalization ability of the large model, and can still complete high-accuracy judgment even if there are very few defect samples; ③ The large model method combines image features and language prompt information through multimodal reasoning, and can accurately output defect anomalies; ④ The large model method has good portability and scalability, and there is no need to redesign the model structure for each new defect or image. It only needs to replace the prompt or a few samples to be applicable to a variety of application scenarios; ⑤ The large model method significantly enhances the model's perception of complex patterns and subtle anomalies by fusing images and texts multimodally, thereby having a higher detection rate and a lower false alarm rate.

[0054] In addition, in this application, in order to achieve more accurate judgment by the large model method, this application also performs multiple iterative training on the defect detection large model AVIGPT, and the specific training method is:

[0055] S11: After obtaining several photos of industrial products (after automatic optical inspection) and annotating them with text, a dataset consisting of several image-text data pairs A is generated; at the same time, the dataset is divided into a training set, a test set, and a validation set.

[0056] Specifically, similar to the above S2, this S11 also obtains the image-text data pair A by combining a computer vision system with text annotation software, and the image-text data pair A is a combination of image data and text data of an industrial product.

[0057] In addition, this application does not impose any restrictions on the specific number and division ratio of the training set, test set, and validation set, which are determined based on processing requirements.

[0058] S12: The data set is input into the defect detection large model AVIGPT, and the image encoder 1 and the text encoder 2 cooperate to generate several intermediate feature representations A each containing image feature information A and text feature information A, and the decoder 3 generates a prediction mask A containing a predicted image mask A and a predicted text mask A based on the intermediate feature representation A.

[0059] Specifically, regarding the intermediate feature representation A and the prediction mask A, reference may be made to the intermediate feature representation and the prediction mask in S3 above, and therefore details will not be repeated here.

[0060] S13: Perform multimodal alignment on the obtained predicted text mask A and the obtained text feature information A to obtain the modal feature representation A; at the same time, calculate the cross-entropy loss A (Cross-Entropy Loss) between the obtained predicted text mask A and the obtained text feature information A, and use the FocalLoss loss function (also known as the focal loss function) to correct the weight distribution of the cross-entropy loss A to minimize the cross-entropy loss A, thereby obtaining the optimized decoder 3, and then obtaining the optimized defect detection large model AVIGPT.

[0061] Specifically, regarding the modal feature representation A, reference may be made to the modal feature representation described in S4 above, so it will not be elaborated here.

[0062] The calculation formula used in this application to calculate the cross entropy loss A is:

[0063]

[0064] Where, L image→text The name of the loss function for the image-to-text task, which is used to quantify the error between the model prediction and the true label; logits is the unnormalized score output by the model in the last layer (usually the fully connected layer); S i,: Represents the model output of the i-th sample, usually a vector; label=i represents the true label of the i-th sample.

[0065] As can be understood, during training, the cross-entropy loss function is used to measure the difference between the text descriptions generated by the defect detection model AVIGPT and the actual text descriptions. By minimizing this cross-entropy loss function, the defect detection model AVIGPT can continuously adjust its parameters, thereby making the generated text descriptions increasingly close to the actual text descriptions.

[0066] The formula of the Focal Loss loss function is:

[0067]

[0068] Where, L focal The name of the focal loss function, which is used to measure the difference between the model prediction result and the true label; (1-p i ) γ is a modulation factor, where p i It represents the probability that the model predicts the i-th sample as a positive class. γ is a hyperparameter and is usually greater than 0.

[0069] It can be understood that the focus loss function is an improved cross entropy loss function. By assigning different weights to different samples, the model pays more attention to those difficult-to-classify samples during training, thereby significantly improving the detection accuracy and recall rate of the model.

[0070] S14: The obtained predicted image mask A is input into the prompt word encoder 4, and after convolution and fully connected network mapping processing, a Prompt vector encoding A with the same dimension as the obtained image feature information A is generated; the obtained Prompt vector encoding A and the obtained modal feature representation A are input into the LLM model 5 to generate a text response; the cross entropy loss B between the obtained text response and the obtained text feature information A is calculated, and the weight gradient of the cross entropy loss B relative to the prompt word encoder 4 is adjusted using an optimization algorithm to minimize the cross entropy loss B, thereby updating and optimizing the parameters of the prompt word encoder 4, thereby obtaining the optimized defect detection large model AVIGPT.

[0071] Specifically, the calculation formula used in this application to calculate the cross entropy loss B is: The meaning of the formula is explained above and will not be repeated here.

[0072] This application uses the gradient descent method or the stochastic gradient descent method to adjust the weight gradient of the cross entropy loss B relative to the prompt word encoder. The gradient descent method and the stochastic gradient descent method are both commonly used gradient optimization algorithms, so they are not described in detail here.

[0073] In addition, it is understandable that the defect detection large model AVIGPT described in the present application also integrates an operation module for performing the calculations of the cross entropy loss A and the cross entropy loss B, and optimizing the weights of the cross entropy loss A and the cross entropy loss B.

[0074] From the above, it can be seen that in this application, the defect detection large model AVIGPT adopts a step-by-step training strategy, namely: first train the decoder 3, and then train the prompt word encoder 4; the main purpose is to first align the text space and the image space, integrate the text information into the output of the decoder 3, provide pixel-level information for the prompt word encoder 4, and then promote the prompt word encoder 4 to generate higher quality text, so that the output results of the defect detection large model AVIGPT are more accurate, more efficient, more generalized, etc.

[0075] Example 2:

[0076] This embodiment 2 provides a device for determining defect data for AVI detection, which includes a processor capable of implementing the steps of the large model method for determining defect data for AVI detection described in the above embodiment 1.

[0077] Specifically, the processor may include one or more processing cores. The processor utilizes various interfaces and circuits to connect various components within the entire device, and executes instructions, programs, code sets, or instruction sets stored in memory, as well as accesses data stored in memory, thereby performing various functions of the device and processing data.

[0078] Furthermore, the processor may be implemented in at least one of the following hardware forms: a digital signal processor (DSP), a field programmable gate array (FPGA), or a programmable logic array (PLA). Alternatively, the processor may integrate one or a combination of a central processing unit (CPU) and a modem.

[0079] In addition, since the device provided in this embodiment 2 corresponds to the large model method provided in the above embodiment 1, and the principle of solving the problem by the device is similar to that of the large model method, the implementation of the device can refer to the implementation process in the above embodiment 1, so it will not be repeated here.

[0080] Example 3:

[0081] This embodiment 3 provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the large model method for AVI detection defect data determination described in the above embodiment 1 are implemented.

[0082] Specifically, the computer program may be stored in a computer-readable storage medium, and the storage medium may include read-only memory ROM, random access memory RAM, programmable read-only memory PROM, erasable programmable read-only memory EPROM, one-time programmable read-only memory OTPROM, electronically erasable rewritable read-only memory EEPROM, read-only compact disc CD-ROM or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.

[0083] In addition, since the storage medium provided in this embodiment 3 corresponds to the large model method provided in the above embodiment 1, and the principle of solving the problem by the storage medium is similar to that of the large model method, the implementation of the storage medium can refer to the implementation process in the above embodiment 1, so it will not be repeated here.

[0084] Finally, the suffixes "A", "B", etc. in the component names in this application description (such as cross entropy loss A, cross entropy loss B, etc.) are only for the convenience of description and are not used to limit the scope of implementation of the patent of this invention.

[0085] In the above description, many specific details are set forth in order to fully understand the present invention. However, the above description is only a preferred embodiment of the present invention. The present invention can be implemented in many other ways different from those described herein, so the present invention is not limited to the specific implementation disclosed above. At the same time, any person skilled in the art can make many possible changes and modifications to the technical solution of the present invention using the methods and technical contents disclosed above without departing from the scope of the technical solution of the present invention, or modify it into an equivalent embodiment of equivalent changes. Any simple modification, equivalent change and modification made to the above embodiment based on the technical essence of the present invention without departing from the content of the technical solution of the present invention still falls within the scope of protection of the technical solution of the present invention.

Claims

1. A large-scale model method for determining AVI inspection defect data, characterized by: include: Build a large defect detection model AVIGPT, which integrates a frozen image encoder, a frozen text encoder, a decoder, a prompt word encoder, and an LLM model; Obtain photos of industrial products, annotate them with text, and generate image-text data pairs; The obtained image-text data pair is input into the defect detection large model AVIGPT, the image encoder and the text encoder cooperate to generate an intermediate feature representation containing image feature information and text feature information, and the decoder generates a prediction mask containing a predicted image mask and a predicted text mask based on the intermediate feature representation; Perform multimodal alignment on the obtained predicted text mask and the obtained text feature information to obtain modal feature representation; Inputting the obtained modal feature representation and the obtained predicted image mask into the prompt word encoder, and generating a prompt vector encoding with the same dimension as the obtained image feature information after convolution and fully connected network mapping processing; The obtained Prompt vector code is input into the LLM model, and after a language question is input into the LLM model, the LLM model outputs a semantic level judgment result.

2. The large-scale model method for determining AVI detection defect data according to claim 1 is characterized in that: Also includes: The defect detection large model AVIGPT is trained.

3. The large-scale model method for determining AVI detection defect data according to claim 2, characterized in that: The method for training the defect detection large model AVIGPT includes: After obtaining several photos of industrial products and annotating them with text, a dataset consisting of several image-text data pairs A is generated; at the same time, the dataset is divided into a training set, a test set, and a validation set; The dataset is input into the large defect detection model AVIGPT, the image encoder and the text encoder cooperate to generate a plurality of intermediate feature representations A each containing image feature information A and text feature information A, and the decoder generates a prediction mask A containing a predicted image mask A and a predicted text mask A based on the intermediate feature representations A; Performing multimodal alignment on the obtained predicted text mask A and the obtained text feature information A to obtain a modal feature representation A; simultaneously calculating a cross entropy loss A between the obtained predicted text mask A and the obtained text feature information A, and correcting the weight distribution of the cross entropy loss A using a FocalLoss loss function to minimize the cross entropy loss A, thereby obtaining the optimized decoder; The obtained predicted image mask A is input into the prompt word encoder, and after convolution and fully connected network mapping processing, a Prompt vector encoding A with the same dimension as the obtained image feature information A is generated; the obtained Prompt vector encoding A and the obtained modal feature representation A are input into the LLM model to generate a text response; the cross entropy loss B between the obtained text response and the obtained text feature information A is calculated, and the weight gradient of the cross entropy loss B relative to the prompt word encoder is adjusted using an optimization algorithm to minimize the cross entropy loss B, thereby updating and optimizing the parameters of the prompt word encoder.

4. The large-scale model method for determining AVI detection defect data according to claim 3 is characterized in that: The text annotation content of the industrial product photo includes: a defect or abnormality at a certain position of the product, or a non-defective or abnormal product.

5. The large-scale model method for determining AVI detection defect data according to claim 3 is characterized in that: The gradient descent method or the stochastic gradient descent method is used to adjust the weight gradient of the cross entropy loss B relative to the prompt word encoder.

6. The large-scale model method for determining AVI inspection defect data according to claim 3 is characterized in that: The formula for calculating the cross entropy loss A and the cross entropy loss B is: The Focal Loss loss function formula is:

7. The large model method for AVI inspection defect data determination according to claim 1 is characterized in that: The language questions input into the LLM model include: whether this product drawing has defects or abnormalities; The semantic level determination result output by the LLM model includes: a defect abnormality at a certain position of the product, or a non-defect abnormality of the product.

8. A device for determining defect data for AVI detection, characterized by: The method comprises a processor capable of implementing the steps of the large model method for AVI detection defect data determination according to any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the large model method for AVI detection defect data determination according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Controllable defect image generation method and equipment based on visual semantic fusion

    CN118097318A

  • Visual segmentation method, device and equipment for defects in welding radiograph and medium

    CN118397284A

  • Model training method and device, industrial image defect detection method and device and storage medium

    CN119068305A

  • Multi-source pavement disease identification method based on language and image large model

    CN119107447A

  • Multi-modal large model vehicle loss assessment method and system based on cue word guidance

    CN119722578A

Cited By

  • Multi-mode industrial defect detection method, device and equipment and storage medium

    CN121437425A