Quality inspection model pre-training system and pre-training method based on data and knowledge fusion

By adopting a pre-training system that integrates data and knowledge in the industrial quality inspection model, the problem of insufficient representation ability of industrial quality inspection model is solved, efficient modeling of a small number of defect samples and accuracy improvement in defect detection, and shortening the quality inspection project cycle.

CN120067687APending Publication Date: 2025-05-30TZTEK TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510139068.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The industrial quality inspection model has shortcomings in its representation capabilities, especially when processing industrial production data and general field data, there are problems such as domain deviation and insufficient labeling data.

Method used

A quality inspection model pre-training system based on the fusion of data and knowledge is adopted, and the multi-task loss function calculation module is calculated through the knowledge graph construction module, the visual encoder module, the text encoder module, the feature fusion module and the loss function calculation module, combined with the masked word prediction loss, the Q&A instruction learning loss and the target cluster coding prediction loss, the calculation of the multi-task loss function is realized.

Benefits of technology

It improves the ability to model a small number of defect samples in actual industrial quality inspection, improves the accuracy of defect detection, shortens the cycle of industrial quality inspection projects, and facilitates promotion and application in the field of industrial quality inspection that requires sample training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067687A_ABST
    Figure CN120067687A_ABST
Patent Text Reader

Abstract

The invention provides a quality inspection model pre-training system and pre-training method based on data and knowledge fusion, and belongs to the field of artificial intelligence and industrial quality inspection.The quality inspection model pre-training system comprises a knowledge graph construction module, a visual encoder module, a text encoder module, a feature fusion module and a loss function calculation module, according to the scheme, related technologies such as data and knowledge fusion and large model pre-training are adopted, and the problem that the industrial quality inspection model is insufficient in representation capability is solved. According to the method, the capability of modeling a small number of flaw samples in actual industrial quality inspection is improved, the flaw detection accuracy is improved, the industrial quality inspection project period is shortened, and the method is convenient to popularize and apply in the industrial quality inspection field needing sample training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of artificial intelligence and industrial quality inspection, and particularly relates to a quality inspection model pre-training system and a pre-training method based on data and knowledge fusion. Background Art

[0002] In the field of artificial intelligence, the relationship between data and knowledge is interactive. Data can be abstracted as the extension of knowledge, and knowledge can be concretized as the connotation of data. When humans perform visual recognition, they not only analyze the data transmitted by the retina into short-term memory, but also utilize the mental images in long-term memory, that is, visual knowledge.

[0003] Pre-trained large models are an efficient method for knowledge integration, with strong data induction capabilities, and can extract some in-depth knowledge from data. In recent years, remarkable progress has been made in data modeling such as language, speech, and images.

[0004] However, there are certain domain biases between industrial production data and general domain data, and the amount of effectively labeled data is relatively small, which leads to insufficient preconditions for directly pre-training large models.

[0005] As an important method for knowledge representation, knowledge graphs have been widely used in various fields. In the field of industrial quality inspection, knowledge graphs generally describe inspection rules, and there are problems such as single representation form, incomplete information, and inconsistent standards. How to integrate industrial knowledge graphs with large model technologies is an important research direction for improving the representation ability of industrial quality inspection large models. Summary of the Invention

[0006] In order to overcome the deficiencies of the prior art, the purpose of the present invention is to provide a quality inspection model pre-training system and a pre-training method based on data and knowledge fusion, which can solve the problem of insufficient representation ability of industrial quality inspection models.

[0007] Design principle: The related technologies of data and knowledge fusion and large model pre-training are adopted. One is the construction method of knowledge graphs for industrial quality inspection; the second is the data and knowledge fusion method; the third is the design and use of the indicator learning and target clustering coding loss functions. The design scheme is as follows.

[0008] A quality inspection model pre-training system based on data and knowledge fusion, the quality inspection model pre-training system includes: a knowledge graph construction module for constructing an industrial quality inspection knowledge graph containing defect types and the defining relationships between defects; a visual encoder module for extracting visual features of industrial product images; a text encoder module for processing text descriptions related to defects to generate text features; a feature fusion module for fusing visual features and text features; and a loss function calculation module for calculating a multi-task loss function including masked word prediction loss, question-answer indication learning loss, and target clustering encoding prediction loss.

[0009] Preferably, the quality inspection model pre-training system further includes a model training and fine-tuning module for freezing the weights of the visual encoder during training and updating the weights of the text encoder using the LoRA fine-tuning technique.

[0010] Preferably, the knowledge graph construction module includes: an attribute knowledge encoding unit for text-encoding the texture attributes, shape attributes, position attributes, and color attributes of defect types; and an association relationship fusion unit for fusing the association relationships of different targets using a cross-attention mechanism.

[0011] Preferably, the texture attributes include material, production process, and surface treatment; the shape attributes include area, length, size, thickness, straightness or curvature.

[0012] Preferably, the feature fusion module fuses visual features and text features based on a linear layer and inputs them into a decoder for decoding.

[0013] Preferably, the loss function calculation module includes: a masked word prediction loss unit; a question-answer indication learning loss unit; and a target clustering encoding prediction loss unit for clustering and encoding the text features, predicting the clustering encoding of the visual features, and taking the distance between the predicted encoded features as the loss value.

[0014] Preferably, the question-answer indication learning loss unit uses text templates including what, where, how big, what color, what does it look like, what treatment is used, what attributes does it have, what is the associated target, etc., and generates multi-turn question-and-answer dialogue data for each image based on ChatGPT to form an interaction sequence of images and texts.

[0015] The present invention also provides a pre-training method for a quality inspection model based on the aforementioned quality inspection model pre-training system. The pre-training method includes the following steps. Obtain images, defect targets, and attribute knowledge; convert the defect targets and their related descriptive knowledge into a form that can be processed by a text encoder, and design a text template to obtain image question-and-answer instruction data; the image passes through a visual encoder to obtain the visual features of the product image, and the visual features and the text features of the text encoder are used as embedding features and input into the Llama decoder; use masked word prediction loss, question-and-answer instruction learning loss, and target clustering encoding prediction loss to calculate the loss function.

[0016] Preferably, during training, freeze the weights of the visual encoder and use the LoRA fine-tuning technique to update the weights of the text encoder.

[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: This application improves the ability to model a small number of defect samples in actual industrial quality inspection and improves the accuracy of defect detection, shortens the industrial quality inspection project cycle, and is convenient for popularization and application in the industrial quality inspection field that requires sample training. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is a schematic diagram of the overall process of online re-verification of the large model of the quality inspection model pre-training system based on data and knowledge fusion of the present invention; Figure 2 It is a schematic diagram of the cross-attention mechanism of associated targets. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0020] Quality inspection model pre-training system A quality inspection model pre-training system based on data and knowledge fusion, see Figure 1 and Figure 2 , the quality inspection model pre-training system includes: a knowledge graph construction module, a visual encoder module, a text encoder module, a feature fusion module, a loss function calculation module, and a model training and fine-tuning module. In a specific example, the industrial quality inspection large model adopts the LLaVA architecture, which is introduced in detail below.

[0021] The knowledge graph construction module is used to construct an industrial quality inspection knowledge graph containing defect types and the delimiting relationships between defects.

[0022] A visual encoder module for extracting visual features of industrial product images.

[0023] A text encoder module for processing text descriptions related to defects and generating text features.

[0024] A feature fusion module for fusing visual features and text features.

[0025] A loss function calculation module for calculating a multi-task loss function including masked word prediction loss, question-answer indication learning loss, and target clustering encoding prediction loss.

[0026] A model training and fine-tuning module for freezing the weights of the visual encoder during training and updating the weights of the text encoder using the LoRA fine-tuning technique.

[0027] Among them, the knowledge graph construction module includes: an attribute knowledge encoding unit and a correlation relationship fusion unit. Specifically, the attribute knowledge encoding unit is used for text encoding of the texture attributes, shape attributes, position attributes, and color attributes of the defect types. The correlation relationship fusion unit is used for fusing the correlation relationships of different targets using a cross-attention mechanism.

[0028] These correlation relationships include the similarity and subordination relationships between defects, and the relationship between defects is defined through this unit.

[0029] Among them, the texture attributes include material, production process, and surface treatment; the shape attributes include area, length, size, thickness, curvature. For these attribute knowledge, they are directly text-encoded; for the correlation relationships, different targets are fused using a cross-attention mechanism, and the cross-attention model structure is as Figure 2 shown. The cross-attention mechanism includes the current target feature (providing feature q), the associated target feature (providing features k and v), and the three features are fused based on cross-attention in a linear layer and a unified encoded feature is output based on the connection.

[0030] Among them, the feature fusion module fuses visual features and text features based on a linear layer and inputs them into a decoder for decoding.

[0031] Among them, the loss function calculation module includes: a masked word prediction loss unit, a question-answer indication learning loss unit, and a target clustering encoding prediction loss unit. Among them, the target clustering encoding prediction loss unit is used for clustering and encoding the text features and predicting the clustering encoding of the visual features, and taking the distance between the predicted encoded features as the loss value.

[0032] Among them, the Q&A instruction learning loss unit adopts text templates including what, where, how big, what color, what it looks like, what processing is adopted, what attributes it has, and what the associated target is, and generates multi-round question-and-answer dialogue data for each image based on ChatGPT to form an interaction sequence of images and texts.

[0033] Through the above construction method, it is ensured that image and text data are input into the model in a unified format, guaranteeing the effective fusion of multi-modal data.

[0034] Quality inspection model pre-training method A quality inspection model pre-training method based on the aforementioned quality inspection model pre-training system, the pre-training method includes the following steps.

[0035] First, obtain images, defect targets, and attribute knowledge.

[0036] Secondly, convert the defect target and its related description knowledge into a form that can be processed by the text encoder, and design a text template to obtain image Q&A instruction data; Thirdly, the image obtains the visual features of the product image through the visual encoder, and the visual features and the text features of the text encoder are used as embedding features and input into the Llama decoder.

[0037] Finally, use the masked word prediction loss, Q&A instruction learning loss, and target clustering coding prediction loss to calculate the loss function.

[0038] It should be noted that during training, the weights of the visual encoder are frozen, and the LoRA fine-tuning technology is used to update the weights of the text encoder.

[0039] Through the above system and method, the small sample modeling ability for downstream applications is improved, and the average demand for real defect samples in industrial quality inspection projects is reduced by about 80%.

[0040] Storage medium A storage medium, on which a computer program is stored. When the computer program is executed by a processor, the aforementioned quality inspection large model pre-training method is implemented.

[0041] Electronic device An electronic device includes a processor and a memory. A computer program is stored in the memory. When the computer program is executed by the processor, the aforementioned quality inspection large model pre-training method is implemented.

[0042] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A quality inspection model pre-training system based on data and knowledge fusion, characterized in that: The quality inspection model pre-training system includes: A knowledge graph construction module is used to construct an industrial quality inspection knowledge graph that includes defect types and the defining relationships between defects; A visual encoder module to extract visual features of industrial product images; The text encoder module is used to process the text description related to the defect and generate text features; Feature fusion module, used to fuse visual features and text features; The loss function calculation module is used to calculate multi-task loss functions including masked word prediction loss, question-answer instruction learning loss, and target cluster encoding prediction loss.

2. The quality inspection model pre-training system according to claim 1, characterized in that: The quality inspection model pre-training system also includes a model training and fine-tuning module, which is used to freeze the visual encoder weights during training and use LoRA fine-tuning technology to update the text encoder weights.

3. The quality inspection model pre-training system according to claim 1, characterized in that: The knowledge graph construction module includes: An attribute knowledge encoding unit, used for text encoding of texture attributes, shape attributes, position attributes and color attributes of defect types; The association fusion unit is used to fuse the associations of different targets using a cross-attention mechanism.

4. The quality inspection model pre-training system according to claim 3, characterized in that: The texture attributes include material, production process and surface treatment; the shape attributes include area, length, size, thickness, and curvature.

5. The quality inspection model pre-training system according to claim 1, characterized in that: The feature fusion module fuses the visual features and text features based on the linear layer and inputs them into the decoder for decoding.

6. The quality inspection model pre-training system according to claim 1, characterized in that: The loss function calculation module includes: Blocked word prediction loss unit; Questions and answers indicate learning loss units; The target cluster encoding prediction loss unit is used to cluster encode text features and predict cluster encoding of visual features, and use the distance between the predicted encoded features as the loss value.

7. The quality inspection model pre-training system according to claim 6, characterized in that: The question-answer instruction learning loss unit adopts a text template including what it is, where it is, how big it is, what color it is, what it looks like, what processing it uses, what attributes it has, and what the associated target is, and generates multiple rounds of question and answer dialogue data for each picture based on ChatGPT to form an interactive sequence of images and texts.

8. A quality inspection model pre-training method based on the quality inspection model pre-training system according to any one of claims 1 to 7, characterized in that: The pre-training method includes the following steps: Acquire images, defect targets and attribute knowledge; Convert defect targets and their related description knowledge into a form that can be processed by a text encoder, and design a text template to obtain image question answering instruction data; The image is passed through the visual encoder to obtain the visual features of the product image, and the visual features and the text features of the text encoder are input into the Llama decoder as embedded features; The loss function is calculated using masked word prediction loss, question-answer instruction learning loss, and target cluster encoding prediction loss.

9. The quality inspection model pre-training method according to claim 8, characterized in that: The visual encoder weights are frozen during training, and the text encoder weights are updated using the LoRA fine-tuning technique.