Defect detection processing method and device, program product and equipment
By combining cross-modal feature extraction and fusion with large and lightweight models to optimize the defect detection model, the problem of identification accuracy in rapid line changes of multi-category industrial products was solved, achieving efficient and accurate defect detection.
Patent Information
- Application Number
- CN202511134480.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-11-18
AI Technical Summary
Existing deep learning-based industrial AI quality inspection solutions have poor recognition accuracy in industrial product production scenarios with multiple categories, small batches, and diverse defect forms, making it difficult to meet the needs of rapid production line changeover and limiting their large-scale application.
A pre-trained cross-modal large model is used for multimodal feature extraction and fusion. Combined with a lightweight initial defect detection model, the defect detection model is optimized through a collaborative optimization mechanism that guides the large model and deploys the small model, so as to achieve flexible production with high accuracy and rapid line changeover.
It achieves high-precision and rapid detection in industrial product inspection, meets the accuracy requirements of high-speed inspection on production lines, and has strong applicability to various scenarios and flexible production capabilities.
Smart Images

Figure CN120976174A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the field of information technology, and in particular, to a defect detection processing method, a defect detection processing apparatus, a computer program product and an electronic device. BACKGROUND
[0002] Industrial products are widely used in packaging, agriculture, medical treatment and electronics fields as key materials, and the surface quality of the industrial products directly affects the performance and safety of the products.
[0003] In the related art, an industrial AI (Artificial Intelligence) quality inspection scheme based on deep learning improves the automation level of industrial product quality detection, but has high dependence on large-scale labeled data, limited model capacity and needs to train special models for different products. In the production scene of multi-category, small batch and defect morphology of industrial products, the recognition accuracy of small defects or complex background is poor, which is difficult to meet the demand of rapid line change in production, and the scale application is limited.
[0004] It should be noted that the information disclosed in the above background section is only used to strengthen the understanding of the background of the present disclosure, and therefore can include information that does not constitute prior art known to those of ordinary skill in the art. SUMMARY
[0005] The present disclosure provides a defect detection processing method, a defect detection processing apparatus, a computer program product and an electronic device to at least partially solve the problems of poor recognition accuracy of industrial product defect detection, difficulty in meeting the demand of rapid line change in production and limited scale application in the related art.
[0006] According to a first aspect of the present disclosure, a defect detection processing method is provided, the method comprising: based on a pre-trained cross-modal large model, performing multi-modal feature extraction and fusion on an industrial sample product to obtain multi-modal fusion features of the industrial sample product; the industrial sample product comprising an industrial product with defects; based on a lightweight initial defect detection model, performing image feature extraction on the industrial sample product to obtain first image features of the industrial sample product; based on the first image features and the multi-modal fusion features of the industrial sample product, optimizing the initial defect detection model to obtain an optimized defect detection model, so as to perform defect detection on a to-be-detected industrial product through the optimized defect detection model.
[0007] According to a second aspect of this disclosure, a defect detection processing apparatus is provided, the apparatus comprising: a first extraction module, configured to extract and fuse multimodal features of an industrial sample product based on a pre-trained cross-modal large model to obtain multimodal fused features of the industrial sample product; the industrial sample product including industrial products with defects; a second extraction module, configured to extract image features of the industrial sample product based on a lightweight initial defect detection model to obtain first image features of the industrial sample product; and a model optimization module, configured to optimize the initial defect detection model based on the first image features of the industrial sample product and the multimodal fused features to obtain an optimized defect detection model, for defect detection of the industrial product to be detected using the optimized defect detection model.
[0008] According to a third aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the defect detection processing method of the first aspect and its possible implementations.
[0009] According to a fourth aspect of this disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform the defect detection processing method of the first aspect and possible implementations thereof by executing the executable instructions.
[0010] The technical solution disclosed herein has the following beneficial effects:
[0011] In the aforementioned industrial product inspection process, based on a pre-trained cross-modal large model, multimodal feature extraction and fusion are performed on industrial sample products to obtain multimodal fused features of the industrial sample products. The industrial sample products include defective products. Based on a lightweight initial defect detection model, image features are extracted from the industrial sample products to obtain the first image features of the industrial sample products. Based on the first image features and the multimodal fused features of the industrial sample products, the initial defect detection model is optimized to obtain an optimized defect detection model, which is then used to perform defect detection on the industrial products to be inspected. This disclosure, through a collaborative optimization mechanism of large model guidance and small model implementation, can both meet the accuracy requirements of high-speed production line inspection and the flexible production needs of rapid line changeover for industrial products, demonstrating strong applicability to various scenarios. Attached Figure Description
[0012] Figure 1 This diagram illustrates a flowchart of a defect detection processing method in this exemplary embodiment;
[0013] Figure 2 This diagram illustrates a defect detection process stage in this exemplary embodiment.
[0014] Figure 3 This illustration shows a flowchart of multimodal feature extraction and fusion of industrial sample products in this exemplary embodiment;
[0015] Figure 4 This exemplary embodiment illustrates a flowchart of an industrial product defect detection process guided by a pre-trained large model.
[0016] Figure 5 This diagram illustrates the deployment environment of a defect detection model according to this exemplary embodiment.
[0017] Figure 6 This diagram illustrates a structural block diagram of a defect detection and processing apparatus according to an exemplary embodiment of the present invention.
[0018] Figure 7 An electronic device for implementing the above-described defect detection processing method is shown in this exemplary embodiment. Detailed Implementation
[0019] Exemplary embodiments of this disclosure will be described more fully below with reference to the accompanying drawings.
[0020] The accompanying drawings are schematic illustrations of this disclosure and are not necessarily drawn to scale. Some block diagrams shown in the drawings may be functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in hardware modules or integrated circuits, or in networks, processors, or microcontrollers. Implementations can be carried out in various forms and should not be construed as limited to the examples set forth herein. The features, structures, or characteristics described in this disclosure can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough description of embodiments of this disclosure. However, those skilled in the art will recognize that one or more specific details may be omitted when implementing the technical solutions of this disclosure, or other methods, components, apparatuses, steps, etc., may be used to replace one or more specific details.
[0021] Among related technologies, industrial AI (Artificial Intelligence) quality inspection solutions based on deep learning have improved the automation level of industrial product quality inspection. However, they are highly dependent on large-scale labeled data, have limited model capacity, and require separate training of dedicated models for different products. In industrial product production scenarios with multiple categories, small batches, and diverse defect forms, this can lead to poor recognition accuracy for minor defects or complex backgrounds, making it difficult to meet the needs of rapid production line changeover and limiting large-scale application.
[0022] In view of the above problems, exemplary embodiments of this disclosure provide a defect detection processing method, a defect detection processing apparatus, a computer program product, and an electronic device.
[0023] Optionally, this disclosure may be applied to, but is not limited to: surface defect detection scenarios where there are few labeled defect samples or where the labor cost of labeling defect samples is high; surface defect detection scenarios where production lines produce a variety of products in small batches with large differences and fast line changes, such as plastic mold quality inspection scenarios, etc. This disclosure does not specifically limit these scenarios.
[0024] In one alternative implementation, refer to Figure 1 The diagram illustrates a defect detection and processing method, specifically including the following steps S110 to S130:
[0025] Step S110: Based on a pre-trained cross-modal large model, multimodal features are extracted and fused from industrial sample products to obtain multimodal fused features of industrial sample products; industrial sample products include defective industrial products.
[0026] Step S120: Based on the lightweight initial defect detection model, image features are extracted from the industrial sample product to obtain the first image features of the industrial sample product.
[0027] Step S130: Based on the first image features and multimodal fusion features of the industrial sample product, optimize the initial defect detection model to obtain the optimized defect detection model, so as to perform defect detection on the industrial product to be inspected through the optimized defect detection model.
[0028] Figure 1 The method shown uses a collaborative optimization mechanism that guides large models and implements small models to achieve both the accuracy requirements of high-speed production line testing and the flexible production needs of rapid line changeover for industrial products, making it highly applicable to various scenarios.
[0029] It should be noted that this disclosure may include a training phase and an inference phase. The training phase may include the following steps: based on a pre-trained cross-modal large model, multimodal feature extraction and fusion are performed on industrial sample products to obtain multimodal fused features of the industrial sample products; based on a lightweight initial defect detection model, image feature extraction is performed on the industrial sample products to obtain the first image features of the industrial sample products; based on the first image features and the multimodal fused features of the industrial sample products, the initial defect detection model is optimized to obtain an optimized defect detection model. The inference phase may include the following steps: defect detection is performed on the industrial product to be detected using the optimized defect detection model.
[0030] The following is about Figure 1 Each step in the process will be explained in detail.
[0031] In step S110, based on the pre-trained cross-modal large model, multimodal features are extracted and fused from industrial sample products to obtain multimodal fused features of industrial sample products; industrial sample products include industrial products with defects.
[0032] Among them, the pre-trained cross-modal large model is a pre-trained large model that can process data from different modalities simultaneously and has strong generalization ability. Optionally, multimodal features may include, but are not limited to, image, text, and other multimodal features.
[0033] The industrial sample products are instances of industrial products provided to the model for learning. These industrial products may have one or more types of defects. Taking a plastic film as an example, possible defects include, but are not limited to, holes, cracks, scratches, adhesions, etc. The specific defects are determined by the plastic mold generation scenario and are not specifically limited here.
[0034] In one optional implementation, the above-mentioned multimodal feature extraction and fusion of industrial sample products based on the pre-trained cross-modal large model to obtain the multimodal fused features of the industrial sample products can be achieved through the following steps: acquiring defect region images and corresponding defect text descriptions of the industrial sample products; extracting features from the defect region images and corresponding defect text descriptions of the industrial sample products based on the pre-trained cross-modal large model to obtain the second image features and text description features of the industrial sample products; and fusing the second image features and text description features of the industrial sample products based on fusion parameters to obtain the multimodal fused features of the industrial sample products.
[0035] Here, the defect area image refers to an image of an industrial product that contains defective areas. For example, an image containing defective areas, obtained by cropping the acquired image corresponding to an industrial sample product, can be used as the defect area image of the industrial sample product.
[0036] The defect text description corresponding to the defect area image can be a description of the type of defect in the defect area image. Taking the industrial sample product as a plastic mold as an example, the defect description text corresponding to the crack type defect area image is "crack"; the defect description text corresponding to the hole type defect area image is "hole".
[0037] After obtaining the defect area image and the corresponding defect text description of the industrial sample product, feature extraction can be performed on the defect area image and the corresponding defect text description of the industrial sample product based on the pre-trained cross-modal large model, so as to obtain the second image feature and text description feature of the industrial sample product.
[0038] The second image feature refers to the image features extracted using the large model; the text description feature refers to the text features extracted using the large model.
[0039] After obtaining the second image features and text description features of industrial sample products using a large model, the second image features and text description features of industrial sample products can be fused based on fusion parameters to obtain the multimodal fusion features of industrial sample products.
[0040] Optionally, the fusion parameters can be learnable and optimized using standard gradient backpropagation; no specific limitations are imposed here. Utilizing learnable fusion parameters to fuse multimodal features of a pre-trained, cross-modal large model enables adaptive extraction of multimodal features.
[0041] Since the features of different modalities can complement each other, but their importance is not equal, by introducing fusion parameters, it is possible to better adapt to industrial inspection scenarios.
[0042] By integrating complementary information from different modal data, the perception capability and robustness of large models can be improved. Through prototype learning of multimodal joint representation of defect features, large models can capture more universal feature patterns and have stronger generalization ability.
[0043] In one optional implementation, the large model includes an image feature extractor and a text feature extractor. The aforementioned pre-trained cross-modal large model extracts features from the defect region image and the corresponding defect text description of the industrial sample product to obtain the second image features and text description features of the industrial sample product. This can be achieved through the following steps: using the image feature extractor to extract features from the defect region image of the industrial sample product to obtain the second image features of the industrial sample product; and using the text feature extractor to extract features from the corresponding defect text description of the defect region image to obtain the text description features of the industrial sample product.
[0044] Multimodal feature extraction using image and text feature extractors from large models can provide a data foundation for further feature fusion.
[0045] It should be noted that, in practical applications, the steps of feature extraction from the defect area image of the industrial sample product and feature extraction from the corresponding defect text description can be performed simultaneously or sequentially. This disclosure does not impose any specific limitations on this.
[0046] For example, such as Figure 2As shown in the adaptive fusion stage based on multimodal large model features during the training phase, the defect region images {x1,...,x} of several industrial sample products (taking plastic film as an example) can be fused together. i The images are input into the image feature extractor to obtain the second image features of the defect area images of each industrial sample product; the defect text descriptions {t1,...,t} corresponding to the defect area images of each industrial sample product can then be used to extract the defect text descriptions {t1,...,t}. i The text features of each industrial sample product are input into the text feature extractor to obtain the text description features. The second image features and text description features of each industrial sample product are then fused and summed to obtain the multimodal fusion features {C1,...,C} of each industrial sample product. i}
[0047] For example, the following formula (1) can be used for feature fusion.
[0048] C i =τ·I(x) i )+(1-τ)·T(t i (1)
[0049] Where, x i The image shows the defect area of the i-th industrial sample product; t i I(x) is the text description of the defect region image corresponding to the i-th industrial sample product; i () is an image feature extractor for pre-trained cross-modal large models targeting image x i The extracted image features, i.e., the second image features of the i-th industrial sample product; T(t) i A text feature extractor for pre-trained cross-modal large models targeting t i The extracted text features are the textual description features of the i-th industrial sample product; τ is the fusion parameter; C i Let be the multimodal fusion feature of the i-th industrial sample product.
[0050] For example, such as Figure 3 As shown, a flowchart for multimodal feature extraction and fusion of industrial sample products is provided, which may include the following steps:
[0051] Step S301: Obtain the defect area image of the industrial sample product and the corresponding defect text description;
[0052] Step S302: Based on the pre-trained cross-modal large model image feature extractor, feature extraction is performed on the defect area image of the industrial sample product to obtain the second image feature of the industrial sample product.
[0053] Step S303: Based on the pre-trained cross-modal large model text feature extractor, feature extraction is performed on the defect text description corresponding to the defect region image to obtain the text description features of the industrial sample product.
[0054] Step S304: Based on the fusion parameters, the second image features and text description features of the industrial sample products are fused to obtain the multimodal fusion features of the industrial sample products.
[0055] Figure 3 In the steps shown, visual features and textual description features of the defect region are extracted by the image encoder and text encoder of the large model, respectively, and the two are fused to form a more discriminative multimodal representation. The cross-modal generalization ability of the large model is used to improve the feature robustness.
[0056] Understandably, multimodal feature fusion utilizes visual modalities to capture spatial features (i.e., secondary image features, such as color and shape) and textual modalities to provide semantic information (i.e., textual descriptive features, such as labels and descriptions). The fusion eliminates the perceptual blind spots of a single modality. When the data quality of one modality is low (e.g., blurry images), other modalities can provide redundant information to ensure system stability and exhibit noise robustness. Simultaneously, cross-modal association can uncover deeper semantics, improving recognition accuracy. Through multimodal joint representation learning, large models can capture more universal feature patterns to meet the high-generalization production needs of industrial quality inspection scenarios, such as rapid line changes.
[0057] In step S120, based on a lightweight initial defect detection model, image features are extracted from the industrial sample product to obtain the first image features of the industrial sample product.
[0058] The initial defect detection model refers to a lightweight defect detection model that has not yet been guided by a large model. It can be a defect detection model trained with a small number of samples. Here, "lightweight" is used to characterize the case where the number of model parameters is small.
[0059] Optionally, different components, such as feature pyramids, can be added to the defect detection model according to the actual application. At the same time, a suitable network framework can be selected according to the accuracy and real-time requirements of the detection task. This disclosure does not impose specific limitations on this.
[0060] The first image feature refers to the image features extracted using the defect detection model.
[0061] For example, such as Figure 2 During the training phase, the defect detection model based on a large model-guided distillation stage is shown. This stage can distill the defect region images {x1,...,x} of several industrial sample products (taking plastic film as an example). iThe region extraction network of the defect detection model is input into each of the following: {R1,...,R} are used to obtain the first image features {R1,...,R} of the defect region images of each industrial sample product. i}
[0062] In step S130, the initial defect detection model is optimized based on the first image features and multimodal fusion features of the industrial sample product to obtain the optimized defect detection model, so as to perform defect detection on the industrial product to be inspected through the optimized defect detection model.
[0063] Among them, the industrial products to be tested can be plastic molds that require defect identification and detection during the actual production process.
[0064] Through optimization, the low-dimensional space of the lightweight defect detection model can be aligned with the high-dimensional feature space of the large model, so that the lightweight defect detection model can maintain high inference speed and significantly improve detection accuracy. This can solve the problem of limited feature extraction capability and high false negative rate of small models, so as to meet the quality inspection requirements of high accuracy and low false negative rate in industrial quality inspection scenarios.
[0065] In one optional implementation, the above-mentioned optimization of the initial defect detection model based on the first image features and multimodal fusion features of the industrial sample product to obtain the optimized defect detection model can be achieved through the following steps: determining the comprehensive loss of the industrial sample product based on the first image features and multimodal fusion features of the industrial sample product; optimizing the initial defect detection model based on the comprehensive loss of the industrial sample product to obtain the optimized defect detection model.
[0066] Among them, the comprehensive loss refers to the total loss after the superposition of multiple types of losses, such as distillation loss and contrast loss.
[0067] For example, the comprehensive loss of an industrial sample product can be determined based on the first image features and multimodal fusion features of the industrial sample product. Based on the comprehensive loss of the industrial sample product, the model parameters of the defect detection model can be optimized through gradient backpropagation, thereby obtaining an optimized defect detection model. This allows the defect detection model with fewer parameters to inherit the powerful semantic understanding and general feature representation of the large model, thus achieving the purpose of knowledge transfer and feature enhancement from the pre-trained large model to the defect detection model with fewer parameters. This can increase the adaptability of the defect detection model to complex scenarios and improve the accuracy of defect detection.
[0068] In one optional implementation, the determination of the comprehensive loss of the industrial sample product based on the first image features and multimodal fusion features of the industrial sample product can be achieved through the following steps: determining the distillation loss of the industrial sample product based on the first image features and multimodal fusion features of the industrial sample product; determining the contrast loss of the industrial sample product based on the first image features and multimodal fusion features of the industrial sample product; and determining the comprehensive loss of the industrial sample product based on the distillation loss and the contrast loss of the industrial sample product.
[0069] For example, the overall loss of an industrial sample product can be obtained by calculating the following formula (2).
[0070] Lcomprehensive = Ldistilled + βLcomparison
[0071] Among them, L 综合 Characterizing the overall loss of industrial sample products; L 蒸馏 Characterizing distillation loss; L 对比 Characterizes the contrast loss; β is a hyperparameter that can be preset based on experience, and is not specifically limited here.
[0072] By calculating distillation loss and contrast loss, distillation and contrast learning are achieved between the multimodal fusion features extracted by the large model and the first image features of the defect detection model, so that the defect detection model can inherit the powerful semantic understanding ability and general feature representation of the large model.
[0073] In an optional implementation, the determination of the distillation loss of the industrial sample product based on the first image features and multimodal fusion features of the industrial sample product can be achieved through the following steps: inputting the first image features and multimodal fusion features of the industrial sample product into the distillation loss function to obtain the distillation loss of the industrial sample product; wherein, the distillation loss function is an L1 norm-based loss function used to align the first image features and the multimodal fusion features.
[0074] For example, the loss function based on the L1 norm is shown in Equation (3) below.
[0075] Ldistillation = ∑ i ||C i -R i || l1 (3)
[0076] Where l1 is the standard L1 norm loss.
[0077] By constraining the distillation loss function, the first image features extracted by the defect detection model with fewer parameters can be aligned with the multimodal features extracted by the large model. This enables the distillation of the multimodal fusion features of the pre-trained cross-modal large model into the defect detection model with fewer parameters, allowing the detection network with fewer parameters to obtain the prior information of the pre-trained large model, thereby ensuring generalization performance.
[0078] In an optional implementation, the determination of the contrast loss of industrial sample products based on the first image features and multimodal fusion features of industrial sample products can be achieved through the following steps: obtaining fusion parameters and balancing parameters; wherein, the fusion parameters are parameters used to fuse the multimodal features of industrial sample products extracted by the large model, and the balancing parameters are parameters used to balance the similarity and class embedding distance between the first image features and multimodal fusion features of industrial sample products; inputting the fusion parameters, balancing parameters, and the first image features and multimodal fusion features of industrial sample products into the contrast loss function to obtain the contrast loss of industrial sample products; wherein, the contrast loss function achieves similarity discrimination of sample pairs by introducing boundary conditions; the boundary conditions are constructed from the multimodal fusion features of different industrial sample products.
[0079] The balance parameter can be a hyperparameter, which can be preset based on experience; no specific limitations are made here.
[0080] For example, in industrial quality inspection applications, the discriminability of regional defect features can be improved by introducing a contrastive loss function with boundary conditions (i.e., a boundary metric loss function) as contrastive training.
[0081] For example, the contrastive loss function with boundary conditions is shown in Equation (4) below.
[0082]
[0083] Among them, D(C) i C j ) represents different multimodal fusion features C i and C j The distance between them is a boundary condition, D(C) i C j )=1-S(C i C j ), S(C i C j )=cos(C i C j ); where S(R) i C i )=cos(R i C i ), λ·D(C iC j In the loss function, for each S(R) i C i ) serves as the adaptive boundary; where λ is the balance parameter and τ is the fusion parameter.
[0084] Understandably, a good defect detection model needs not only to increase the similarity of features within the same defect category, but also to reduce the similarity between features of the defect category and features of the normal category. Otherwise, the defect detection model will not optimize the classification boundary from this data and will stop learning prematurely, which will limit the model's representation ability to some extent. This disclosure, by introducing a contrastive loss function with boundary conditions, can force the similarity of multimodal fusion features of the same category to be higher than the similarity of features of different categories. Since different defects may have similar semantic expressions and relatively small semantic distances, the adaptive term D(C) is used. i C j With a lower value, the contrast loss is smaller, allowing for weak optimization of such classification boundaries; for features of the normal category and features of the defect category, the semantic distance is relatively large, in which case the adaptive term D(C) is more suitable. i C j With a larger value, the contrast loss is higher. Strong optimization can be performed on this type of classification boundary to further reduce the distance between the features of the defect category and the normal category, so that the model has a clear classification boundary.
[0085] The contrastive loss function with boundary conditions fully utilizes the detailed knowledge of semantic relationships in multimodal fusion features, which can optimize the similarity of samples of the same category when different defects are distinguished by a defect detection model with a small number of parameters, thereby improving the robustness of the defect detection model.
[0086] After the defect detection model is optimized, such as Figure 2 In the inference stage based on the lightweight defect detection model, the image of the industrial product to be inspected (taking plastic film as an example) can be input into the optimized defect detection model to identify defects in the industrial product to be inspected.
[0087] During the inference phase, retaining only the defect detection model with smaller parameters can significantly improve the detection accuracy of small target defects in industrial scenarios while maintaining detection efficiency.
[0088] By using a defect detection model with fewer parameters during the inference phase and removing the pre-trained large model portion from the training phase, the defect detection model trained on a small number of samples can leverage the features of the pre-trained large model for enhancement, while maintaining high inference efficiency. In industrial scenarios, the low-parameter defect detection model has a high inference speed, which can better adapt to the high-speed operation of production lines and meet the speed requirements of industrial production quality inspection.
[0089] For example, such as Figure 4 As shown, a flowchart of industrial product defect detection processing based on a pre-trained large model is provided, which may include the following steps:
[0090] Step S401: During the training phase, based on the pre-trained cross-modal large model, multimodal features of industrial sample products are extracted and fused to obtain the multimodal fused features of industrial sample products.
[0091] Step S402: Based on the lightweight initial defect detection model, image features are extracted from the industrial sample product to obtain the first image features of the industrial sample product.
[0092] Step S403: Determine the distillation loss of the industrial sample product based on the first image features and multimodal fusion features of the industrial sample product;
[0093] Step S404: Determine the contrast loss of the industrial sample products based on the first image features and multimodal fusion features of the industrial sample products;
[0094] Step S405: Determine the overall loss of the industrial sample products based on the distillation loss and the comparative loss of the industrial sample products.
[0095] Step S406: Based on the comprehensive loss of industrial sample products, optimize the initial defect detection model to obtain the optimized defect detection model;
[0096] Step S407: In the inference stage, defect detection is performed on the industrial product to be inspected using the optimized defect detection model.
[0097] Figure 4 In the steps shown, during the training phase, the pre-trained cross-modal large model multimodal features are used for adaptive fusion to fully utilize the prior information of the large model multimodal features and guide the distillation and contrastive learning of the defect detection model. During the inference phase, a low-parameter defect detection model is used for inference, which can achieve high computational efficiency and detection accuracy, meet the real-time and accuracy requirements of high-speed detection on the production line, and improve generalization.
[0098] like Figure 5The diagram illustrates a deployment environment for a defect detection model. It includes modules such as a visual data acquisition system, an industrial production line, an industrial control computer (ICC), and a display platform. The visual data acquisition system supports the connection of various sensors, enabling image acquisition of product objects and transmitting the acquired data to the ICC. The industrial production line is tightly coupled with industrial equipment, controlling the operation of the equipment and data communication with the ICC. For example, communication protocols such as Modbus and TCP can be used, without specific limitations. The ICC can process data uploads from multiple devices via industrial protocols, HTTP protocols, TCP protocols, etc., aggregating statistical results from multiple production lines to the display platform. The display platform includes an application platform (providing data acquisition management, inference optimization deployment, and defect statistical classification) and a system platform (providing model optimization, algorithm models, and data processing). The display platform can display information such as production line number, operating status, defect type, product model, and defect statistical information; this disclosure does not impose specific limitations on this aspect.
[0099] Exemplary embodiments of this disclosure also provide a defect detection processing apparatus. (See reference...) Figure 6 As shown, the defect detection and processing device 600 may include the following program modules:
[0100] The first extraction module 610 is used to extract and fuse multimodal features of industrial sample products based on a pre-trained cross-modal large model to obtain the multimodal fused features of the industrial sample products; the industrial sample products include industrial products with defects;
[0101] The second extraction module 620 is used to extract image features from industrial sample products based on a lightweight initial defect detection model to obtain the first image features of the industrial sample products.
[0102] The model optimization module 630 is used to optimize the initial defect detection model based on the first image features and multimodal fusion features of the industrial sample product, so as to obtain the optimized defect detection model and perform defect detection on the industrial product to be inspected.
[0103] In an optional implementation, based on the aforementioned scheme, the first extraction module 610 includes: an acquisition module, used to acquire defect region images of industrial sample products and corresponding defect text descriptions; a multimodal feature extraction module, used to extract features from the defect region images and corresponding defect text descriptions of industrial sample products based on a pre-trained cross-modal large model, to obtain second image features and text description features of the industrial sample products; and a multimodal feature fusion module, used to fuse the second image features and text description features of the industrial sample products based on fusion parameters, to obtain multimodal fused features of the industrial sample products.
[0104] In an optional implementation, based on the aforementioned scheme, the large model includes an image feature extractor and a text feature extractor. The multimodal feature extraction module can be configured to: extract features from the defect area image of the industrial sample product using the image feature extractor to obtain the second image features of the industrial sample product; and extract features from the defect text description corresponding to the defect area image using the text feature extractor to obtain the text description features of the industrial sample product.
[0105] In an optional implementation, based on the aforementioned scheme, the model optimization module 630 may include: a comprehensive loss determination module, used to determine the comprehensive loss of the industrial sample product based on the first image features and multimodal fusion features of the industrial sample product; and a detection model optimization module, used to optimize the initial defect detection model based on the comprehensive loss of the industrial sample product to obtain the optimized defect detection model.
[0106] In an optional implementation, based on the aforementioned scheme, the comprehensive loss determination module further includes: a distillation loss determination module, used to determine the distillation loss of the industrial sample product based on the first image features and multimodal fusion features of the industrial sample product; a contrast loss determination module, used to determine the contrast loss of the industrial sample product based on the first image features and multimodal fusion features of the industrial sample product; and a loss synthesis module, used to determine the comprehensive loss of the industrial sample product based on the distillation loss and the contrast loss of the industrial sample product.
[0107] In an optional implementation, based on the aforementioned scheme, the distillation loss determination module can be configured to: input the first image features and multimodal fusion features of the industrial sample product into the distillation loss function to obtain the distillation loss of the industrial sample product; wherein, the distillation loss function is an L1 norm-based loss function used to align the first image features and the multimodal fusion features.
[0108] In an optional implementation, based on the aforementioned scheme, the contrast loss determination module can be configured to: acquire fusion parameters and balancing parameters; wherein, the fusion parameters are parameters used to fuse the multimodal features of industrial sample products extracted by the large model, and the balancing parameters are parameters used to balance the similarity and category embedding distance between the first image features of the industrial sample products and the multimodal fusion features; the fusion parameters, balancing parameters, and the first image features and multimodal fusion features of the industrial sample products are input into the contrast loss function to obtain the contrast loss of the industrial sample products; wherein, the contrast loss function achieves similarity discrimination of sample pairs by introducing boundary conditions; the boundary conditions are constructed from the multimodal fusion features of different industrial sample products.
[0109] The specific details of each part of the aforementioned defect detection and processing device 600 have been described in detail in the method section of the implementation plan. Any undisclosed details can be found in the implementation plan of the method section, and therefore will not be repeated here.
[0110] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to exemplary embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0111] An exemplary embodiment of this disclosure also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the above-described defect detection processing method.
[0112] In one embodiment, the computer program product can be a tangible product containing a computer program, such as a computer-readable storage medium storing the computer program. The readable storage medium can be a storage medium based on electrical, magnetic, optical, electromagnetic, infrared, or other signals, including but not limited to: random access memory (RAM), read-only memory (ROM), magnetic tape, floppy disk, flash memory, hard disk drive (HDD), solid-state drive (SSD), etc. For example, the computer program product can be implemented as a non-volatile storage medium storing the computer program, such as read-only memory, NAND flash memory, etc.
[0113] In one implementation, the computer program product can be an intangible product containing a computer program. For example, the computer program product can be implemented as a virtual digital product, such as an executable file, installation package, or other digital file storing the computer program.
[0114] Computer program code can be written in one or more programming languages. Examples of programming languages include C, Java, and C++. Program code can execute entirely on the user's computing device, partially on the user's computing device, or as a standalone software package. It can also execute partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, such as a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via an internet connection provided by a mobile network operator).
[0115] Computer programs can be carried or transmitted via signals such as electricity, magnetism, light, electromagnetic radiation, and infrared rays. Electronic devices can convert signals carrying computer programs into digital signals, thereby running the computer programs. When a computer program runs on an electronic device, its code causes the electronic device to execute (more specifically, its processor) the method steps of various exemplary embodiments of this disclosure, such as the methods described above, which include the following steps:
[0116] Based on a pre-trained cross-modal large model, multimodal features are extracted and fused from industrial sample products to obtain multimodal fused features of industrial sample products; industrial sample products include defective industrial products;
[0117] Based on a lightweight initial defect detection model, image features are extracted from industrial sample products to obtain the first image features of the industrial sample products.
[0118] Based on the first image features and multimodal fusion features of industrial sample products, the initial defect detection model is optimized to obtain the optimized defect detection model, which is then used to detect defects in the industrial products to be inspected.
[0119] In an optional implementation, based on the aforementioned scheme, the above-mentioned multimodal feature extraction and fusion of industrial sample products based on the pre-trained cross-modal large model to obtain the multimodal fused features of the industrial sample products can be achieved through the following steps: acquiring defect region images and corresponding defect text descriptions of the industrial sample products; extracting features from the defect region images and corresponding defect text descriptions of the industrial sample products based on the pre-trained cross-modal large model to obtain the second image features and text description features of the industrial sample products; and fusing the second image features and text description features of the industrial sample products based on fusion parameters to obtain the multimodal fused features of the industrial sample products.
[0120] In an optional implementation, based on the aforementioned scheme, the large model includes an image feature extractor and a text feature extractor. The aforementioned pre-trained cross-modal large model extracts features from the defect region image and the corresponding defect text description of the industrial sample product to obtain the second image features and text description features of the industrial sample product. This can be achieved through the following steps: using the image feature extractor to extract features from the defect region image of the industrial sample product to obtain the second image features of the industrial sample product; using the text feature extractor to extract features from the corresponding defect text description of the defect region image to obtain the text description features of the industrial sample product.
[0121] In an optional implementation, based on the aforementioned scheme, the optimization of the initial defect detection model based on the first image features and multimodal fusion features of the industrial sample product to obtain the optimized defect detection model can be achieved through the following steps: determining the comprehensive loss of the industrial sample product based on the first image features and multimodal fusion features of the industrial sample product; optimizing the initial defect detection model based on the comprehensive loss of the industrial sample product to obtain the optimized defect detection model.
[0122] In an optional implementation, based on the aforementioned scheme, the determination of the comprehensive loss of the industrial sample product based on the first image features and multimodal fusion features of the industrial sample product can be achieved through the following steps: determining the distillation loss of the industrial sample product based on the first image features and multimodal fusion features of the industrial sample product; determining the contrast loss of the industrial sample product based on the first image features and multimodal fusion features of the industrial sample product; and determining the comprehensive loss of the industrial sample product based on the distillation loss and the contrast loss of the industrial sample product.
[0123] In an optional implementation, based on the aforementioned scheme, the determination of the distillation loss of the industrial sample product based on the first image features and multimodal fusion features of the industrial sample product can be achieved through the following steps: inputting the first image features and multimodal fusion features of the industrial sample product into the distillation loss function to obtain the distillation loss of the industrial sample product; wherein, the distillation loss function is an L1 norm-based loss function used to align the first image features and the multimodal fusion features.
[0124] In an optional implementation, based on the aforementioned scheme, the determination of the contrast loss of industrial sample products based on the first image features and multimodal fusion features of industrial sample products can be achieved through the following steps: obtaining fusion parameters and balancing parameters; wherein, the fusion parameters are parameters used to fuse the multimodal features of industrial sample products extracted by the large model, and the balancing parameters are parameters used to balance the similarity and class embedding distance between the first image features and multimodal fusion features of industrial sample products; inputting the fusion parameters, balancing parameters, and the first image features and multimodal fusion features of industrial sample products into the contrast loss function to obtain the contrast loss of industrial sample products; wherein, the contrast loss function achieves similarity discrimination of sample pairs by introducing boundary conditions; the boundary conditions are constructed from the multimodal fusion features of different industrial sample products.
[0125] In the aforementioned industrial product testing and processing process, the collaborative optimization mechanism of large-scale model guidance and small-scale model implementation can not only meet the accuracy requirements of high-speed production line testing, but also meet the flexible production needs of rapid line changeover for industrial products, making it highly applicable to various scenarios.
[0126] An exemplary embodiment of this disclosure also provides an electronic device capable of implementing the above-described defect detection processing method. The electronic device may include a processor and a memory. The memory stores executable instructions of the processor, such as program code. The processor executes the executable instructions to perform the method of this exemplary embodiment. Furthermore, the electronic device may also include a display for displaying a graphical user interface.
[0127] The following is for reference. Figure 7 The electronic device is illustrated by way of a general-purpose computing device. It should be understood that... Figure 7 The electronic device 700 shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments disclosed herein.
[0128] like Figure 7 As shown, the electronic device 700 may include: a processor 710, a memory 720, a bus 730, an I / O (input / output) interface 740, a network adapter 750, and a display 760.
[0129] The memory 720 may include volatile memory, such as RAM 721 and cache unit 722, and may also include non-volatile memory, such as ROM 723. The memory 720 may also include one or more program modules 724, including but not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. For example, program module 724 may include the modules described above.
[0130] The processor 710 may include one or more processing units, such as an AP (Application Processor), a modem processor, a GPU (Graphics Processing Unit), an ISP (Image Signal Processor), a controller, an encoder, a decoder, a DSP (Digital Signal Processor), a baseband processor, and / or an NPU (Neural-Network Processing Unit).
[0131] The processor 710 can be used to execute executable instructions stored in the memory 720, such as performing any one or more method steps in this exemplary embodiment.
[0132] For example, processor 710 may perform the following steps:
[0133] Based on a pre-trained cross-modal large model, multimodal features are extracted and fused from industrial sample products to obtain multimodal fused features of industrial sample products; industrial sample products include defective industrial products;
[0134] Based on a lightweight initial defect detection model, image features are extracted from industrial sample products to obtain the first image features of the industrial sample products.
[0135] Based on the first image features and multimodal fusion features of industrial sample products, the initial defect detection model is optimized to obtain the optimized defect detection model, which is then used to detect defects in the industrial products to be inspected.
[0136] In an optional implementation, based on the aforementioned scheme, the above-mentioned multimodal feature extraction and fusion of industrial sample products based on the pre-trained cross-modal large model to obtain the multimodal fused features of the industrial sample products can be achieved through the following steps: acquiring defect region images and corresponding defect text descriptions of the industrial sample products; extracting features from the defect region images and corresponding defect text descriptions of the industrial sample products based on the pre-trained cross-modal large model to obtain the second image features and text description features of the industrial sample products; and fusing the second image features and text description features of the industrial sample products based on fusion parameters to obtain the multimodal fused features of the industrial sample products.
[0137] In an optional implementation, based on the aforementioned scheme, the large model includes an image feature extractor and a text feature extractor. The aforementioned pre-trained cross-modal large model extracts features from the defect region image and the corresponding defect text description of the industrial sample product to obtain the second image features and text description features of the industrial sample product. This can be achieved through the following steps: using the image feature extractor to extract features from the defect region image of the industrial sample product to obtain the second image features of the industrial sample product; using the text feature extractor to extract features from the corresponding defect text description of the defect region image to obtain the text description features of the industrial sample product.
[0138] In an optional implementation, based on the aforementioned scheme, the optimization of the initial defect detection model based on the first image features and multimodal fusion features of the industrial sample product to obtain the optimized defect detection model can be achieved through the following steps: determining the comprehensive loss of the industrial sample product based on the first image features and multimodal fusion features of the industrial sample product; optimizing the initial defect detection model based on the comprehensive loss of the industrial sample product to obtain the optimized defect detection model.
[0139] In an optional implementation, based on the aforementioned scheme, the determination of the comprehensive loss of the industrial sample product based on the first image features and multimodal fusion features of the industrial sample product can be achieved through the following steps: determining the distillation loss of the industrial sample product based on the first image features and multimodal fusion features of the industrial sample product; determining the contrast loss of the industrial sample product based on the first image features and multimodal fusion features of the industrial sample product; and determining the comprehensive loss of the industrial sample product based on the distillation loss and the contrast loss of the industrial sample product.
[0140] In an optional implementation, based on the aforementioned scheme, the determination of the distillation loss of the industrial sample product based on the first image features and multimodal fusion features of the industrial sample product can be achieved through the following steps: inputting the first image features and multimodal fusion features of the industrial sample product into the distillation loss function to obtain the distillation loss of the industrial sample product; wherein, the distillation loss function is an L1 norm-based loss function used to align the first image features and the multimodal fusion features.
[0141] In an optional implementation, based on the aforementioned scheme, the determination of the contrast loss of industrial sample products based on the first image features and multimodal fusion features of industrial sample products can be achieved through the following steps: obtaining fusion parameters and balancing parameters; wherein, the fusion parameters are parameters used to fuse the multimodal features of industrial sample products extracted by the large model, and the balancing parameters are parameters used to balance the similarity and class embedding distance between the first image features and multimodal fusion features of industrial sample products; inputting the fusion parameters, balancing parameters, and the first image features and multimodal fusion features of industrial sample products into the contrast loss function to obtain the contrast loss of industrial sample products; wherein, the contrast loss function achieves similarity discrimination of sample pairs by introducing boundary conditions; the boundary conditions are constructed from the multimodal fusion features of different industrial sample products.
[0142] In the aforementioned industrial product testing and processing process, the collaborative optimization mechanism of large-scale model guidance and small-scale model implementation can not only meet the accuracy requirements of high-speed production line testing, but also meet the flexible production needs of rapid line changeover for industrial products, making it highly applicable to various scenarios.
[0143] Bus 730 is used to connect different components of electronic device 700 and may include a data bus, an address bus and a control bus.
[0144] Electronic device 700 can communicate with one or more external devices 800 (such as keyboard, mouse, external controller, etc.) through I / O interface 740.
[0145] Electronic device 700 can communicate with one or more networks via network adapter 750. For example, network adapter 750 can provide mobile communication solutions such as 3G / 4G / 5G, or wireless communication solutions such as wireless LAN, Bluetooth, and near-field communication. Network adapter 750 can communicate with other modules of electronic device 700 via bus 730.
[0146] Electronic device 700 can display a graphical user interface, etc., via display 760.
[0147] although Figure 7 As not shown in the diagram, other hardware and / or software modules may also be configured in the electronic device 700, including but not limited to: a display, microcode, device driver, redundant processor, external disk drive array, RAID (Redundant Arrays of Independent Disks) system, tape drive, and data backup storage system.
[0148] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to exemplary embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0149] Those skilled in the art will understand that various aspects of this disclosure can be implemented as systems, methods, or program products. Therefore, various aspects of this disclosure can be embodied in entirely hardware implementations, entirely software implementations (including firmware, microcode, etc.), or implementations combining hardware and software aspects, collectively referred to herein as “circuit,” “module,” or “system.” Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.
[0150] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is defined only by the appended claims.
Claims
1. A defect detection and processing method, characterized in that, The method includes: Based on a pre-trained cross-modal large model, multimodal features are extracted and fused from industrial sample products to obtain the multimodal fused features of the industrial sample products; the industrial sample products include industrial products with defects. Based on a lightweight initial defect detection model, image features are extracted from the industrial sample product to obtain the first image features of the industrial sample product. Based on the first image features and the multimodal fusion features of the industrial sample product, the initial defect detection model is optimized to obtain an optimized defect detection model, which is then used to perform defect detection on the industrial product to be inspected.
2. The method according to claim 1, characterized in that, The pre-trained cross-modal large model extracts and fuses multimodal features from industrial sample products to obtain the multimodal fused features of the industrial sample products, including: Acquire images of defective areas in industrial sample products and corresponding textual descriptions of the defects. Based on a pre-trained cross-modal large model, feature extraction is performed on the defect region image of the industrial sample product and the corresponding defect text description, respectively, to obtain the second image feature and text description feature of the industrial sample product. Based on the fusion parameters, the second image features and the text description features of the industrial sample product are fused to obtain the multimodal fusion features of the industrial sample product.
3. The method according to claim 2, characterized in that, The large model includes an image feature extractor and a text feature extractor. The pre-trained cross-modal large model extracts features from the defect region image of the industrial sample product and the corresponding defect text description, respectively, to obtain the second image features and text description features of the industrial sample product, including: The image feature extractor is used to extract features from the defect area image of the industrial sample product to obtain the second image features of the industrial sample product. The text feature extractor extracts features from the defect text description corresponding to the defect region image to obtain the text description features of the industrial sample product.
4. The method according to claim 1, characterized in that, The optimization of the initial defect detection model based on the first image features of the industrial sample product and the multimodal fusion features to obtain the optimized defect detection model includes: Based on the first image features of the industrial sample product and the multimodal fusion features, the comprehensive loss of the industrial sample product is determined; Based on the comprehensive loss of the industrial sample products, the initial defect detection model is optimized to obtain the optimized defect detection model.
5. The method according to claim 4, characterized in that, The determination of the comprehensive loss of the industrial sample product based on the first image features and the multimodal fusion features includes: Based on the first image features of the industrial sample product and the multimodal fusion features, the distillation loss of the industrial sample product is determined; Based on the first image features of the industrial sample product and the multimodal fusion features, the contrast loss of the industrial sample product is determined; The overall loss of the industrial sample product is determined based on the distillation loss and the comparative loss of the industrial sample product.
6. The method according to claim 5, characterized in that, The determination of the distillation loss of the industrial sample product based on the first image features and the multimodal fusion features includes: The first image feature and the multimodal fusion feature of the industrial sample product are input into the distillation loss function to obtain the distillation loss of the industrial sample product; wherein, the distillation loss function is an L1 norm-based loss function used to align the first image feature and the multimodal fusion feature.
7. The method according to claim 5, characterized in that, The step of determining the contrast loss of the industrial sample product based on the first image features and the multimodal fusion features of the industrial sample product includes: Obtain fusion parameters and balancing parameters; wherein, the fusion parameters are parameters used to fuse the multimodal features of the industrial sample products extracted by the large model, and the balancing parameters are parameters used to balance the similarity and category embedding distance between the first image features of the industrial sample products and the multimodal fused features; The fusion parameters, the balancing parameters, and the first image features of the industrial sample products, along with the multimodal fusion features, are input into the contrast loss function to obtain the contrast loss of the industrial sample products. The contrast loss function achieves similarity discrimination of sample pairs by introducing boundary conditions. The boundary conditions are constructed from the multimodal fusion features of different industrial sample products.
8. A defect detection and processing device, characterized in that, The device includes: The first extraction module is used to extract and fuse multimodal features of industrial sample products based on a pre-trained cross-modal large model to obtain the multimodal fused features of the industrial sample products; the industrial sample products include industrial products with defects. The second extraction module is used to extract image features from the industrial sample product based on a lightweight initial defect detection model to obtain the first image features of the industrial sample product. The model optimization module is used to optimize the initial defect detection model based on the first image features and the multimodal fusion features of the industrial sample product to obtain an optimized defect detection model, so as to perform defect detection on the industrial product to be inspected through the optimized defect detection model.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method described in any one of claims 1 to 7.
10. An electronic device, characterized in that, include: processor; as well as Memory for storing the executable instructions of the processor; The processor is configured to execute the method of any one of claims 1 to 7 by executing the executable instructions.