Discrete manufacturing product defect AI visual inspection method based on multi-modal fusion
Through multimodal fusion technology, combined with visible light and thermal infrared images, the problem of insufficient accuracy of machine vision inspection in complex environments is solved, and efficient detection of small and internal defects is achieved, the detection accuracy and stability are improved, and it adapts to complex lighting conditions. It is suitable for automated quality inspection in the discrete manufacturing industry.
Patent Information
- Application Number
- CN202511292444.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-11
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-09-11
AI Technical Summary
Existing machine vision inspection methods lack accuracy in complex industrial environments and find it difficult to fully capture tiny and internal defects in products. Single-modality inspection is also susceptible to changes in lighting and environmental interference, and cannot meet the modern manufacturing industry's demand for high-precision inspection.
A multimodal fusion method is adopted to combine visible light images and thermal infrared images. Through photometric distortion correction, radiometric calibration correction and spatial registration, multi-level DWT and DS evidence theory are used for feature-level and decision-level fusion. Combined with the dynamic attention gating module, multi-scale feature aggregation and global attention fusion, a visual detection model is constructed, and real-time detection is achieved on an embedded AI computing platform.
It significantly improves the ability to detect tiny surface and internal defects, improves detection accuracy and stability, reduces false detection rate, has anti-interference and high reliability, adapts to complex lighting conditions, and realizes batch automated detection.
Smart Images

Figure CN120807500A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the technical field of product detection, and in particular to a discrete manufacturing product defect AI visual detection method based on multi-modal fusion. BACKGROUND
[0002] In discrete manufacturing, the appearance quality of products not only concerns the external image of the products, but also directly affects their performance and market competitiveness. A small appearance defect, such as a scratch, a crack, or a flaw, may cause a malfunction in the use of the product, thereby affecting the quality and reliability of the entire product. Consumers have increasingly high requirements for product appearance, and products with excellent appearance quality are more likely to gain market favor. Therefore, ensuring that the appearance quality of products is defect-free has become a key link that discrete manufacturing enterprises must pay attention to.
[0003] For a long time, traditional product appearance quality detection has mainly relied on manual (quality inspectors) detection. Quality inspectors rely on naked eye observation and simple tool measurement to check the appearance of products item by item. However, this manual detection method has many drawbacks. On the one hand, manual detection is relatively inefficient, and in the case of large-scale production, it is difficult to meet the demand for rapid delivery; on the other hand, manual detection is prone to errors, and long-term repetitive work can cause visual fatigue of quality inspectors, greatly reducing the detection accuracy. At the same time, there are differences in subjective judgment standards between different quality inspectors, which increases the uncertainty of the detection results and seriously affects the stability of product quality.
[0004] With the continuous progress of science and technology, machine vision detection technology has emerged, bringing new solutions for appearance detection in discrete manufacturing. Machine vision detection uses industrial cameras, lenses, and image processing algorithms to quickly capture and analyze the appearance images of products. Compared with manual detection, machine vision detection has significantly improved in speed and can achieve automated batch detection.
[0005] However, the current machine vision detection method still has many shortcomings.
[0006] In a complex industrial production environment, single-modal visual detection faces serious challenges. Factors such as changes in light within the factory, vibration of equipment, and interference from the surrounding environment can all affect the accuracy of visual detection. Moreover, single-modal visual detection can only obtain information about one aspect of the appearance of the product, making it difficult to fully capture the defect information of the product. For some surface micro-cracks or internal defects, relying solely on visible light image detection may not be able to accurately identify them, limiting the detection accuracy.
[0007] In the context of increasingly stringent product quality requirements in modern manufacturing, single-modal visual inspection techniques have been unable to meet the needs of enterprises for high-precision detection of product appearance defects, and there is an urgent need for scientific and reasonable multi-modal visual inspection techniques to improve detection accuracy. SUMMARY
[0008] To solve the above technical problems, embodiments of the present application propose a discrete manufacturing product defect AI visual inspection method based on multi-modal fusion, which takes advantage of the complementary nature of multi-modal information, fully leverages the strengths of each modality, enhances the detection capability of micro-defects and internal defects, improves adaptability under complex lighting conditions, and realizes batched and automated detection, providing more reliable quality detection assurance for the discrete manufacturing industry.
[0009] To achieve the above purpose, embodiments of the present application propose a discrete manufacturing product defect AI (Artificial Intelligence) visual inspection method based on multi-modal fusion, which is suitable for appearance defect detection of discrete manufacturing products, comprising: acquiring time-aligned visible light images and thermal infrared images of a target product to be detected, performing photometric distortion correction on the visible light images, performing radiation calibration correction on the thermal infrared images, and then performing spatial registration on the photometric distortion corrected visible light images and the radiation calibration corrected thermal infrared images to obtain registered visible light images and registered thermal infrared images; performing feature level fusion based on multi-level DWT (Discrete Wavelet Transform) and decision level fusion based on DS (Dempster Shafer) evidence theory on the registered visible light images and the registered thermal infrared images to obtain a fusion image; inputting the fusion image into a pre-trained visual inspection model running on an embedded AI computing platform to obtain a detection result of the target product output by the visual inspection model; wherein the visual inspection model is composed of a dynamic attention gate module, a multi-scale feature aggregation module, a global attention fusion module and an output module, the dynamic attention gate module dynamically enhances the defect area of the fusion image based on a dynamic attention gate mechanism, the multi-scale feature aggregation module captures defect features of different scales based on the enhanced fusion image, the global attention fusion module performs global attention fusion on the defect features of different scales to obtain global fusion features, and the output module outputs the detection result of the target product based on the global fusion features.
[0010] To achieve the above object, the embodiment of the present application also proposes a discrete manufacturing product defect AI visual detection system based on multi-modal fusion, which is suitable for appearance defect detection of discrete manufacturing products, and comprises a visible light imaging module, a thermal infrared imaging module, a synchronous control module, a data preprocessing module, a multi-modal fusion module, a model construction module, a model training module, an embedded AI computing platform and a model use module; the synchronous control module is used for synchronously starting and controlling the visible light imaging module and the thermal infrared imaging module based on a timestamp synchronization mechanism to shoot a target product to be detected, so as to obtain time-aligned visible light images and thermal infrared images of the target product; the data preprocessing module is used for performing photometric distortion correction on the visible light images, performing radiation calibration correction on the thermal infrared images, and then performing spatial registration on the visible light images after photometric distortion correction and the thermal infrared images after radiation calibration correction, to obtain registered visible light images and registered thermal infrared images; the multi-modal fusion module is used for performing feature-level fusion based on multi-level DWT and decision-level fusion based on DS evidence theory on the registered visible light images and the registered thermal infrared images, to obtain a fusion image; the model construction module is used for constructing a visual detection model composed of a dynamic attention gate module, a multi-scale feature aggregation module, a global attention fusion module and an output module; the model training module is used for iteratively training the visual detection model based on an industrial defect detection standard dataset until convergence, to obtain a trained visual detection model; the embedded AI computing platform is used for running the trained visual detection model; and the model use module is used for inputting the fusion image into the trained visual detection model running on the embedded AI computing platform, to obtain a detection result of the target product output by the visual detection model; wherein the dynamic attention gate module dynamically enhances a defect area of the fusion image based on a dynamic attention gate mechanism, the multi-scale feature aggregation module captures defect features of different scales based on the enhanced fusion image, the global attention fusion module performs global attention fusion on the defect features of different scales to obtain global fusion features, and the output module outputs the detection result of the target product based on the global fusion features.
[0011] To achieve the above object, the embodiment of the present application also proposes an electronic device, which comprises at least one processor and a memory communicatively connected with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform a discrete manufacturing product defect AI visual detection method based on multi-modal fusion as described above.
[0012] To achieve the above object, the embodiment of the present application also proposes a computer readable storage medium, which stores a computer program, and the computer program can realize the above-mentioned multi-modal fusion based AI visual detection method for discrete manufacturing product defects when executed by a processor.
[0013] The multi-modal fusion based AI visual detection method for discrete manufacturing product defects proposed in the present application effectively integrates the complementary information of visible light images and thermal infrared images through multi-modal data fusion technology, significantly improves the detection capability for surface micro-defects and internal hidden defects, has higher comprehensive recognition accuracy compared with traditional single modal detection methods, and the introduction of dynamic attention mechanism enhances the focusing capability of the visual detection model on the defect area, and shows more stable detection performance under complex lighting conditions. The multi-modal fusion strategy, i.e. feature level fusion based on multi-level DWT and decision level fusion based on DS evidence theory, effectively improves the quality of multi-modal fusion, effectively overcomes the sensitivity of single modal to light changes and noise interference, and still maintains high reliability in complex environments in industrial sites. The visual detection model composed of dynamic attention gate module, multi-scale feature aggregation module, global attention fusion module and output module has very strong anti-interference ability, and has very strong robustness to image noise, mechanical vibration and other interference factors, and reduces the false detection rate. The multi-scale feature fusion mechanism set in the multi-scale feature aggregation module effectively enhances the capture ability for defect morphology diversity, which can cover a wider range of defect types. The embedded AI computing platform provides a hardware acceleration scheme for the running of the visual detection model, greatly reduces the computing resource demand, realizes real-time and efficient inference of edge devices, and improves the deployment flexibility. In summary, the present application realizes batch and automatic detection, and provides more reliable quality detection guarantee for the discrete manufacturing industry.
[0014] Optionally, the visible light image is obtained by a visible light imaging module, and the thermal infrared image is obtained by a thermal infrared imaging module. The visible light imaging module is equipped with a Sony IMX586 CMOS sensor, with a resolution of 4800*3200 and a frame rate of 90fps. The visible light imaging module is also integrated with a ring-shaped LED fill light system, and the wavelength adjustment range of the ring-shaped LED fill light system is 500nm to 700nm, and the illumination uniformity is greater than 92%. The thermal infrared imaging module is equipped with a FLIR Tau2 640 thermal imager, with a thermal sensitivity less than 30mK@30℃, a temperature measurement range of-20℃ to 350℃, and a polarization filter with an extinction ratio greater than 1000:1. The visible light image is subjected to photometric distortion correction, the thermal infrared image is subjected to radiometric calibration correction, and then the visible light image subjected to photometric distortion correction and the thermal infrared image subjected to radiometric calibration correction are subjected to spatial registration to obtain a registered visible light image and a registered thermal infrared image, comprising: The visible light image is subjected to photometric distortion correction by using a polynomial distortion model with an order greater than 5 to compensate for the radial distortion of the lens, to obtain a visible light image subjected to photometric distortion correction; A temperature-gray mapping relationship is established based on a blackbody radiation source, and the thermal infrared image is subjected to radiometric calibration correction by using the temperature-gray mapping relationship, to obtain a thermal infrared image subjected to radiometric calibration correction; The visible light image subjected to photometric distortion correction and the thermal infrared image subjected to radiometric calibration correction are subjected to sub-pixel level spatial registration based on a SURF feature matching algorithm, to obtain a registered visible light image and a registered thermal infrared image with a registration error less than 0.5 pixel.
[0015] Optionally, the registered visible light image and the registered thermal infrared image are subjected to feature level fusion based on multi-level DWT and decision level fusion based on DS evidence theory to obtain a fused image, comprising: Basic probability values of the visible light image and basic probability values of the thermal infrared image are calculated respectively based on the registered visible light image and the registered thermal infrared image, and a conflict measurement factor is calculated based on the basic probability values of the visible light image and the basic probability values of the thermal infrared image; It is judged whether the calculated conflict measurement factor is greater than a preset conflict measurement threshold value; If the calculated conflict measurement factor is greater than the conflict measurement threshold value, a decision level fusion weight is calculated based on the basic probability values of the visible light image, the basic probability values of the thermal infrared image and the calculated conflict measurement factor, and the registered visible light image and the registered thermal infrared image are subjected to decision level fusion weighting based on the decision level fusion weight to obtain a visible light image subjected to decision level fusion weighting and a thermal infrared image subjected to decision level fusion weighting; The visible light image subjected to decision level fusion weighting and the thermal infrared image subjected to decision level fusion weighting are subjected to multi-level DWT respectively, multi-level low frequency low frequency components, multi-level high frequency low frequency components and multi-level low frequency high frequency components are determined based on the multi-level DWT results, and the spliced results of the multi-level low frequency low frequency components, the multi-level high frequency low frequency components and the multi-level low frequency high frequency components are subjected to IDWT to obtain a fused image; If the calculated conflict measurement factor is less than or equal to the conflict measurement threshold value, the registered visible light image and the registered thermal infrared image are directly subjected to multi-level DWT respectively, multi-level low frequency low frequency components, multi-level high frequency low frequency components and multi-level low frequency high frequency components are determined based on the multi-level DWT results, and the spliced results of the multi-level low frequency low frequency components, the multi-level high frequency low frequency components and the multi-level low frequency high frequency components are subjected to IDWT to obtain a fused image.
[0016] Optionally, the conflict metric factor is calculated based on the visible light basic probability value and the thermal infrared basic probability value, and is implemented by the following formula: ; in, represents the registered visible light image, represents the thermal infrared image after registration, represents the basic probability value of visible light, represents the basic probability value of thermal infrared, represents the conflict measurement factor; Based on the basic probability value of visible light, the basic probability value of thermal infrared and the calculated conflict measurement factor, the decision-level fusion weight is calculated and implemented by the following formula: ; in, represents the decision-level fusion weight; Based on the decision-level fusion weight, the registered visible light image and the registered thermal infrared image are fused at the decision level to obtain the decision-level fusion weighted visible light image and the decision-level fusion weighted thermal infrared image, which is achieved by the following formula: ; ; in, represents the visible light image after decision-level fusion weighting, Represents the thermal infrared image after decision-level fusion weighting.
[0017] Optionally, the multi-level low-frequency low-frequency component, the multi-level high-frequency low-frequency component and the multi-level low-frequency high-frequency component are determined based on the multi-level DWT result, which is achieved by the following formula: ; ; ; ; in, Indicates the first Level DWT, is the total number of levels, 、 、 、 All are The dynamic weight coefficient of the level, represents the ReLU function, represents the Softmax function, 、 and respectively represent the low-frequency low-frequency component, the high-frequency low-frequency component and the low-frequency high-frequency component of the first level, respectively represent the low-frequency low-frequency component, the high-frequency low-frequency component and the low-frequency high-frequency component of the first level, the spliced results of the multi-level low-frequency low-frequency component, the multi-level high-frequency low-frequency component and the multi-level low-frequency high-frequency component are subjected to IDWT to obtain a fusion image, which is realized by the following formula: ; wherein, , and respectively represent the low-frequency low-frequency component, the high-frequency low-frequency component and the low-frequency high-frequency component of the first level, , and respectively represent the low-frequency low-frequency component, the high-frequency low-frequency component and the low-frequency high-frequency component of the first level, respectively represent the low-frequency low-frequency component, the high-frequency low-frequency component and the low-frequency high-frequency component of the first level, represents performing IDWT, represents a fusion image.
[0018] Optionally, the visual detection model is obtained by training through the following steps: loading an industrial defect detection standard dataset, the industrial defect detection standard dataset containing a plurality of product appearance images labeled with defect labels, the defect types including three types of scratches, bubbles and stains; randomly performing preprocessing including combined application of affine transformation, adding Gaussian noise and adding salt and pepper noise on each product appearance image in the industrial defect detection standard dataset, taking the preprocessed product appearance image as a training sample to form a training dataset; designing a total loss function based on cross-entropy loss, Dice loss and structural similarity loss; inputting the training sample in the training dataset into the visual detection model to obtain a detection result of the training sample output by the visual detection model, calculating a loss value based on the total loss function, the defect label corresponding to the training sample and the detection result of the training sample output by the visual detection model; based on the loss value, performing back propagation to update the network parameters of the visual detection model until the visual detection model is trained to convergence, obtaining a trained visual detection model; the total loss function is represented by the formula: ; wherein, is the total loss function, is the cross-entropy loss, is the Dice loss, is the structural similarity loss, , , respectively 、 、 Corresponding loss weight coefficients, the loss weight coefficients are searched based on a Bayesian optimization algorithm, and F1 score, micro defect recall rate and model convergence speed are used as evaluation indexes during the search.
[0019] Optionally, the embedded AI computing platform selects NVIDIA Jetson AGX Xavier with a computing power of 32 TOPS and a memory size of 32 GB, and realizes inference acceleration of the visual detection model through the TensorRT framework, including INT8 quantization of the visual detection model, reduction of inference delay rate, adoption of channel pruning and weight sharing mechanism, and reduction of the volume of the visual detection model, so as to achieve the purpose of memory optimization. BRIEF DESCRIPTION OF DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the drawings needed to be used in the description of the embodiments of the present application or the related art will be briefly introduced. Obviously, the following drawings are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor, and the drawings described herein are only used to explain the present application, and not to limit the present application.
[0021] Figure 1 is a flowchart of a multi-modal fusion-based discrete manufacturing product defect AI visual detection method provided in an embodiment of the present application; Figure 2 is a structural schematic diagram of a visual detection model provided in an embodiment of the present application; Figure 3 is a structural schematic diagram of a multi-modal fusion-based discrete manufacturing product defect AI visual detection system provided in another embodiment of the present application; Figure 4 is a structural schematic diagram of an electronic device provided in another embodiment of the present application. DETAILED DESCRIPTION
[0022] In order to make the purpose, technical solutions and advantages of the embodiments of the present application more clear, the embodiments of the present application will be described in detail below with reference to the drawings. Those skilled in the art can understand that in the embodiments of the present application, many technical details are proposed in order to make the reader better understand. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in the present application can be implemented. The following embodiments are divided for the purpose of description, and should not constitute any limitation on the specific implementation of the present application, and the following embodiments can be combined with each other under the premise of no contradiction.
[0023] Some common product defect detection methods in the field are introduced below.
[0024] Single visible light image-based machine vision detection technology is a common detection method. This technology uses an industrial camera to capture a visible light image of the product under specific lighting conditions, removes noise through preprocessing operations such as grayscale transformation and filtering, then extracts the image edge profile using edge detection algorithms such as the Canny algorithm, calculates the image gradient amplitude and direction, uses non-maximum suppression and double threshold processing to determine the edge, and finally compares the extracted edge with the standard template to determine whether the product has defects.
[0025] Single visible light image-based machine vision detection technology has the following shortcomings.
[0026] First, small defects are easily missed. Small defects have weak features in visible light images, making it difficult for edge detection methods such as the Canny algorithm to accurately extract their edge information, which can lead to missed detection.
[0027] Second, poor light adaptability. Changes in light intensity and angle can cause uneven reflection on the product surface, affecting image quality and interfering with edge detection and defect judgment.
[0028] Third, unable to detect internal defects. Visible light images can only reflect the information on the product surface, and cannot detect internal defects such as cracks and bubbles.
[0029] Single-modality image detection technology based on deep learning has been widely applied in recent years. Taking CNN-based detection as an example, first collect a large number of product appearance images as training data, label the images (defect type, location, etc.), and CNN automatically extracts image features through convolutional layers, pooling layers, and fully connected layers. The trained CNN can be used to detect defects in new product images.
[0030] Single-modality image detection technology based on deep learning has the following shortcomings.
[0031] First, data dependence is severe. The detection accuracy is highly dependent on the size and diversity of the training data. When the data is insufficient, the model's generalization ability is poor, making it difficult to accurately detect defects that have not been seen before.
[0032] Second, single modality limitation. Only using a single image modality, it cannot integrate other useful information, and the detection effect for complex defects is not good.
[0033] Third, large computational resource requirements. The structure of CNN is complex, and the training and inference process requires strong computational resource support, which limits its application in resource-constrained scenarios.
[0034] To solve the above technical problems, one embodiment of the present application proposes a discrete manufacturing product defect AI visual detection method based on multi-modal fusion, which is suitable for appearance defect detection of discrete manufacturing products and applied to an electronic device, which can be a terminal or a server. In this and the following embodiments, a server is taken as an example for illustration. The implementation details of the discrete manufacturing product defect AI visual detection method based on multi-modal fusion proposed in this embodiment are described below. The following content only provides related implementation details for easy understanding and is not necessary for implementing the present solution.
[0035] The specific process of the discrete manufacturing product defect AI visual detection method based on multi-modal fusion proposed in this embodiment can be as shown in Figure 1 The specific process of the discrete manufacturing product defect AI visual detection method based on multi-modal fusion proposed in this embodiment can be as shown in Step 11, acquiring time-aligned visible light images and thermal infrared images of the target product to be detected, performing photometric distortion correction on the visible light images, performing radiation calibration correction on the thermal infrared images, and then performing spatial registration on the photometric distortion corrected visible light images and the radiation calibration corrected thermal infrared images to obtain registered visible light images and registered thermal infrared images.
[0036] In a specific implementation, the server first needs to perform multi-modal data acquisition, that is, acquire time-aligned visible light images and thermal infrared images of the target product to be detected. Due to the reasons of the acquisition device, the visible light images and the thermal infrared images may have distortion and deviation, so it is necessary to perform photometric distortion correction on the visible light images, perform radiation calibration correction on the thermal infrared images, and then perform spatial registration on the photometric distortion corrected visible light images and the radiation calibration corrected thermal infrared images to obtain registered visible light images and registered thermal infrared images.
[0037] In one example, the visible light images are acquired by a visible light imaging module, and the thermal infrared images are acquired by a thermal infrared imaging module.
[0038] The visible light imaging module is equipped with a Sony IMX586 CMOS sensor with a resolution of 4800x3200 and a frame rate of 90fps. The visible light imaging module also integrates a ring-shaped LED fill light system, and the wavelength adjustment range of the ring-shaped LED fill light system is 500nm to 700nm, and the illumination uniformity is greater than 92%.
[0039] The thermal infrared imaging module is equipped with a FLIR Tau2 640 thermal imager with a thermal sensitivity less than 30mK@30℃ and a temperature measurement range of -20℃ to 350℃, and is equipped with a polarization filter with an extinction ratio greater than 1000:1.
[0040] The visible light imaging module and the thermal infrared imaging module are controlled by a synchronization control module, the synchronization control module is based on a timestamp synchronization mechanism of FPGA, and the visible light imaging module and the thermal infrared imaging module are started synchronously to capture the target product to be detected, so that the time-aligned visible light image and the thermal infrared image are obtained.
[0041] In one example, the server compensates the radial distortion of the lens by using a polynomial distortion model with an order greater than 5 to perform photometric distortion correction on the visible light image to obtain a visible light image after photometric distortion correction.
[0042] In one example, the server establishes a temperature-gray mapping relationship based on a blackbody radiation source, and performs radiation calibration correction on the thermal infrared image by using the temperature-gray mapping relationship to obtain a thermal infrared image after radiation calibration correction.
[0043] In one example, the server performs sub-pixel level spatial registration on the visible light image after photometric distortion correction and the thermal infrared image after radiation calibration correction based on a SURF feature matching algorithm to obtain a registered visible light image and a registered thermal infrared image with a registration error less than 0.5 pixel.
[0044] Step 12, performing feature level fusion based on multi-level DWT and decision level fusion based on DS evidence theory on the registered visible light image and the registered thermal infrared image to obtain a fusion image.
[0045] In a specific implementation, after the server completes the multi-modal data acquisition and obtains the registered visible light image and the registered thermal infrared image, the server can perform feature level fusion based on multi-level DWT and decision level fusion based on DS evidence theory on the registered visible light image and the registered thermal infrared image to obtain a fusion image.
[0046] In one example, during the entire multi-modal fusion process, the server needs to first determine whether to perform decision level fusion based on DS evidence theory, if needed, first perform decision level fusion based on DS evidence theory, and then perform feature level fusion based on multi-level DWT, if not needed, directly perform feature level fusion based on multi-level DWT.
[0047] In one example, the server first performs basic probability assignment on the registered visible light image and the registered thermal infrared image respectively to obtain visible light basic probability values and thermal infrared basic probability values, and calculates a conflict measurement factor based on the visible light basic probability values and the thermal infrared basic probability values.
[0048] Then, it is determined whether the calculated conflict measurement factor is greater than a preset conflict measurement threshold.
[0049] If the calculated conflict measure factor is greater than the conflict measure threshold, a decision-level fusion weight is calculated based on the visible light basic probability value, the thermal infrared basic probability value and the calculated conflict measure factor, and the registered visible light image and the registered thermal infrared image are fused at a decision level based on the decision-level fusion weight to obtain a visible light image fused at a decision level and a thermal infrared image fused at a decision level.
[0050] Next, the visible light image fused at a decision level and the thermal infrared image fused at a decision level are subjected to multi-level DWT respectively, multi-level low-frequency low-frequency components, multi-level high-frequency low-frequency components and multi-level low-frequency high-frequency components are determined based on the multi-level DWT results, and the spliced results of the multi-level low-frequency low-frequency components, the multi-level high-frequency low-frequency components and the multi-level low-frequency high-frequency components are subjected to IDWT to obtain a fused image.
[0051] If the calculated conflict measure factor is less than or equal to the conflict measure threshold, the registered visible light image and the registered thermal infrared image are directly subjected to multi-level DWT respectively, multi-level low-frequency low-frequency components, multi-level high-frequency low-frequency components and multi-level low-frequency high-frequency components are determined based on the multi-level DWT results, and the spliced results of the multi-level low-frequency low-frequency components, the multi-level high-frequency low-frequency components and the multi-level low-frequency high-frequency components are subjected to IDWT to obtain a fused image.
[0052] In one example, the preset conflict measure threshold is set to 0.7, that is, when the calculated conflict measure factor is greater than 0.7, decision-level fusion based on DS evidence theory is needed.
[0053] In one example, the conflict measure factor is calculated based on the visible light basic probability value and the thermal infrared basic probability value, and is realized by the following formula: ; Wherein, represents the registered visible light image, represents the registered thermal infrared image, represents the visible light basic probability value, represents the thermal infrared basic probability value, represents the conflict measure factor.
[0054] In one example, the decision-level fusion weight is calculated based on the visible light basic probability value, the thermal infrared basic probability value and the calculated conflict measure factor, and is realized by the following formula: ; Wherein, represents the decision-level fusion weight.
[0055] In one example, the registered visible light image and the registered thermal infrared image are decision-level fusion weighted based on a decision-level fusion weight, to obtain a decision-level fusion weighted visible light image and a decision-level fusion weighted thermal infrared image, which is realized by the following formula: ; ;
[0056] In one example, the multi-level low-frequency low-frequency component, the multi-level high-frequency low-frequency component and the multi-level low-frequency high-frequency component are determined based on the multi-level DWT result, which is realized by the following formula: ; ; ; ;
[0057] In one example, the multi-level low-frequency low-frequency component, the multi-level high-frequency low-frequency component and the multi-level low-frequency high-frequency component are determined based on the multi-level DWT result, which is realized by the following formula: ; The fusion image is input to a pre-trained visual detection model running on an embedded AI computing platform to obtain a detection result of the target product output by the visual detection model.
[0058] At step 13, the fusion image is input to a pre-trained visual detection model running on an embedded AI computing platform to obtain a detection result of the target product output by the visual detection model. The visual detection model is composed of a dynamic attention gate module, a multi-scale feature aggregation module, a global attention fusion module, and an output module. The dynamic attention gate module dynamically enhances the defect area of the fusion image based on a dynamic attention gate mechanism. The multi-scale feature aggregation module captures defect features of different scales based on the enhanced fusion image. The global attention fusion module globally fuses defect features of different scales to obtain global fusion features. The output module outputs the detection result of the target product based on the global fusion features.
[0059] In a specific implementation, after obtaining the fusion image, the server can input the fusion image to a pre-trained visual detection model running on an embedded AI computing platform to obtain a detection result of the target product output by the visual detection model. The visual detection model is composed of a dynamic attention gate module, a multi-scale feature aggregation module, a global attention fusion module, and an output module. The dynamic attention gate module dynamically enhances the defect area of the fusion image based on a dynamic attention gate mechanism. The multi-scale feature aggregation module captures defect features of different scales based on the enhanced fusion image. The global attention fusion module globally fuses defect features of different scales to obtain global fusion features. The output module outputs the detection result of the target product based on the global fusion features.
[0060] In one example, the specific structure of the visual detection model can be as shown in Figure 2 The dynamic attention gate module calculates spatial attention and channel attention through parallel convolution flow to achieve dynamic enhancement of the defect area. The multi-scale feature aggregation module uses an atrous spatial pyramid pooling (ASPP) with dilation rates of 6, 12, and 18 to capture defect features of different scales.
[0061] In one example, the visual detection model is trained by the following steps.
[0062] First, load the industrial defect detection standard dataset (MVTec AD Dataset). The industrial defect detection standard dataset contains a plurality of product appearance images labeled with defect labels. The defect types included in the industrial defect detection standard dataset include scratches, bubbles, and stains.
[0063] Next, each product appearance image in the standard industrial defect detection dataset was randomly preprocessed using a combination of affine transformations, Gaussian noise, and salt-and-pepper noise. These preprocessed product appearance images served as training samples to form the training dataset. The affine transformations applied included rotation and scaling, with rotation angles ranging from -15° to 15° and scaling factors ranging from 0.8x to 1.2x.
[0064] Subsequently, the overall loss function is designed based on cross entropy loss, Dice loss and structural similarity loss.
[0065] After that, the training samples in the training data set are input into the visual inspection model to obtain the detection results of the training samples output by the visual inspection model. The loss value is calculated based on the overall loss function, the defect labels corresponding to the training samples and the detection results of the training samples output by the visual inspection model.
[0066] Finally, back propagation is performed based on the loss value to update the network parameters of the visual detection model until the visual detection model is trained to convergence, and a trained visual detection model is obtained.
[0067] In one example, the overall loss function is formulated as: ; in, is the overall loss function, is the cross entropy loss, is the Dice loss, is the structural similarity loss, 、 、 They are 、 、 The corresponding loss weight coefficient is obtained based on the Bayesian optimization algorithm search. During the search, the F1 score, minor defect recall rate and model convergence speed are used as evaluation indicators.
[0068] In one example, the optimal loss weight coefficient 、 、 , which can be determined through simulation experiments, that is, based on the Bayesian optimization algorithm search. The main evaluation indicator in the search process is the F1 score, and the auxiliary evaluation indicators are the minor defect recall rate and the model convergence speed.
[0069] When searching, you need to 、 、 Set the value range, search step and constraints.
[0070] the value range of is 0.5, and the constraint condition is .
[0071] the value range of is 0.5, and the constraint condition is nonlinearly related to .
[0072] the value range of is 0.2, and the constraint condition is to suppress the overfitting risk of structural similarity loss.
[0073] When searching, first, initialization is performed, that is, 10 groups of loss weight coefficient combinations are randomly sampled, Latin hypercube sampling can be used, then 5 rounds of searching are performed, the FI score needs to be calculated in each round, and the surrogate model (Gaussian process) is updated, and finally the loss weight coefficient combination with the highest F1 score is taken.
[0074] Through 5 independent repeated searches (with fixed random seeds), the optimal loss weight coefficient combination finally determined is: , , .
[0075] In an example, the embedded AI computing platform selects the NVIDIA Jetson AGX Xavier with a computing power of 32 TOPS and a memory size of 32 GB, and realizes inference acceleration of the visual detection model through the TensorRT framework, including INT8 quantization of the visual detection model, reduction of inference delay rate, adoption of channel pruning and weight sharing mechanism, and reduction of the volume of the visual detection model, so as to achieve the purpose of memory optimization.
[0076] In an example, in addition to the improvement in hardware, the embodiment also innovatively improves the software design. An online incremental learning mechanism is set, real-time uploading and labeling of defect samples are supported, a knowledge distillation technology is adopted to realize online updating of the model, which significantly shortens the updating cycle, a dynamic resource allocation mechanism is set, and computing power resources are automatically allocated according to the priority of the detection task (such as 60% of the computing power for micro defect detection), to ensure the real-time performance of key tasks. A cross-platform deployment mechanism is set, Docker containerized deployment is supported, and Linux / Windows industrial control systems can be compatible.
[0077] The embodiment proposes a discrete manufacturing product defect AI visual detection method based on multi-modal fusion. Through multi-modal data fusion technology, the complementary information of visible light images and thermal infrared images is effectively integrated, and the detection capability of surface micro-defects and internal hidden defects is significantly improved. Compared with the traditional single modal detection method, the comprehensive recognition accuracy is higher. The introduction of dynamic attention mechanism enhances the focusing ability of the visual detection model on the defect area, and shows more stable detection performance under complex lighting conditions. The multi-modal fusion strategy, i.e. feature-level fusion based on multi-level DWT and decision-level fusion based on DS evidence theory, effectively improves the quality of multi-modal fusion, and effectively overcomes the sensitivity of single modal to light changes and noise interference. In the complex environment of industrial field, it can still maintain high reliability. The visual detection model composed of dynamic attention gate module, multi-scale feature aggregation module, global attention fusion module and output module has very strong anti-interference ability, and has strong robustness to image noise, mechanical vibration and other interference factors, reducing the false detection rate. The multi-scale feature fusion mechanism set in the multi-scale feature aggregation module effectively enhances the capture ability of defect morphology diversity, which can cover a wider range of defect types. The embedded AI computing platform provides a hardware acceleration scheme for the operation of the visual detection model, greatly reducing the computing resource demand, realizing real-time and efficient inference of edge devices, and improving the deployment flexibility. In summary, the embodiment realizes batch and automatic detection, and provides more reliable quality detection guarantee for the discrete manufacturing industry.
[0078] The step division of the above method is only for the purpose of clear description, and when implemented, it can be combined into one step or some steps can be divided into multiple steps, as long as the same logical relationship is included, and it is within the protection scope of the present application. Adding insignificant modifications or introducing insignificant designs in the algorithm or process does not change the core design of the algorithm and process, and is within the protection scope of the present application.
[0079] Another embodiment of the present application proposes a discrete manufacturing product defect AI visual detection system based on multi-modal fusion, which is suitable for appearance defect detection of discrete manufacturing products. The details of the discrete manufacturing product defect AI visual detection system based on multi-modal fusion proposed in the embodiment will be specifically explained below. The following content only provides implementation details for easy understanding, and is not necessary for implementing the embodiment.
[0080] The specific structure of the discrete manufacturing product defect AI visual detection system based on multi-modal fusion proposed in the embodiment can be as follows Figure 3As shown, it comprises: a visible light imaging module 21, a thermal infrared imaging module 22, a synchronous control module 23, a data preprocessing module 24, a multi-modal fusion module 25, a model construction module 26, a model training module 27, an embedded AI computing platform 28 and a model use module 29.
[0081] The synchronous control module 23 is used to synchronize the start and control of the visible light imaging module 21 and the thermal infrared imaging module 22 based on a timestamp synchronization mechanism to take pictures of the target product to be detected, thereby obtaining time-aligned visible light images and thermal infrared images of the target product.
[0082] The data preprocessing module 24 is used to perform photometric distortion correction on the visible light images, radiation calibration correction on the thermal infrared images, and spatial registration on the photometric distortion corrected visible light images and the radiation calibration corrected thermal infrared images, to obtain registered visible light images and registered thermal infrared images.
[0083] The multi-modal fusion module 25 is used to perform feature level fusion based on multi-level DWT and decision level fusion based on DS evidence theory on the registered visible light images and the registered thermal infrared images, to obtain a fusion image.
[0084] The model construction module 26 is used to construct a visual detection model composed of a dynamic attention gate module, a multi-scale feature aggregation module, a global attention fusion module and an output module.
[0085] The model training module 27 is used to iteratively train the visual detection model based on an industrial defect detection standard dataset until convergence, to obtain a trained visual detection model.
[0086] The embedded AI computing platform 28 is used to run the trained visual detection model.
[0087] The model use module 29 is used to input the fusion image into the trained visual detection model running on the embedded AI computing platform to obtain the detection result of the target product output by the visual detection model; wherein the dynamic attention gate module dynamically enhances the defect area of the fusion image based on a dynamic attention gate mechanism, the multi-scale feature aggregation module captures defect features of different scales based on the enhanced fusion image, the global attention fusion module performs global attention fusion on the defect features of different scales to obtain global fusion features, and the output module outputs the detection result of the target product based on the global fusion features.
[0088] It can be found that the embodiment is a system embodiment corresponding to the method embodiment described above, and the embodiment can be implemented in cooperation with the method embodiment described above. The related technical details and technical effects mentioned in the method embodiment are still valid in the embodiment. In order to reduce repetition, they will not be described here. Accordingly, the related technical details mentioned in the embodiment can also be applied to the method embodiment.
[0089] It is worth mentioning that each module and module involved in the embodiment is a logical module. In actual application, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. In addition, in order to highlight the innovative part of the present application, units not closely related to solving the technical problems proposed in the present application are not introduced in the embodiment. However, this does not mean that there are no other units in the embodiment.
[0090] Another embodiment of the present application provides an electronic device, as shown in Figure 4 The electronic device includes at least one processor 31 and a memory 32 connected with the at least one processor 31. The memory 32 stores instructions executable by the at least one processor 31. The instructions are executed by the at least one processor 31 to enable the at least one processor 31 to perform a method for AI visual detection of discrete manufacturing product defects based on multi-modal fusion as described in the method embodiment.
[0091] The memory and the processor are connected in a bus mode. The bus includes any number of interconnected buses and bridges. The bus connects various circuits of one or more processors and memories together. The bus can also connect various other circuits such as peripheral devices, voltage regulators and power management circuits, etc. together, which are well known in the art, and therefore will not be further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be one element or multiple elements such as multiple receivers and transmitters, which provide a unit for communicating with various other devices on a transmission medium. Data processed by the processor is transmitted on a wireless medium through the antenna, and further, the antenna also receives data and transmits the data to the processor.
[0092] The processor is responsible for managing the bus and general processing, and can also provide various functions including timing, peripheral interface, voltage regulation, power management and other control functions. The memory can be used to store data used by the processor during execution.
[0093] Another embodiment of the present application provides a computer readable storage medium, wherein a computer program is stored in the computer readable storage medium, and the computer program, when executed by a processor, enables a method for AI visual detection of discrete manufacturing product defects based on multi-modal fusion to be implemented.
[0094] That is, a person skilled in the art can understand that all or part of the steps in the above method embodiments can be completed by programs instructing related hardware, and the programs are stored in a storage medium and include a plurality of instructions for enabling a device (such as a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the method described in the method embodiments of the present application. The foregoing storage medium includes a U disk, a mobile hard disk, a read-only memory, a random access memory, a magnetic disk or an optical disk, and various media that can store program codes.
[0095] A person skilled in the art can understand that each of the above embodiments is a specific embodiment for implementing the present application, and in actual application, various changes can be made in form and details without departing from the spirit and scope of the present application. For those skilled in the art, a number of improvements and refinements can be made without departing from the principles of the present application, and these improvements and refinements are also considered to be within the protection scope of the present application.
Claims
1. An AI visual inspection method for discrete manufacturing product defects based on multimodal fusion, suitable for appearance defect detection of discrete manufacturing products, characterized by: The method comprises: Obtain time-aligned visible light images and thermal infrared images of the target product to be inspected, perform photometric distortion correction on the visible light image, perform radiometric calibration correction on the thermal infrared image, and then perform spatial registration on the photometric distortion-corrected visible light image and the radiometric calibration-corrected thermal infrared image to obtain a registered visible light image and a registered thermal infrared image; The registered visible light image and the registered thermal infrared image are subjected to feature-level fusion based on multi-level DWT and decision-level fusion based on DS evidence theory to obtain a fused image. The fused image is input into a pre-trained visual inspection model running on an embedded AI computing platform to obtain the detection result of the target product output by the visual inspection model; wherein, the visual inspection model consists of a dynamic attention gating module, a multi-scale feature aggregation module, a global attention fusion module and an output module. The dynamic attention gating module dynamically enhances the defect area of the fused image based on the dynamic attention gating mechanism. The multi-scale feature aggregation module captures defect features of different scales based on the enhanced fused image. The global attention fusion module performs global attention fusion on defect features of different scales to obtain global fusion features. The output module outputs the detection result of the target product based on the global fusion features.
2. The method for AI visual detection of discrete manufacturing product defects based on multimodal fusion according to claim 1 is characterized in that: Visible light images are acquired by the visible light imaging module, and thermal infrared images are acquired by the thermal infrared imaging module; The visible light imaging module is equipped with a Sony IMX586 CMOS sensor with a resolution of 4800×3200 and a frame rate of 90fps. It also integrates a ring LED fill light system with an adjustable wavelength range of 500nm to 700nm and an illumination uniformity greater than 92%. The thermal infrared imaging module is equipped with a FLIR Tau2 640 thermal imager with a thermal sensitivity of less than 30mK @ 30°C, a temperature measurement range of -20°C to 350°C, and a polarizing filter with an extinction ratio greater than 1000:
1. Perform photometric distortion correction on the visible light image, perform radiometric calibration correction on the thermal infrared image, and then perform spatial registration on the photometric distortion corrected visible light image and the radiometric calibration corrected thermal infrared image to obtain the registered visible light image and the registered thermal infrared image, including: A polynomial distortion model with an order greater than 5 is used to compensate for the radial distortion of the lens to correct the photometric distortion of the visible light image, thereby obtaining a visible light image after photometric distortion correction. A temperature-grayscale mapping relationship is established based on a blackbody radiation source, and the thermal infrared image is subjected to radiation calibration correction using the temperature-grayscale mapping relationship to obtain a thermal infrared image after radiation calibration correction; Based on the SURF feature matching algorithm, sub-pixel spatial registration is performed on the visible light image after photometric distortion correction and the thermal infrared image after radiometric calibration correction, and the registered visible light image and the registered thermal infrared image are obtained with a registration error less than 0.5 pixel.
3. The method for AI visual detection of discrete manufacturing product defects based on multimodal fusion according to claim 1 is characterized in that: The registered visible light image and the registered thermal infrared image are subjected to feature-level fusion based on multi-level DWT and decision-level fusion based on DS evidence theory to obtain a fused image, including: The registered visible light image and the registered thermal infrared image are assigned basic probabilities respectively to obtain a visible light basic probability value and a thermal infrared basic probability value, and a conflict measurement factor is calculated based on the visible light basic probability value and the thermal infrared basic probability value; Determine whether the calculated conflict metric factor is greater than a preset conflict metric threshold; If the calculated conflict measurement factor is greater than the conflict measurement threshold, the decision-level fusion weight is calculated based on the visible light basic probability value, the thermal infrared basic probability value and the calculated conflict measurement factor, and the registered visible light image and the registered thermal infrared image are subjected to decision-level fusion weighting based on the decision-level fusion weight to obtain a decision-level fusion weighted visible light image and a decision-level fusion weighted thermal infrared image; Perform multi-level DWT on the weighted visible light image and the weighted thermal infrared image after decision-level fusion, determine the multi-level low-frequency low-frequency component, the multi-level high-frequency low-frequency component and the multi-level low-frequency high-frequency component based on the multi-level DWT results, and then perform IDWT on the splicing results of the multi-level low-frequency low-frequency component, the multi-level high-frequency low-frequency component and the multi-level low-frequency high-frequency component to obtain the fused image; If the calculated conflict metric factor is less than or equal to the conflict metric threshold, the multi-level DWT is directly performed on the registered visible light image and the registered thermal infrared image respectively. The multi-level low-frequency low-frequency component, the multi-level high-frequency low-frequency component and the multi-level low-frequency high-frequency component are determined based on the multi-level DWT results. Then, the IDWT is performed on the splicing results of the multi-level low-frequency low-frequency component, the multi-level high-frequency low-frequency component and the multi-level low-frequency high-frequency component to obtain the fused image.
4. The method for AI visual detection of discrete manufacturing product defects based on multimodal fusion according to claim 3 is characterized in that: The conflict measurement factor is calculated based on the visible light basic probability value and the thermal infrared basic probability value, and is implemented by the following formula: ; in, represents the registered visible light image, represents the thermal infrared image after registration, represents the basic probability value of visible light, represents the basic probability value of thermal infrared, represents the conflict measurement factor; Based on the basic probability value of visible light, the basic probability value of thermal infrared and the calculated conflict measurement factor, the decision-level fusion weight is calculated and implemented by the following formula: ; in, represents the decision-level fusion weight; Based on the decision-level fusion weight, the registered visible light image and the registered thermal infrared image are fused at the decision level to obtain the decision-level fusion weighted visible light image and the decision-level fusion weighted thermal infrared image, which is achieved by the following formula: ; ; in, represents the visible light image after decision-level fusion weighting, Represents the thermal infrared image after decision-level fusion weighting.
5. The method for AI visual detection of discrete manufacturing product defects based on multimodal fusion according to claim 4 is characterized in that: Based on the multi-level DWT results, the multi-level low-frequency components, multi-level high-frequency components and multi-level low-frequency components are determined by the following formula: ; ; ; ; in, Indicates the first Level DWT, is the total number of levels, 、 、 、 All are The dynamic weight coefficient of the level, represents the ReLU function, represents the Softmax function, 、 and Respectively represent The low-frequency low-frequency component, high-frequency low-frequency component and low-frequency high-frequency component of the level; IDWT is performed on the splicing results of multi-level low-frequency low-frequency components, multi-level high-frequency low-frequency components and multi-level low-frequency high-frequency components to obtain a fused image, which is achieved by the following formula: ; in, 、 and Represent the low-frequency low-frequency component, high-frequency low-frequency component and low-frequency high-frequency component of the first level respectively. 、 and Respectively represent The low-frequency low-frequency component, high-frequency low-frequency component and low-frequency high-frequency component of the level, Indicates IDWT. Represents the fused image.
6. The method for AI visual inspection of discrete manufacturing product defects based on multimodal fusion according to claim 1, characterized in that: The visual detection model is trained through the following steps: Load the standard industrial defect detection dataset, which contains several product appearance images labeled with defects. The defect types include scratches, bubbles, and stains. Each product appearance image in the industrial defect detection standard dataset is randomly preprocessed, including the combined application of affine transformation, the addition of Gaussian noise, and the addition of salt and pepper noise. The preprocessed product appearance images are used as training samples to form a training dataset. Design the overall loss function based on cross entropy loss, Dice loss and structural similarity loss; Input the training samples in the training data set into the visual inspection model, obtain the detection results of the training samples output by the visual inspection model, and calculate the loss value based on the overall loss function, the defect labels corresponding to the training samples, and the detection results of the training samples output by the visual inspection model; Backpropagation is performed based on the loss value to update the network parameters of the visual detection model until the visual detection model is trained to convergence, thereby obtaining a trained visual detection model; The overall loss function is expressed as: ; in, is the overall loss function, is the cross entropy loss, is the Dice loss, is the structural similarity loss, 、 、 They are 、 、 The corresponding loss weight coefficient is obtained based on the Bayesian optimization algorithm search. During the search, the F1 score, minor defect recall rate and model convergence speed are used as evaluation indicators.
7. The AI visual detection method for discrete manufacturing product defects based on multimodal fusion according to any one of claims 1 to 6, characterized in that: The embedded AI computing platform uses the NVIDIA Jetson AGX Xavier with a computing power of 32TOPS and a memory size of 32GB. It uses the TensorRT framework to accelerate the inference of the visual detection model, including INT8 quantization of the visual detection model to reduce the inference latency rate, and adopts channel pruning and weight sharing mechanisms to reduce the size of the visual detection model, thereby achieving the purpose of memory optimization.
8. An AI visual inspection system for discrete manufacturing product defects based on multimodal fusion, suitable for appearance defect detection of discrete manufacturing products, characterized by: The system includes: a visible light imaging module, a thermal infrared imaging module, a synchronization control module, a data preprocessing module, a multimodal fusion module, a model building module, a model training module, an embedded AI computing platform and a model use module; A synchronization control module is used to synchronously start and control the visible light imaging module and the thermal infrared imaging module based on a timestamp synchronization mechanism to capture the target product to be inspected, thereby obtaining time-aligned visible light images and thermal infrared images of the target product; A data preprocessing module is used to perform photometric distortion correction on the visible light image, perform radiometric calibration correction on the thermal infrared image, and then perform spatial registration on the photometric distortion corrected visible light image and the radiometric calibration corrected thermal infrared image to obtain a registered visible light image and a registered thermal infrared image; The multimodal fusion module is used to perform feature-level fusion based on multi-level DWT and decision-level fusion based on DS evidence theory on the registered visible light image and the registered thermal infrared image to obtain a fused image; The model building module is used to build a visual detection model consisting of a dynamic attention gating module, a multi-scale feature aggregation module, a global attention fusion module, and an output module; The model training module is used to iteratively train the visual inspection model based on the industrial defect detection standard dataset until convergence, thereby obtaining a trained visual inspection model; Embedded AI computing platform for running trained visual inspection models; The model uses a module to input the fused image into a trained visual inspection model running on an embedded AI computing platform to obtain the detection results of the target product output by the visual inspection model; among them, the dynamic attention gating module dynamically enhances the defect area of the fused image based on the dynamic attention gating mechanism, the multi-scale feature aggregation module captures defect features of different scales based on the enhanced fused image, the global attention fusion module performs global attention fusion on defect features of different scales to obtain global fusion features, and the output module outputs the detection results of the target product based on the global fusion features.
9. An electronic device, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute an AI visual detection method for discrete manufacturing product defects based on multimodal fusion as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it can implement an AI visual detection method for discrete manufacturing product defects based on multimodal fusion as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Electronic component surface defect detection method and device based on multispectral fusion
CN117191816A
Visible light and infrared image fusion target detection method based on deep learning
CN118887453A
Multi-spectral imaging photovoltaic defect identification method and system based on deep learning
CN120375180A
A fabric defect detection method based on multi-modal deep learning
US20220414856A1
Cited By
Physical and chemical enhancement control method and system for multi-source low-quality hazardous waste-based artificial stone
CN121052809A
Galvanometer optical system lens state online monitoring method and device based on AI vision
CN121598323A
Mirror optical system lens state online monitoring method and device based on AI vision
CN121598323B
Ironing sole plate quality detection method and system based on image processing
CN122468730A