Unsupervised Learning Tablet Defect Detection Method and Related Devices Based on Diffusion Model
Through an unsupervised learning method based on diffusion model, the diffusion model and auxiliary reconstruction network are used to reconstruct tablet image and feature extraction, which solves the problems of poor reconstruction effect and low detection accuracy in tablet defect detection, and achieves more efficient tablet defect detection.
Patent Information
- Application Number
- CN202510650415.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-20
AI Technical Summary
The existing tablet defect detection methods have poor reconstruction effect and low detection accuracy. Especially when there are fewer defect samples and diverse forms during the tablet production process, traditional methods are difficult to meet the requirements of accuracy and real-time.
Unsupervised learning method based on diffusion model is adopted, and the diffusion model denoising network and auxiliary reconstruction network are constructed. The semantic information of the pill image is extracted using the auxiliary reconstruction network, and the auxiliary diffusion model denoising network is reconstructed. The residual network and feature pyramid network are combined for feature extraction, the cosine distance and Euclidean distance are calculated, and the abnormal scores are weighted for detection.
It improves the reconstruction quality and accuracy of tablet defect detection, can better identify tablet defects, adapt to the diversity and complexity in the tablet production process, and improves the effectiveness and accuracy of the detection.
Smart Images

Figure CN120182256B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of tablet defect detection, and specifically relates to an unsupervised learning tablet defect detection method and related device based on a diffusion model. Background Art
[0002] The quality of drugs is directly related to the health and safety of patients and the brand reputation of pharmaceutical companies. In recent years, drug safety problems have occurred frequently, and the number of drug safety accidents has continued to rise. Among them, the problem of drug surface defects is an important part, which is not only a key quality problem that needs to be detected urgently in the production process, but also the most intuitive manifestation of the affected product quality. Therefore, detecting drug surface defects is of great significance for ensuring drug quality, improving drug use safety, and reducing the frequency of accidents.
[0003] Tablets are the most common form of drugs. Tablet defect detection generally aims to detect defects such as scratches, defects, foreign object occlusion, color contamination, and holes on the tablet surface. Traditional manual visual inspection methods are inefficient and subjective, and it is difficult to meet the requirements of high-speed production lines for accuracy and real-time performance. Traditional image processing techniques achieve anomaly recognition through low-dimensional feature extraction and rule design, but they have insufficient generalization ability for complex texture defects and are easily affected by illumination changes and background interference.
[0004] In recent years, deep learning techniques have significantly improved the accuracy and robustness of tablet defect detection with their powerful feature learning ability and have gradually become the mainstream method. Among them, although the supervised method has high accuracy and good adaptability, this method requires a large amount of labeled defect data and cannot cope with the tablet production process, where the defect samples generated during tablet production are few and the defect forms are diverse.
[0005] Therefore, it is more appropriate to use unsupervised methods such as DRAEM (discriminatively trained reconstruction anomaly embedding model), PatchCore (Patch-based Core-set Sampling for Anomaly Detection), and LDM (Latent Diffusion Model). Among them, the reconstruction-based method is a major research hotspot. The core of the reconstruction-based method is that in the training stage, the model only learns the feature distribution from normal images, and in the testing stage, the trained model reconstructs abnormal images into normal images, so as to determine the abnormal position by comparing the reconstructed images with the input images.
[0006] Currently, the existing tablet defect detection methods have poor reconstruction effects and low detection accuracy. Summary of the Invention
[0007] The object of the present invention is to provide an unsupervised learning tablet defect detection method and related device based on a diffusion model, which are used to solve the problems of poor reconstruction effect and low detection accuracy in the prior art.
[0008] To achieve the above object, the present invention adopts the following technical solutions:
[0009] In a first aspect, the present invention provides an unsupervised learning tablet defect detection method based on a diffusion model, including the following steps:
[0010] Obtain the original tablet image;
[0011] Encode the original tablet image to obtain a latent space vector;
[0012] Construct a diffusion model denoising network and an auxiliary reconstruction network. First, use the auxiliary reconstruction network to extract the semantic information of the original tablet image, and then use the extracted semantic information of the original tablet image to assist the diffusion model denoising network to reconstruct the obtained latent space vector to obtain a reconstructed tablet image;
[0013] Extract features from the original tablet image and the reconstructed tablet image respectively to obtain the original tablet image features and the reconstructed tablet image features;
[0014] Use the original tablet image features and the reconstructed tablet image features to calculate the cosine distance and Euclidean distance at different scales, where the scale refers to the size of the original tablet image features and the reconstructed tablet image features;
[0015] Weight the cosine distance and Euclidean distance at different scales to obtain an anomaly score;
[0016] Detect the tablet defects according to the anomaly score to obtain the tablet defect detection result.
[0017] A further improvement of the present invention is that in the step of constructing a diffusion model denoising network and an auxiliary reconstruction network, first using the auxiliary reconstruction network to extract the semantic information of the original tablet image, and then using the extracted semantic information of the original tablet image to assist the diffusion model denoising network to reconstruct the obtained latent space vector to obtain a reconstructed tablet image, specifically includes:
[0018] Perform noise addition processing on the obtained latent space vector to obtain a noise-added latent space vector;
[0019] Construct a diffusion model denoising network and an auxiliary reconstruction network;
[0020] Input the noisy latent space vector into the denoising network of the diffusion model for denoising to obtain the denoised latent space vector;
[0021] Input the original pill image and the noisy latent space vector into the auxiliary reconstruction network. First, use the auxiliary reconstruction network to extract the semantic information of the original pill image, and then use the extracted semantic information of the original pill image to assist the diffusion model denoising network in reconstructing the denoised latent space vector to obtain the reconstructed latent space vector;
[0022] Perform decoding processing on the reconstructed latent space vector to obtain the reconstructed pill image.
[0023] A further improvement of the present invention is that the original pill image and the reconstructed pill image are respectively subjected to feature extraction to obtain the original pill image features and the reconstructed pill image features, which specifically includes:
[0024] Use a residual network to respectively extract features from the original pill image and the reconstructed pill image to obtain a number of initially original pill image features and initially reconstructed pill image features of different scales;
[0025] Use a feature pyramid network to respectively fuse a number of initially original pill image features and initially reconstructed pill image features of different scales to obtain a number of finally original pill image features and finally reconstructed pill image features of different scales.
[0026] A further improvement of the present invention is that in the step of using a residual network to respectively extract features from the original pill image and the reconstructed pill image to obtain a number of initially original pill image features and initially reconstructed pill image features of different scales, a convolutional attention module is also added to each residual block of the residual network.
[0027] A further improvement of the present invention is that in the step of calculating the cosine distance and Euclidean distance of different scales using the original pill image features and the reconstructed pill image features, where the scale refers to the size of the original pill image features and the reconstructed pill image features, the calculation formula for the cosine distance of different scales is:
[0028]
[0029] Where, is the cosine distance of the i th scale, is the cosine similarity of the i th scale, is the original pill image feature, is the reconstructed pill image feature, iis the number of scales of the original tablet image features and the reconstructed tablet image features;
[0030] The calculation formula for the Euclidean distance at different scales is:
[0031]
[0032] where, is the Euclidean distance at the i th scale, is the original tablet image feature, is the reconstructed tablet image feature, is the original tablet image feature of the j th element value, is the reconstructed tablet image feature of the j th element value, i is the number of scales of the original tablet image features and the reconstructed tablet image features, n is the original tablet image feature contains the number of elements, j is an index variable used to traverse the element values in the original tablet image feature and the reconstructed tablet image feature ; the value range of j is [1, n .
[0033] A further improvement of the present invention lies in that in the step of weighting the cosine distance and the Euclidean distance at different scales to obtain an anomaly score, the calculation formula for the anomaly score is:
[0034]
[0035] where, represents the anomaly score, is the upsampling factor at the i th scale, is the Euclidean distance at the i th scale, is the cosine distance at the i th scale, i is the number of scales of the original tablet image features and the reconstructed tablet image features, n is the original tablet image feature contains the number of elements, is the weight coefficient, the value of
[0036] A further improvement of the present invention lies in that the detection of tablet defects based on the anomaly score to obtain the tablet defect detection result is specifically as follows:
[0037] Compare the anomaly score with the set anomaly threshold. When the anomaly score is greater than the set anomaly threshold, it indicates that the tablet has defects; when the anomaly score is less than or equal to the set anomaly threshold, it indicates that the tablet has no defects.
[0038] In a second aspect, the present invention provides an unsupervised learning tablet defect detection system based on a diffusion model, including a data acquisition module, a latent space vector acquisition module, an image reconstruction module, a feature extraction module, a distance determination module, an anomaly score determination module, and a tablet defect detection module;
[0039] The data acquisition module is used to acquire the original tablet image;
[0040] The latent space vector acquisition module is used to perform encoding processing on the original tablet image to obtain a latent space vector;
[0041] The image reconstruction module is used to construct a diffusion model denoising network and an auxiliary reconstruction network. First, use the auxiliary reconstruction network to extract the semantic information of the original tablet image, and then use the extracted semantic information of the original tablet image to assist the diffusion model denoising network to reconstruct the obtained latent space vector to obtain a reconstructed tablet image;
[0042] The feature extraction module is used to perform feature extraction on the original tablet image and the reconstructed tablet image respectively to obtain the original tablet image features and the reconstructed tablet image features;
[0043] The distance determination module is used to calculate the cosine distance and Euclidean distance at different scales by using the original tablet image features and the reconstructed tablet image features, where the scale refers to the size of the original tablet image features and the reconstructed tablet image features;
[0044] The anomaly score determination module is used to weight the cosine distance and Euclidean distance at different scales to obtain an anomaly score;
[0045] The tablet defect detection module is used to detect tablet defects based on the anomaly score to obtain the tablet defect detection result.
[0046] In a third aspect, the present invention provides an electronic device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the above-described unsupervised learning tablet defect detection method based on a diffusion model are implemented.
[0047] Fourthly, the present invention provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the unsupervised learning tablet defect detection method based on the diffusion model introduced above are implemented.
[0048] Compared with the prior art, the present invention has the following beneficial effects:
[0049] The present invention belongs to an improved invention. Compared with the existing tablet defect detection methods, on the one hand, the present invention first uses an auxiliary reconstruction network to extract the semantic information of the original tablet image, and then uses the extracted semantic information of the original tablet image to assist the diffusion model denoising network to reconstruct the obtained latent space vector. The auxiliary reconstruction network has strong feature extraction ability and feature representation ability. By using the semantic information of the original tablet image extracted by the auxiliary reconstruction network to assist the diffusion model denoising network to reconstruct the obtained latent space vector, the quality of the reconstructed tablet image can be improved, making the reconstructed result closer to the normal tablet, so as to ensure the effectiveness of the comparison between the original tablet image and the reconstructed tablet image. On the other hand, the present invention weights the cosine distance and Euclidean distance of different scales to obtain an anomaly score. The weighting process can comprehensively consider the two factors of the Euclidean distance and cosine distance of different scales, making the tablet defect detection result closer to the actual situation, thus effectively solving the problems of poor reconstruction effect and low detection accuracy in the prior art.
[0050] Furthermore, the present invention discloses the specific process of obtaining the reconstructed tablet image. The original tablet image is encoded to obtain a latent space vector, and the latent space vector is noise-added to obtain a noise-added latent space vector; a diffusion model denoising network and an auxiliary reconstruction network are constructed; the noise-added latent space vector is input into the diffusion model denoising network for denoising to obtain a denoised latent space vector; the original tablet image and the noise-added latent space vector are jointly input into the auxiliary reconstruction network. First, the auxiliary reconstruction network is used to extract the semantic information of the original tablet image, and then the extracted semantic information of the original tablet image is used to assist the diffusion model denoising network to reconstruct the denoised latent space vector to obtain a reconstructed latent space vector; the reconstructed latent space vector is decoded to obtain a reconstructed tablet image. It can be seen that the reconstructed tablet image obtained by the present invention contains both latent space features and original semantic features, and contains various factors, thus improving the quality of the reconstructed tablet image.
[0051] Furthermore, the present invention discloses the specific process of extracting features from the original tablet image and the reconstructed tablet image respectively to obtain the features of the original tablet image and the features of the reconstructed tablet image. The residual network is used to extract features from the original tablet image and the reconstructed tablet image respectively, and a number of initially original tablet image features and initially reconstructed tablet image features of different scale sizes are obtained; the feature pyramid network is used to fuse the initially original tablet image features and initially reconstructed tablet image features of different scale sizes respectively, and a number of finally original tablet image features and finally reconstructed tablet image features of different scale sizes are obtained. It can be seen that the present invention can make full use of the feature outputs of each stage of the residual network, so as to more perfectly characterize the difference between the input image (original tablet image) and the reconstructed image (reconstructed tablet image). BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 is a flowchart of the unsupervised learning tablet defect detection method based on the diffusion model of the present invention;
[0053] Figure 2 is a schematic diagram of the unsupervised learning tablet defect detection system based on the diffusion model of the present invention;
[0054] Figure 3 is a flowchart of the unsupervised learning tablet defect detection method based on the diffusion model in Embodiment 3 of the present invention;
[0055] Figure 4 is a schematic diagram of the reconstruction principle of the original tablet image of the present invention;
[0056] Figure 5 is a structural diagram of the auxiliary reconstruction network of the present invention;
[0057] Figure 6 is a structural diagram of the feature fusion and extraction module of the present invention;
[0058] Figure 7 is a structural diagram of the convolutional attention module of the present invention;
[0059] Figure 8 is a comparison result diagram of the present invention and the existing algorithms;
[0060] Figure 9 is a reconstruction effect diagram and a result diagram of tablet defect detection of the present invention;
[0061] Figure 10 is a schematic structural diagram of the electronic device of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0062] To further understand the content of the present invention, the following provides a detailed description of the present invention in conjunction with the accompanying drawings and specific embodiments. It should be understood that the embodiments are only for explaining the present invention rather than limiting it.
[0063] The unsupervised learning pill defect detection method based on the diffusion model proposed by the present invention obtains the original pill image; encodes the original pill image to obtain a latent space vector; constructs a diffusion model denoising network and an auxiliary reconstruction network. First, the auxiliary reconstruction network is used to extract the semantic information of the original pill image, and then the extracted semantic information of the original pill image is used to assist the diffusion model denoising network to reconstruct the obtained latent space vector to obtain a reconstructed pill image; feature extraction is performed on the original pill image and the reconstructed pill image respectively to obtain the original pill image features and the reconstructed pill image features; the Euclidean distance and cosine distance at different scales are calculated using the original pill image features and the reconstructed pill image features; the cosine distances and Euclidean distances at different scales are weighted to obtain an anomaly score; the pill defect is detected based on the anomaly score to obtain the pill defect detection result. Compared with the prior art, the present invention effectively solves the problems of poor reconstruction effect and low detection accuracy in the prior art.
[0064] Embodiment 1:
[0065] The flow chart of the unsupervised learning pill defect detection method based on the diffusion model of the present invention is as Figure 1 shown. The unsupervised learning pill defect detection method based on the diffusion model of the present invention includes the following steps:
[0066] S1. Obtain the original pill image;
[0067] S2. Encode the original pill image to obtain a latent space vector;
[0068] S3. Construct a diffusion model denoising network and an auxiliary reconstruction network. First, use the auxiliary reconstruction network to extract the semantic information of the original pill image, and then use the extracted semantic information of the original pill image to assist the diffusion model denoising network to reconstruct the obtained latent space vector to obtain a reconstructed pill image;
[0069] S4. Perform feature extraction on the original pill image and the reconstructed pill image respectively to obtain the original pill image features and the reconstructed pill image features;
[0070] S5. Calculate the cosine distance and Euclidean distance at different scales using the original pill image features and the reconstructed pill image features, where the scale refers to the size of the original pill image features and the reconstructed pill image features;
[0071] S6. Weight the cosine distances and Euclidean distances at different scales to obtain an anomaly score;
[0072] S7. Detect tablet defects based on the anomaly score to obtain the tablet defect detection result.
[0073] Example 2:
[0074] The schematic diagram of the unsupervised learning tablet defect detection system based on the diffusion model of the present invention is as Figure 2 shown. The unsupervised learning tablet defect detection system based on the diffusion model of the present invention includes a data acquisition module, a latent space vector acquisition module, an image reconstruction module, a feature extraction module, a distance determination module, an anomaly score determination module, and a tablet defect detection module.
[0075] The data acquisition module is used to acquire the original tablet image.
[0076] The latent space vector acquisition module is used to encode the original tablet image to obtain the latent space vector.
[0077] The image reconstruction module is used to construct a diffusion model denoising network and an auxiliary reconstruction network. First, use the auxiliary reconstruction network to extract the semantic information of the original tablet image, and then use the extracted semantic information of the original tablet image to assist the diffusion model denoising network to reconstruct the obtained latent space vector to obtain the reconstructed tablet image.
[0078] The feature extraction module is used to extract features from the original tablet image and the reconstructed tablet image respectively to obtain the original tablet image features and the reconstructed tablet image features.
[0079] The distance determination module is used to calculate the Euclidean distance and cosine distance at different scales using the original tablet image features and the reconstructed tablet image features, where the scale refers to the size of the original tablet image features and the reconstructed tablet image features.
[0080] The anomaly score determination module is used to weight the cosine distance and Euclidean distance at different scales to obtain the anomaly score.
[0081] The tablet defect detection module is used to detect tablet defects based on the anomaly score to obtain the tablet defect detection result.
[0082] Example 3:
[0083] The flowchart of the unsupervised learning tablet defect detection method based on the diffusion model of the present invention is as Figure 3 shown. The unsupervised learning tablet defect detection method based on the diffusion model of the present invention includes the following steps:
[0084] S1. Acquire the original tablet image.
[0085] S2. Encode the original pill image to obtain a latent space vector.
[0086] Encode the original pill image using an encoder to obtain a latent space vector .
[0087] S3. Construct a diffusion model denoising network and an auxiliary reconstruction network. First, use the auxiliary reconstruction network to extract the semantic information of the original pill image, and then use the extracted semantic information of the original pill image to assist the diffusion model denoising network in reconstructing the obtained latent space vector to obtain a reconstructed pill image.
[0088] Construct a diffusion model denoising network and an auxiliary reconstruction network (the auxiliary reconstruction network refers to a neural network that assists the diffusion model in reconstruction through the semantic information of the original pill image in the reconstruction task to improve the reconstruction effect and quality. The auxiliary reconstruction network is also called the auxiliary reconstruction module). First, use the auxiliary reconstruction network to extract the semantic information of the original pill image, and then use the extracted semantic information of the original pill image to assist the diffusion model denoising network in reconstructing the obtained latent space vector. The schematic diagram of the reconstruction of the original pill image is as Figure 4 shown, and the specific process of obtaining the reconstructed pill image is as follows:
[0089] A. Perform noise addition processing (also called adding noise) on the obtained latent space vector to obtain a noise-added latent space vector (also called a latent space vector in the form of random Gaussian noise );
[0090] The latent space vector in the form of random Gaussian noise has the following calculation formula:
[0091]
[0092] where is the diffusion time step, is the signal retention ratio at time step t during the diffusion process, is the cumulative noise retention ratio from the initial time to time step t , , is the preset variance schedule, determines the amount of noise added, is an identity matrix with the same dimension as the initial value of the input image (the original pill image), is the noise sampled from the standard Gaussian distribution, s s is a loop variable, sThe value range of t is (1, is the s value at time . is the at the 0th time, is a normal distribution.
[0093] B. Construct a diffusion model denoising network and an auxiliary reconstruction network;
[0094] C. Input the noisy latent space vector into the diffusion model denoising network for denoising to obtain the denoised latent space vector;
[0095] D. Input the original tablet image and the noisy latent space vector into the auxiliary reconstruction network together. Use the auxiliary reconstruction network to extract the semantic information of the original tablet image, and then use the extracted semantic information of the original tablet image to assist the diffusion model denoising network to reconstruct the denoised latent space vector to obtain the reconstructed latent space vector;
[0096] E. Decode the reconstructed latent space vector using a decoder to obtain the reconstructed tablet image.
[0097] The following is a detailed description of the auxiliary reconstruction network:
[0098] The structural diagram of the auxiliary reconstruction network is as shown in Figure 5 . The auxiliary reconstruction network first performs feature extraction on the input image (the original tablet image) by a convolutional neural network , and reduces the dimension of the extracted original features through one layer of convolution. The output after dimension reduction is combined with the latent space vector in the form of random Gaussian noise , and encoding operations of four encoding blocks in sequence and context information extraction and feature refinement operations of the ARM (Assisted Reconstruction Mid-block, the middle block of the auxiliary reconstruction network, Figure 5 the middle block of the auxiliary reconstruction network in is represented by the auxiliary reconstruction middle block) are performed to obtain the semantic information of the original tablet image. The processing result is sent to the MID (Mid-block, the middle block) of the diffusion model denoising network Figure 5 the MID of the diffusion model denoising network in is represented by the denoising middle block. At the same time, the output of each encoding block will be added to the decoding block of the denoising network at the corresponding scale to assist the diffusion model denoising network to reconstruct the denoised latent space vector. The reconstructed output (the reconstructed tablet image) contains both latent space features and original semantic features, and can also pay attention to multi-scale information, and can obtain better reconstruction results.
[0099] S4. Feature extraction is performed on the original tablet image and the reconstructed tablet image respectively to obtain the original tablet image features and the reconstructed tablet image features.
[0100] The specific process of performing feature extraction on the original tablet image and the reconstructed tablet image respectively to obtain the original tablet image features and the reconstructed tablet image features is as follows:
[0101] The Residual Network (ResNet) is used to perform feature extraction on the original tablet image and the reconstructed tablet image respectively to obtain a number of original tablet image features and reconstructed tablet image features of different scale sizes at the beginning;
[0102] The Feature Pyramid Network is used to fuse the original tablet image features and the reconstructed tablet image features of different scale sizes at the beginning respectively to obtain a number of final original tablet image features and final reconstructed tablet image features of different scale sizes.
[0103] In this step, a convolutional attention module is also added to each residual block of the residual network.
[0104] In this embodiment, the original tablet image features and the reconstructed tablet image features are specifically obtained through the FFEM (Feature Fusion Extraction module), and the Feature Fusion Extraction module is specifically described as follows:
[0105] The structure diagram of the Feature Fusion Extraction module is as Figure 6 shown, Figure 6 where C represents the feature maps (feature maps are also called features) of each stage of the residual network. The size of the feature maps in each stage gradually decreases, and the number of channels gradually increases. The shallow stage can extract low-level features such as the edges and textures of the image, and the subsequent stages can extract high-level semantic features. In this embodiment, the input image size is fixed at 256×256. After the initial convolution and max pooling processes in the C1 stage, preliminary feature extraction and two consecutive downsampling operations are performed, making the size of the feature maps become 64×64, enhancing the robustness of the features. After completing the initial convolution and pooling operations, the feature maps enter 4 main residual block stages (C2 - C5). Each stage is stacked by multiple residual blocks, and the number of residual blocks and the size of the feature maps are different in different stages. In the C2 - C5 stages, the residual network continuously learns the high-level features of the image through the residual blocks, and further downsamples through the convolutional operation with a stride of 2, gradually reducing the size of the feature maps and increasing the number of channels of the feature maps. Figure 6In this, P represents the feature fusion part of the top-down network of the FPN (Feature Pyramid Network). In this part, feature maps of different scales are fused, and then through lateral connection operations, F feature maps of corresponding sizes ( Figure 6 represented by F2, F3, F4, and F5 in this case) are formed, thus obtaining a rich multi-scale feature representation. The combination of the residual network and the FPN structure can construct a feature pyramid at different scales, fully integrating and utilizing the feature outputs of each stage in the residual network, so as to more perfectly represent the difference between the original image and the reconstructed image. M represents detecting tablet defects using the features of the original tablet image and the reconstructed tablet image features.
[0106] Among them, the shallow stage refers to the starting part of the residual network, which mainly preprocesses the input image, extracts low-level features such as the edges and textures of the image, and at the same time reduces the data volume through the pooling layer to prepare for feature extraction in subsequent stages.
[0107] The subsequent stage refers to the part of the residual network after the shallow stage, which contains multiple residual blocks and more complex network structures. Through the shortcut connection of the residual blocks and continuous feature fusion, high-level semantic features of the image are gradually extracted, thereby realizing the optimization and refinement of the features.
[0108] The specific description of the convolutional attention module is as follows:
[0109] CBAM (Convolutional Block Attention Module) is a lightweight attention module, which consists of a channel attention module and a spatial attention module (Spatial Attention Module, SAM). CBAM can adaptively adjust the feature responses in the channel and spatial dimensions of the feature map. The structure diagram of the convolutional attention module is as Figure 7 shown.
[0110] Among them, the channel attention module mainly focuses on the importance of different channels in the feature map. For the input feature map, it performs global average pooling and global max pooling operations respectively to obtain two different global feature descriptors. These two global feature descriptors are processed by a shared multi-layer perceptron (MLP). The MLP contains a dimensionality reduction layer and a dimensionality increase layer. The purpose of the dimensionality reduction layer is to reduce the computational complexity. Finally, the two feature descriptors processed by the MLP are added together, and a channel attention map is obtained through the Sigmoid function. The channel attention map is multiplied element-wise with the input feature map to achieve attention adjustment for the channels.
[0111] The spatial attention module focuses on the importance of different spatial positions in the feature map. For the feature map processed by the channel attention module, average pooling and max pooling operations are performed respectively in the channel dimension to obtain two one-dimensional spatial feature descriptors. These two spatial feature descriptors are concatenated in the channel dimension and then passed through a 7×7 convolutional layer for feature fusion and dimensionality reduction to obtain a spatial attention map. Finally, the spatial attention map is multiplied element-wise with the input feature map to complete the attention adjustment for the spatial positions. Adding it after each layer of the residual network can extract rich feature information while maintaining the deep feature representation of small features, process feature information of different scales, enable the feature fusion extraction model to simultaneously focus on local and global features and solve the problem of missing target details, thereby making the feature extraction network have a stronger extraction ability. Figure 7 Among them, H, W, and C respectively represent the height, width, and number of channels of the feature map (where the full name of H is Height, the full name of W is Width, and the full name of C is Channel), Sigmoid is the Sigmoid activation function, X is the feature map input to CBAM, Y is the feature map processed by the channel attention module, and Z is the output feature map after being processed by the channel attention module and the spatial attention module in sequence.
[0112] S5. Calculate the cosine distance and Euclidean distance at different scales using the original tablet image features and the reconstructed tablet image features.
[0113] Calculate the cosine distance and Euclidean distance at different scales using the original tablet image features and the reconstructed tablet image features. Here, the scale refers to the size of the original tablet image features and the reconstructed tablet image features.
[0114] The calculation formula for the cosine distance at different scales is:
[0115]
[0116] Among them, is the cosine distance of the i th scale, is the cosine similarity of the i th scale, is the original pill image feature, is the reconstructed pill image feature, i is the number of scales of the original pill image feature and the reconstructed pill image feature.
[0117] The calculation formula for the Euclidean distance of different scales is:
[0118]
[0119] Among them, is the Euclidean distance of the i th scale, is the original pill image feature of the j th element value, is the reconstructed pill image feature of the j th element value, i is the number of scales of the original pill image feature and the reconstructed pill image feature, n is the original pill image feature containing the number of elements, j is the index variable used to traverse the original pill image feature and the reconstructed pill image feature in the element values, j ranges from [1, n .
[0120] S6. Weight the cosine distance and Euclidean distance of different scales to obtain the anomaly score.
[0121] The calculation formula for the anomaly score is:
[0122]
[0123] Among them, represents the anomaly score, is the upsampling factor of the i th scale, is the Euclidean distance of the i th scale, is the cosine distance of the i th scale, i is the number of scales of the original pill image feature and the reconstructed pill image feature, n is the original pill image feature The number of elements included, is the weight coefficient, whose value ranges from 0 to 1. In this embodiment, the weight coefficient is 0.5. The weight coefficient can be adjusted according to actual needs.
[0124] S7. Detect the tablet defects based on the anomaly score to obtain the tablet defect detection result.
[0125] The specific process of tablet defect detection in this step is as follows:
[0126] Compare the anomaly score with the set anomaly threshold. When the anomaly score is greater than the set anomaly threshold, it indicates that the tablet has a defect. When the anomaly score is less than or equal to the set anomaly threshold, it indicates that the tablet has no defect. The anomaly threshold in this embodiment is set according to actual needs.
[0127] To verify the effectiveness of the unsupervised learning tablet defect detection method based on the diffusion model proposed in the present invention, this embodiment uses the MVTEC dataset (MVTec Anomaly Detection Dataset, anomaly detection dataset) combined with the tablet defect dataset augmented by actual pictures collected on the industrial front line (specifically including 1000 normal tablets of various shapes and 2000 abnormal photos of various types of defects) for verification.
[0128] The operating system used in this embodiment is ubuntu20.04, the CPU (Central Processing Unit) model is 14 vCPU Intel(R) Xeon(R) Platinum 8362 CPU @ 2.80GHz, the GPU (Graphics Processing Unit) model is RTX3090, the video memory size is 48GB, and the memory size is 64GB.
[0129] All models used in this embodiment are implemented based on Pytorch, and the cuda version is 11.3.
[0130] This embodiment uses F1 score, AP (Average Precision), PRO (per-region–overlap), and AUROC (Area Under the Receiver Operating Characteristic Curve) to evaluate the tablet defect detection result. The following is a detailed description of F1 score, AP, PRO, and AUROC:
[0131] The F1-score combines the harmonic mean of precision (the proportion of correctly predicted positive examples among those predicted as positive) and recall (the proportion of correctly predicted positive examples among the actual positive examples). The F1-score is high only when both precision and recall are high. The F1-score is commonly used for performance evaluation in binary and multi-class classification tasks to help assess the overall performance of the model in terms of precision and recall.
[0132] AP calculates the average precision at different recall levels of the model, reflecting the accuracy of the model in ranking relevant samples. The higher the AP value, the better the performance of the model in retrieving or detecting relevant objects.
[0133] PRO is the degree of overlap between different regions in the image. It is used to evaluate the accuracy of the model's detection results. By calculating the overlap ratio, it judges the matching degree between the predicted region and the actual target region, and further measures the performance of the model in locating the target.
[0134] The ROC curve is plotted with the false positive rate on the x-axis and the true positive rate on the y-axis. AUROC is the area under the ROC curve, with a value range between 0 and 1. The larger the AUROC value, the stronger the ability of the model to distinguish positive and negative examples, that is, the better the classification performance of the model.
[0135] In this embodiment, the method proposed in the present invention is compared with four algorithms, namely LDM (Latent Diffusion Model), PaDiM (Patch Distribution Modeling), SimpleNet (a simple and easy-to-apply neural network for detecting and locating anomalies), and DRAEM (Discriminatively Trained Reconstruction Anomaly Embedding Model) on the tablet defect dataset expanded by combining the MVTEC dataset and actual pictures collected from the industrial front line. The comparison results are shown in Table 1 and Figure 8 as shown. In Table 1, ↑ indicates that the higher the evaluation index, the better the performance of the algorithm model.
[0136] Table 1 Comparison Results
[0137]
[0138] The data in Table 1 show that the method proposed in the present invention has the best detection effect, and all four evaluation indicators are the highest. Therefore, the method proposed in the present invention performs better than the four algorithms of LDM, PaDiM, SimpleNet, and DRAEM, and can complete the tablet defect detection task more excellently.
[0139] In order to verify the effectiveness of the feature fusion extraction module, different feature extraction networks (Resnet50, Resnet101, VGG16, VGG19, Efficient net and Inception net) are compared with the feature fusion extraction module. In this embodiment, the feature fusion extraction module of the method proposed in the present invention is replaced by Resnet50, Resnet101, VGG16, VGG19, Efficient net and Inception net respectively. The ablation experiment results are shown in Table 2. ↑ in Table 2 indicates that the higher the evaluation index, the better the algorithm model performance.
[0140] Table 2 Ablation experiment results
[0141]
[0142] The data in Table 2 show that the feature fusion extraction module in the method proposed in the present invention is superior to the other six feature extraction network combination algorithms in all indicators, which shows that the method proposed in the present invention can achieve a more ideal effect in tablet defect detection. The average accuracy of the method proposed in the present invention is 2.5% higher than the best data in the other six algorithms, the F1 score is 2.1% higher than the best data in other algorithms, the area under the receiver operating characteristic curve is 0.9% higher than the best data in other algorithms, and the area overlap is 1.1% higher than the best data in other algorithms, which further proves the advantages of the method proposed in the present invention in the task of tablet defect detection.
[0143] The reconstruction effect diagram and tablet defect detection result diagram of the tablet defect dataset expanded by combining the MVTEC dataset with the actual images collected from the industrial front line in this embodiment are shown in the figure below. Figure 9 As shown, Figure 9 The left side shows the reconstruction effect of 4 types of contaminated tablets and the detection effect of abnormal areas (also called defect areas), and the right side shows the reconstruction effect of 4 types of defective tablets and the detection effect of abnormal areas. Figure 9 It can be seen that for tablets of different types and with different defect types, the reconstructed images of the method proposed in the present invention are similar to normal tablets, and can achieve good reconstruction effects. In terms of tablet defect detection effects, whether it is a pure color tablet with simple texture or a multicolored tablet with more complex texture, the method proposed in the present invention can perform good distinction and detection, and can also perform abnormal heat map ( Figure 9 The anomaly heat map is used for visual display.
[0144] Embodiment 4:
[0145] See also Figure 10As shown in the figure, the present invention also provides an electronic device 100 for an unsupervised learning tablet defect detection method based on a diffusion model; the electronic device 100 includes a memory 101, at least one processor 102, a computer program 103 stored in the memory 101 and executable on the at least one processor 102, and at least one communication bus 104.
[0146] The memory 101 can be used to store the computer program 103. The processor 102 realizes the steps of the unsupervised learning tablet defect detection method based on the diffusion model described in Embodiment 1 by running or executing the computer program stored in the memory 101 and calling the data stored in the memory 101. The memory 101 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the electronic device 100 (such as audio data, etc.). In addition, the memory 101 can include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices.
[0147] The at least one processor 102 can be a Central Processing Unit (CPU), or can also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 102 can be a microprocessor, or the processor 102 can also be any conventional processor, etc. The processor 102 is the control center of the electronic device 100, and connects various parts of the entire electronic device 100 through various interfaces and lines.
[0148] The memory 101 in the electronic device 100 stores multiple instructions to implement the unsupervised learning tablet defect detection method based on the diffusion model. The processor 102 can execute the multiple instructions to implement:
[0149] Obtain the original tablet image;
[0150] Encode the original pill image to obtain a latent space vector;
[0151] Construct a diffusion model denoising network and an auxiliary reconstruction network. First, use the auxiliary reconstruction network to extract the semantic information of the original pill image, and then use the extracted semantic information of the original pill image to assist the diffusion model denoising network to reconstruct the obtained latent space vector to obtain a reconstructed pill image;
[0152] Extract features from the original pill image and the reconstructed pill image respectively to obtain the original pill image features and the reconstructed pill image features;
[0153] Use the original pill image features and the reconstructed pill image features to calculate the Euclidean distance and cosine distance at different scales, where the scale refers to the size of the original pill image features and the reconstructed pill image features;
[0154] Weight the cosine distance and Euclidean distance at different scales to obtain an anomaly score;
[0155] Detect pill defects based on the anomaly score to obtain the pill defect detection result.
[0156] Example 5:
[0157] If the modules / units integrated in the electronic device 100 are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above method embodiments of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, and read-only memory (ROM, Read-Only Memory).
[0158] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0159] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0160] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that realize the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0161] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0162] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: the specific implementation manners of the present invention can still be modified or equivalently replaced, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the protection scope of the claims of the present invention.
Claims
1. An unsupervised learning method for detecting tablet defects based on a diffusion model, characterized in that, It includes the following steps: Obtain the original pill image; Perform encoding processing on the original pill image to obtain a latent space vector; Construct a diffusion model denoising network and an auxiliary reconstruction network. First, use the auxiliary reconstruction network to extract the semantic information of the original pill image, and then use the extracted semantic information of the original pill image to assist the diffusion model denoising network to reconstruct the obtained latent space vector to obtain a reconstructed pill image; Extract features from the original pill image and the reconstructed pill image respectively to obtain the original pill image features and the reconstructed pill image features; The step of extracting features from the original pill image and the reconstructed pill image respectively to obtain the original pill image features and the reconstructed pill image features specifically includes: Use a residual network to extract features from the original pill image and the reconstructed pill image respectively to obtain a number of initial original pill image features and initial reconstructed pill image features with different scale sizes; Use a feature pyramid network to fuse the number of initial original pill image features and initial reconstructed pill image features with different scale sizes respectively to obtain the original pill image features and the reconstructed pill image features; In the step of using a residual network to extract features from the original pill image and the reconstructed pill image respectively to obtain a number of initial original pill image features and initial reconstructed pill image features with different scale sizes, a convolutional attention module is also added to each residual block of the residual network; Use the original pill image features and the reconstructed pill image features to calculate the cosine distance and Euclidean distance at different scales, where the scale refers to the size of the original pill image features and the reconstructed pill image features; Weight the cosine distance and Euclidean distance at different scales to obtain an anomaly score; Detect the pill defect based on the anomaly score to obtain the pill defect detection result.
2. The unsupervised learning tablet defect detection method based on a diffusion model according to claim 1, characterized in that, The step of constructing a diffusion model denoising network and an auxiliary reconstruction network, first using the auxiliary reconstruction network to extract the semantic information of the original pill image, and then using the extracted semantic information of the original pill image to assist the diffusion model denoising network to reconstruct the obtained latent space vector to obtain a reconstructed pill image specifically includes: Perform noise addition processing on the obtained latent space vector to obtain a noise-added latent space vector; Construct a diffusion model denoising network and an auxiliary reconstruction network; Input the noise-added latent space vector into the diffusion model denoising network for denoising processing to obtain a denoised latent space vector; Input the original pill image and the noise-added latent space vector into the auxiliary reconstruction network together. First, use the auxiliary reconstruction network to extract the semantic information of the original pill image, and then use the extracted semantic information of the original pill image to assist the diffusion model denoising network to reconstruct the denoised latent space vector to obtain a reconstructed latent space vector; Perform decoding processing on the reconstructed latent space vector to obtain a reconstructed pill image.
3. The unsupervised learning tablet defect detection method based on a diffusion model according to claim 1, characterized in that, In the step of using the original pill image features and the reconstructed pill image features to calculate the cosine distance at different scales, the calculation formula for the cosine distance at different scales is: Among them, is the cosine distance of the i th scale, is the cosine similarity of the i th scale, is the feature of the original tablet image, is the feature of the reconstructed tablet image, i is the number of scales of the feature of the original tablet image and the feature of the reconstructed tablet image; The calculation formula for the Euclidean distance at different scales is: Among them, is the Euclidean distance of the i -th scale, is the -th element value of the original tablet image feature j , is the -th element value of the reconstructed tablet image feature j , i is the number of scales of the original tablet image feature and the reconstructed tablet image feature, n is the original tablet image feature including the number of elements, j is an index variable used to traverse the element values in the original tablet image feature and the reconstructed tablet image feature , j ranges from 1 to n .
4. The unsupervised learning tablet defect detection method based on a diffusion model according to claim 3, wherein In the step of weighting the cosine distance and Euclidean distance at different scales to obtain the anomaly score, the calculation formula of the anomaly score is as follows: Among them, represents the anomaly score, is the upsampling factor of the i -th scale, is the Euclidean distance of the i -th scale, is the cosine distance of the i -th scale, i is the number of scales of the original pill image features and the reconstructed pill image features, n is the original pill image features contains the number of elements, is the weight coefficient, takes values between 0 and 1.
5. The unsupervised learning tablet defect detection method based on a diffusion model according to claim 1, wherein The step of detecting tablet defects based on the anomaly score to obtain the tablet defect detection result is specifically as follows: Compare the anomaly score with the set anomaly threshold. When the anomaly score is greater than the set anomaly threshold, it indicates that there are defects in the tablet; when the anomaly score is less than or equal to the set anomaly threshold, it indicates that there are no defects in the tablet.
6. An unsupervised learning tablet defect detection system based on a diffusion model, characterized in that, It includes a data acquisition module, a latent space vector acquisition module, an image reconstruction module, a feature extraction module, a distance determination module, an anomaly score determination module, and a tablet defect detection module; The data acquisition module is used to acquire the original tablet image; The latent space vector acquisition module is used to perform encoding processing on the original tablet image to obtain a latent space vector; The image reconstruction module is used to construct a diffusion model denoising network and an auxiliary reconstruction network. First, use the auxiliary reconstruction network to extract the semantic information of the original tablet image, and then use the extracted semantic information of the original tablet image to assist the diffusion model denoising network to reconstruct the obtained latent space vector to obtain a reconstructed tablet image; The feature extraction module is used to perform feature extraction on the original tablet image and the reconstructed tablet image respectively to obtain the original tablet image features and the reconstructed tablet image features; The step of performing feature extraction on the original tablet image and the reconstructed tablet image respectively to obtain the original tablet image features and the reconstructed tablet image features specifically includes: Use a residual network to perform feature extraction on the original tablet image and the reconstructed tablet image respectively to obtain several initially original tablet image features and initially reconstructed tablet image features of different scale sizes; Use a feature pyramid network to fuse several initially original tablet image features and initially reconstructed tablet image features of different scale sizes respectively to obtain the original tablet image features and the reconstructed tablet image features; In the step of using a residual network to perform feature extraction on the original tablet image and the reconstructed tablet image respectively to obtain several initially original tablet image features and initially reconstructed tablet image features of different scale sizes, a convolutional attention module is also added to each residual block of the residual network; The distance determination module is used to calculate the cosine distance and Euclidean distance at different scales by using the original tablet image features and the reconstructed tablet image features, where the scale refers to the size of the original tablet image features and the reconstructed tablet image features; The anomaly score determination module is used to weight the cosine distance and Euclidean distance at different scales to obtain the anomaly score; The tablet defect detection module is used to detect tablet defects based on the anomaly score to obtain the tablet defect detection result.
7. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the unsupervised learning tablet defect detection method based on the diffusion model according to any one of claims 1 to 5.
8. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the unsupervised learning tablet defect detection method based on the diffusion model according to any one of claims 1 to 5.
Citation Information
Patent Citations
Defect detection method, device, system and equipment based on image processing, medium and product
CN118297884A
Multi-class defect detection method and device based on variational auto-encoder and denoising network
CN119516249A