Unsupervised learning tablet defect detection method based on diffusion model and related device
By adopting an unsupervised learning method based on diffusion model in the detection of tablet defects, a diffusion model denoising network and auxiliary reconstruction network are constructed, the verbal information of the tablet image is extracted and the potential spatial vector is reconstructed, and the cosine distance and Euclidean distance are weighted to obtain abnormal scores for defect detection, which solves the problem of poor reconstruction effect and low detection accuracy in the existing technology, and achieves more efficient tablet defect detection.
Patent Information
- Application Number
- CN202510650415.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-20
AI Technical Summary
The existing pill defect detection methods have poor reconstruction effect and low detection accuracy.
Using an unsupervised learning method based on diffusion model, by constructing a diffusion model denoising network and auxiliary reconstruction network, semantic information of the pill image is extracted and potential spatial vectors are reconstructed, and the cosine distance and Euclidean distance are weighted to obtain abnormal scores for defect detection.
It improves the reconstruction quality and detection accuracy of tablet defect detection, can more effectively identify tablet defects, and reduces the misjudgment rate.
Smart Images

Figure CN120182256A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of tablet defect detection, and specifically relates to an unsupervised learning tablet defect detection method and related device based on a diffusion model. Background Art
[0002] The quality of drugs is directly related to the health and safety of patients and the brand reputation of pharmaceutical companies. In recent years, drug safety issues have occurred frequently, and the number of drug safety accident problems has continued to rise. Among them, the problem of drug surface defects is an important part. It is not only a key quality problem that needs to be detected urgently in the production process, but also the most intuitive manifestation of the affected product quality. Therefore, detecting drug surface defects is of great significance for ensuring drug quality, improving drug use safety, and reducing the frequency of accidents.
[0003] Tablets are the most common form of drugs. Tablet defect detection generally aims to detect defects such as scratches, defects, foreign object occlusion, color contamination, and holes on the tablet surface. Traditional manual visual inspection methods are inefficient and subjective, and it is difficult to meet the requirements of high-speed production lines for accuracy and real-time performance. Traditional image processing techniques achieve anomaly recognition through low-dimensional feature extraction and rule design, but they have insufficient generalization ability for complex texture defects and are easily affected by illumination changes and background interference.
[0004] In recent years, deep learning techniques have significantly improved the accuracy and robustness of tablet defect detection with their powerful feature learning ability and have gradually become the mainstream method. Among them, although the supervised method has high accuracy and good adaptability, this method requires a large amount of defect data to be labeled and cannot cope with the tablet production process. The defect samples generated during the tablet production process are few, and the defect forms are diverse.
[0005] Therefore, it is more appropriate to use unsupervised methods such as DRAEM (discriminatively trained reconstruction anomaly embedding model), PatchCore (Patch-based Core-set Sampling for Anomaly Detection), and LDM (Latent Diffusion Model). Among them, the reconstruction-based method is a major research hotspot. The core of the reconstruction-based method is that in the training stage, the model only learns the feature distribution from normal images, and in the testing stage, the trained model reconstructs abnormal images into normal images, so as to determine the abnormal position by comparing the reconstructed images with the input images.
[0006] Currently, the existing tablet defect detection methods have poor reconstruction effects and low detection accuracy. Summary of the Invention
[0007] The object of the present invention is to provide an unsupervised learning tablet defect detection method and related device based on a diffusion model, which is used to solve the problems of poor reconstruction effect and low detection accuracy in the prior art.
[0008] To achieve the above object, the present invention adopts the following technical solutions: In the first aspect, the present invention provides an unsupervised learning tablet defect detection method based on a diffusion model, including the following steps: Obtain the original tablet image; Perform encoding processing on the original tablet image to obtain a latent space vector; Construct a diffusion model denoising network and an auxiliary reconstruction network. First, use the auxiliary reconstruction network to extract the semantic information of the original tablet image, and then use the extracted semantic information of the original tablet image to assist the diffusion model denoising network to reconstruct the obtained latent space vector to obtain a reconstructed tablet image; Extract features from the original tablet image and the reconstructed tablet image respectively to obtain the original tablet image features and the reconstructed tablet image features; Use the original tablet image features and the reconstructed tablet image features to calculate the cosine distance and Euclidean distance at different scales, where the scale refers to the size of the original tablet image features and the reconstructed tablet image features; Weight the cosine distance and Euclidean distance at different scales to obtain an anomaly score; Detect tablet defects according to the anomaly score to obtain the tablet defect detection result.
[0009] A further improvement of the present invention is that in the step of constructing a diffusion model denoising network and an auxiliary reconstruction network, first using the auxiliary reconstruction network to extract the semantic information of the original tablet image, and then using the extracted semantic information of the original tablet image to assist the diffusion model denoising network to reconstruct the obtained latent space vector to obtain a reconstructed tablet image, specifically includes: Perform noise addition processing on the obtained latent space vector to obtain a noise-added latent space vector; Construct a diffusion model denoising network and an auxiliary reconstruction network; Input the noise-added latent space vector into the diffusion model denoising network for denoising processing to obtain a denoised latent space vector; Input the original tablet image and the noise-added latent space vector into the auxiliary reconstruction network together. First, use the auxiliary reconstruction network to extract the semantic information of the original tablet image, and then use the extracted semantic information of the original tablet image to assist the diffusion model denoising network to reconstruct the denoised latent space vector to obtain a reconstructed latent space vector; Decode the reconstructed latent space vector to obtain the reconstructed pill image.
[0010] A further improvement of the present invention lies in that the original pill image and the reconstructed pill image are respectively subjected to feature extraction to obtain the original pill image features and the reconstructed pill image features, which specifically include: Use a residual network to respectively perform feature extraction on the original pill image and the reconstructed pill image to obtain a number of initially original pill image features and initially reconstructed pill image features of different scale sizes; Use a feature pyramid network to respectively fuse a number of initially original pill image features and initially reconstructed pill image features of different scale sizes to obtain a number of finally original pill image features and finally reconstructed pill image features of different scale sizes.
[0011] A further improvement of the present invention lies in that in the step of using a residual network to respectively perform feature extraction on the original pill image and the reconstructed pill image to obtain a number of initially original pill image features and initially reconstructed pill image features of different scale sizes, a convolutional attention module is further added to each residual block of the residual network.
[0012] A further improvement of the present invention lies in that in the step of calculating the cosine distance and Euclidean distance of different scales by using the original pill image features and the reconstructed pill image features, where the scale refers to the size of the original pill image features and the reconstructed pill image features, the calculation formula for the cosine distance of different scales is:
[0013] Wherein, is the cosine distance of the i th scale, is the cosine similarity of the i th scale, is the original pill image feature, is the reconstructed pill image feature, i is the number of scales of the original pill image feature and the reconstructed pill image feature; The calculation formula for the Euclidean distance of different scales is:
[0014] Wherein, is the Euclidean distance of the i th scale, is the original pill image feature, is the reconstructed pill image feature, is the th of the original pill image featurej The value of the th element of the reconstructed tablet image feature is the j th element value, i where n is the number of scales of the original tablet image feature and the reconstructed tablet image feature, is the number of elements contained in the original tablet image feature j is an index variable used to traverse the element values in the original tablet image feature and the reconstructed tablet image feature , and the value range of j is [1, n .
[0015] A further improvement of the present invention lies in that in the step of weighting the cosine distance and the Euclidean distance at different scales to obtain the anomaly score, the calculation formula of the anomaly score is:
[0016] where represents the anomaly score, is the upsampling factor of the i th scale, is the Euclidean distance of the i th scale, is the cosine distance of the i th scale, i is the number of scales of the original tablet image feature and the reconstructed tablet image feature, n is the number of elements contained in the original tablet image feature , is the weight coefficient, and the value range of
[0017] is between 0 and 1. A further improvement of the present invention lies in that the detection of tablet defects according to the anomaly score to obtain the tablet defect detection result is specifically:
[0018] Compare the anomaly score with the set anomaly threshold. When the anomaly score is greater than the set anomaly threshold, it indicates that the tablet has defects; when the anomaly score is less than or equal to the set anomaly threshold, it indicates that the tablet has no defects. The data acquisition module is used to acquire the original tablet image; The potential space vector acquisition module is used to encode the original pill image to obtain a potential space vector; The image reconstruction module is used to construct a diffusion model denoising network and an auxiliary reconstruction network. First, the auxiliary reconstruction network is used to extract the semantic information of the original pill image, and then the semantic information of the original pill image is used to assist the diffusion model denoising network to reconstruct the obtained potential space vector to obtain a reconstructed pill image; The feature extraction module is used to extract features from the original pill image and the reconstructed pill image respectively to obtain the original pill image features and the reconstructed pill image features; The distance determination module is used to calculate the cosine distance and Euclidean distance at different scales by using the original pill image features and the reconstructed pill image features, where the scale refers to the size of the original pill image features and the reconstructed pill image features; The anomaly score determination module is used to weight the cosine distance and Euclidean distance at different scales to obtain an anomaly score; The pill defect detection module is used to detect pill defects according to the anomaly score to obtain a pill defect detection result.
[0019] In a third aspect, the present invention provides an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above-described unsupervised learning pill defect detection method based on a diffusion model are implemented.
[0020] In a fourth aspect, the present invention provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-described unsupervised learning pill defect detection method based on a diffusion model are implemented.
[0021] Compared with the prior art, the present invention has the following beneficial effects: The present invention belongs to an improved invention. Compared with the existing tablet defect detection methods, on the one hand, the present invention first uses an auxiliary reconstruction network to extract the semantic information of the original tablet image, and then uses the extracted semantic information of the original tablet image to assist the diffusion model denoising network to reconstruct the obtained latent space vector. The auxiliary reconstruction network has strong feature extraction ability and feature representation ability. By using the semantic information of the original tablet image extracted by the auxiliary reconstruction network to assist the diffusion model denoising network to reconstruct the obtained latent space vector, the quality of the reconstructed tablet image can be improved, making the reconstructed result closer to a normal tablet, so as to ensure the effectiveness of the comparison between the original tablet image and the reconstructed tablet image. On the other hand, the present invention weights the cosine distance and Euclidean distance at different scales to obtain an anomaly score. The weighting process can comprehensively consider the two factors of Euclidean distance and cosine distance at different scales, making the tablet defect detection result closer to the actual situation, thus effectively solving the problems of poor reconstruction effect and low detection accuracy in the prior art.
[0022] Further, the present invention discloses the specific process of obtaining the reconstructed tablet image. The original tablet image is encoded to obtain a latent space vector, and the latent space vector is subjected to noise addition processing to obtain a noise-added latent space vector; a diffusion model denoising network and an auxiliary reconstruction network are constructed; the noise-added latent space vector is input into the diffusion model denoising network for denoising processing to obtain a denoised latent space vector; the original tablet image and the noise-added latent space vector are jointly input into the auxiliary reconstruction network. First, the auxiliary reconstruction network is used to extract the semantic information of the original tablet image, and then the extracted semantic information of the original tablet image is used to assist the diffusion model denoising network to reconstruct the denoised latent space vector to obtain a reconstructed latent space vector; the reconstructed latent space vector is decoded to obtain the reconstructed tablet image. It can be seen that the reconstructed tablet image obtained by the present invention contains both latent space features and original semantic features, and contains diverse factors, thus improving the quality of the reconstructed tablet image.
[0023] Furthermore, the present invention discloses the specific process of extracting features from the original tablet image and the reconstructed tablet image respectively to obtain the features of the original tablet image and the reconstructed tablet image. The residual network is used to extract features from the original tablet image and the reconstructed tablet image respectively, and several initially original tablet image features and initially reconstructed tablet image features with different scale sizes are obtained; the feature pyramid network is used to fuse the several initially original tablet image features and initially reconstructed tablet image features with different scale sizes respectively, and several finally original tablet image features and finally reconstructed tablet image features with different scale sizes are obtained. It can be seen that the present invention can make full use of the feature outputs of each stage of the residual network, so as to be able to more perfectly represent the difference between the input image (original tablet image) and the reconstructed image (reconstructed tablet image). BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is a flowchart of the unsupervised learning tablet defect detection method based on the diffusion model of the present invention; Figure 2 is a schematic diagram of the unsupervised learning tablet defect detection system based on the diffusion model of the present invention; Figure 3 is a flowchart of the unsupervised learning tablet defect detection method based on the diffusion model in Embodiment 3 of the present invention; Figure 4 is a schematic diagram of the reconstruction of the original tablet image of the present invention; Figure 5 is a structural diagram of the auxiliary reconstruction network of the present invention; Figure 6 is a structural diagram of the feature fusion and extraction module of the present invention; Figure 7 is a structural diagram of the convolutional attention module of the present invention; Figure 8 is a comparison result diagram between the present invention and existing algorithms; Figure 9 is a reconstruction effect diagram and a result diagram of tablet defect detection of the present invention; Figure 10 is a schematic structural diagram of the electronic device of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] To further understand the content of the present invention, the following will describe the present invention in detail with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments are only for explaining the present invention and not for limiting it.
[0026] The unsupervised learning tablet defect detection method based on the diffusion model proposed by the present invention obtains the original tablet image; encodes the original tablet image to obtain a latent space vector; constructs a diffusion model denoising network and an auxiliary reconstruction network. First, the auxiliary reconstruction network is used to extract the semantic information of the original tablet image, and then the extracted semantic information of the original tablet image is used to assist the diffusion model denoising network to reconstruct the obtained latent space vector to obtain a reconstructed tablet image; feature extraction is performed on the original tablet image and the reconstructed tablet image respectively to obtain the original tablet image features and the reconstructed tablet image features; the Euclidean distance and cosine distance at different scales are calculated using the original tablet image features and the reconstructed tablet image features; the cosine distance and Euclidean distance at different scales are weighted to obtain an anomaly score; the tablet defect is detected according to the anomaly score to obtain the tablet defect detection result. Compared with the prior art, the present invention effectively solves the problems of poor reconstruction effect and low detection accuracy in the prior art.
[0027] Example 1: The flowchart of the unsupervised learning tablet defect detection method based on the diffusion model of the present invention is as Figure 1 shown. The unsupervised learning tablet defect detection method based on the diffusion model of the present invention includes the following steps: S1. Obtain the original tablet image; S2. Encode the original tablet image to obtain a latent space vector; S3. Construct a diffusion model denoising network and an auxiliary reconstruction network. First, use the auxiliary reconstruction network to extract the semantic information of the original tablet image, and then use the extracted semantic information of the original tablet image to assist the diffusion model denoising network to reconstruct the obtained latent space vector to obtain a reconstructed tablet image; S4. Perform feature extraction on the original tablet image and the reconstructed tablet image respectively to obtain the original tablet image features and the reconstructed tablet image features; S5. Use the original tablet image features and the reconstructed tablet image features to calculate the cosine distance and Euclidean distance at different scales, where the scale refers to the size of the original tablet image features and the reconstructed tablet image features; S6. Weight the cosine distance and Euclidean distance at different scales to obtain an anomaly score; S7. Detect the tablet defect according to the anomaly score to obtain the tablet defect detection result.
[0028] Example 2: The schematic diagram of the unsupervised learning tablet defect detection system based on the diffusion model of the present invention is as Figure 2As shown in the figure, the unsupervised learning tablet defect detection system based on the diffusion model of the present invention includes a data acquisition module, a latent space vector acquisition module, an image reconstruction module, a feature extraction module, a distance determination module, an anomaly score determination module, and a tablet defect detection module.
[0029] The data acquisition module is used to acquire the original tablet image.
[0030] The latent space vector acquisition module is used to encode the original tablet image to obtain a latent space vector.
[0031] The image reconstruction module is used to construct a diffusion model denoising network and an auxiliary reconstruction network. First, the semantic information of the original tablet image is extracted by the auxiliary reconstruction network, and then the semantic information of the original tablet image is used to assist the diffusion model denoising network to reconstruct the obtained latent space vector to obtain a reconstructed tablet image.
[0032] The feature extraction module is used to extract features from the original tablet image and the reconstructed tablet image respectively to obtain the original tablet image features and the reconstructed tablet image features.
[0033] The distance determination module is used to calculate the Euclidean distance and cosine distance at different scales by using the original tablet image features and the reconstructed tablet image features, where the scale refers to the size of the original tablet image features and the reconstructed tablet image features.
[0034] The anomaly score determination module is used to weight the cosine distance and Euclidean distance at different scales to obtain an anomaly score.
[0035] The tablet defect detection module is used to detect tablet defects according to the anomaly score to obtain the tablet defect detection result.
[0036] Embodiment 3: The flowchart of the unsupervised learning tablet defect detection method based on the diffusion model of the present invention is as Figure 3 As shown in the figure, the unsupervised learning tablet defect detection method based on the diffusion model of the present invention includes the following steps: S1. Acquire the original tablet image.
[0037] S2. Encode the original tablet image to obtain a latent space vector.
[0038] The original tablet image is encoded by an encoder to obtain a latent space vector .
[0039] S3. Construct a diffusion model denoising network and an auxiliary reconstruction network. First, use the auxiliary reconstruction network to extract the semantic information of the original pill image, and then use the extracted semantic information of the original pill image to assist the diffusion model denoising network in reconstructing the obtained latent space vector to obtain the reconstructed pill image.
[0040] Construct a diffusion model denoising network and an auxiliary reconstruction network (the auxiliary reconstruction network refers to a neural network that assists the diffusion model in reconstruction through the semantic information of the original pill image in the reconstruction task to improve the reconstruction effect and quality. The auxiliary reconstruction network is also called the auxiliary reconstruction module). First, use the auxiliary reconstruction network to extract the semantic information of the original pill image, and then use the extracted semantic information of the original pill image to assist the diffusion model denoising network in reconstructing the obtained latent space vector. The schematic diagram of the reconstruction of the original pill image is as Figure 4 shown. The specific process of obtaining the reconstructed pill image is as follows: A. Perform noise addition processing (also called adding noise) on the obtained latent space vector to obtain the latent space vector after noise addition (also called the latent space vector in the form of random Gaussian noise ); The formula for the latent space vector in the form of random Gaussian noise is:
[0041] where, is the diffusion time step, is the signal retention ratio at the time step t in the diffusion process, is the cumulative noise retention ratio from the initial time to the time step t , , is the preset variance schedule, determines the amount of noise added, is the same-dimensional identity matrix as the initial value of the input image (the original pill image), is the noise sampled from the standard Gaussian distribution, s s is the loop variable, t s
[0042] B. Construct a diffusion model denoising network and an auxiliary reconstruction network; C. Input the noisy latent space vector into the diffusion model denoising network for denoising to obtain a denoised latent space vector; D. Input the original pill image and the noisy latent space vector into the auxiliary reconstruction network together. Use the auxiliary reconstruction network to extract the semantic information of the original pill image, and then use the extracted semantic information of the original pill image to assist the diffusion model denoising network in reconstructing the denoised latent space vector to obtain a reconstructed latent space vector; E. Decode the reconstructed latent space vector using a decoder to obtain a reconstructed pill image.
[0043] The following is a detailed description of the auxiliary reconstruction network: The structural diagram of the auxiliary reconstruction network is as Figure 5 shown. The auxiliary reconstruction network first performs feature extraction on the input image (the original pill image) by a convolutional neural network and performs dimensionality reduction on the extracted original features through one layer of convolution. The output after dimensionality reduction is combined with the latent space vector in the form of random Gaussian noise and undergoes encoding operations in four consecutive encoding blocks and context information extraction and feature refinement operations of the ARM (Assisted Reconstruction Mid-block, the middle block of the auxiliary reconstruction network, Figure 5 where the middle block of the auxiliary reconstruction network is represented by the assisted reconstruction middle block) to obtain the semantic information of the original pill image. The processing result is sent to the MID (Mid-block, the middle block) of the diffusion model denoising network Figure 5 where the MID of the diffusion model denoising network is represented by the denoising middle block. At the same time, the output of each encoding block will be added to the decoding block of the denoising network at the corresponding scale to assist the diffusion model denoising network in reconstructing the denoised latent space vector. The reconstructed output (the reconstructed pill image) contains both latent space features and original semantic features, and can also focus on multi-scale information, enabling a better reconstruction effect.
[0044] S4. Perform feature extraction on the original pill image and the reconstructed pill image respectively to obtain the original pill image features and the reconstructed pill image features.
[0045] The specific process of performing feature extraction on the original pill image and the reconstructed pill image respectively to obtain the original pill image features and the reconstructed pill image features is as follows: The original pill image and the reconstructed pill image are respectively subjected to feature extraction using a Residual Network (ResNet) to obtain several initial original pill image features and initial reconstructed pill image features of different scale sizes; The Feature Pyramid Network is used to fuse the several initial original pill image features and initial reconstructed pill image features of different scale sizes respectively, to obtain several final original pill image features and final reconstructed pill image features of different scale sizes.
[0046] In this step, a convolutional attention module is also added to each residual block of the Residual Network.
[0047] In this embodiment, the original pill image features and the reconstructed pill image features are specifically obtained through the FFEM (Feature Fusion Extraction module), and the Feature Fusion Extraction module is specifically described as follows: The structural diagram of the Feature Fusion Extraction module is as Figure 6 shown, Figure 6 In it, C represents the feature maps (feature maps are also called features) of each stage of the Residual Network. The size of the feature maps in each stage gradually decreases, and the number of channels gradually increases. The shallow stage can extract low-level features such as the edges and textures of the image, and the subsequent stages can extract high-level semantic features. In this embodiment, the input image size is fixed at 256×256. After the initial convolution and max-pooling process of the C1 stage, preliminary feature extraction and two consecutive downsampling operations are performed, so that the size of the feature map becomes 64×64, enhancing the robustness of the features. After the initial convolution and pooling operations are completed, the feature map enters 4 main residual block stages (C2-C5). Each stage is stacked by multiple residual blocks, and the number of residual blocks and the size of the feature maps in different stages are different. In the C2-C5 stages, the Residual Network continuously learns the high-level features of the image through the residual blocks, and further downsamples through the convolutional operation with a stride of 2, gradually reducing the size of the feature map and increasing the number of channels of the feature map. Figure 6 In it, P represents the feature fusion part of the top-down network of the FPN (Feature Pyramid Network) structure. In this part, feature maps of different scales are fused, and then corresponding-sized F feature maps ( Figure 6 represented by F2, F3, F4, and F5 in it) are formed through lateral connection operations, so as to obtain a rich multi-scale feature representation. The combination of the Residual Network and the FPN structure can construct a feature pyramid at different scales, fully fuse and utilize the feature outputs of each stage in the Residual Network, so as to more perfectly represent the difference between the original image and the reconstructed image. M represents detecting pill defects using the original pill image features and the reconstructed pill image features.
[0048] The shallow stage refers to the starting part of the residual network, which mainly performs preliminary processing on the input image, extracts low-level features such as edges and textures of the image, and at the same time reduces the data volume through the pooling layer to prepare for feature extraction in subsequent stages.
[0049] The subsequent stage refers to the part of the residual network after the shallow stage, which contains multiple residual blocks and more complex network structures. Through the shortcut connection of the residual blocks and continuous feature fusion, high-level semantic features of the image are gradually extracted, thereby realizing the optimization and refinement of the features.
[0050] The convolutional attention module is specifically described as follows: CBAM (Convolutional Block Attention Module) is a lightweight attention module composed of a channel attention module and a spatial attention module (SAM). CBAM can adaptively adjust the feature responses in the channel and spatial dimensions of the feature map. The structure diagram of the convolutional attention module is as Figure 7 shown.
[0051] Among them, the channel attention module mainly focuses on the importance of different channels in the feature map. For the input feature map, it performs global average pooling and global max pooling operations respectively to obtain two different global feature descriptors. These two global feature descriptors are processed by a shared multi-layer perceptron (MLP), which contains a dimensionality reduction layer and a dimensionality increase layer. The purpose of the dimensionality reduction layer is to reduce the computational amount. Finally, the two feature descriptors processed by the MLP are added together, and a channel attention map is obtained through the Sigmoid function. It is multiplied element-wise with the input feature map to achieve attention adjustment for the channels.
[0052] The spatial attention module focuses on the importance of different spatial positions in the feature map. For the feature map processed by the channel attention module, average pooling and max pooling operations are respectively performed on the channel dimension to obtain two one-dimensional spatial feature descriptors. These two spatial feature descriptors are concatenated on the channel dimension and then passed through a 7×7 convolutional layer for feature fusion and dimensionality reduction to obtain the spatial attention map. Finally, the spatial attention map is multiplied element-wise with the input feature map to complete the attention adjustment of the spatial position. Adding it after each layer of the residual network can extract rich feature information while maintaining the deep feature representation of small features, process feature information of different scales, enable the feature fusion extraction model to simultaneously focus on local and global features and solve the problem of missing object details, thus making the feature extraction network have a stronger extraction ability. Figure 7 In this, H, W, and C respectively represent the height, width, and number of channels of the feature map (where the full name of H is Height, the full name of W is Width, and the full name of C is Channel), Sigmoid is the Sigmoid activation function, X is the feature map input to CBAM, Y is the feature map processed by the channel attention module, and Z is the output feature map after being processed by the channel attention module and the spatial attention module in sequence.
[0053] S5. Calculate the cosine distance and Euclidean distance at different scales using the original pill image features and the reconstructed pill image features.
[0054] Calculate the cosine distance and Euclidean distance at different scales using the original pill image features and the reconstructed pill image features. Here, the scale refers to the size of the original pill image features and the reconstructed pill image features.
[0055] The calculation formula for the cosine distance at different scales is:
[0056] Among them, is the cosine distance at the i th scale, is the cosine similarity at the i th scale, is the original pill image feature, is the reconstructed pill image feature, i is the number of scales of the original pill image features and the reconstructed pill image features.
[0057] The calculation formula for the Euclidean distance at different scales is:
[0058] Among them, is the iEuclidean distance at a scale, is the feature of the original tablet image The j element value of is the feature of the reconstructed tablet image The j element value of i is the number of scales of the features of the original tablet image and the reconstructed tablet image, n is the feature of the original tablet image The number of elements included j is an index variable used to traverse the elements in the features of the original tablet image and the features of the reconstructed tablet image The element values in j The value range of n is [1,
[0059] S6. Weight the cosine distances at different scales and the Euclidean distance to obtain an anomaly score.
[0060] The calculation formula for the anomaly score is:
[0061] where, represents the anomaly score, is the upsampling factor at the i th scale, is the Euclidean distance at the i th scale, is the cosine distance at the i th scale, i is the number of scales of the features of the original tablet image and the reconstructed tablet image, n is the feature of the original tablet image The number of elements included is the weight coefficient, The value of is 0.5. The weight coefficient can be adjusted according to actual needs.
[0062] S7. Detect tablet defects based on the anomaly score to obtain the tablet defect detection result.
[0063] The specific process of tablet defect detection in this step is: Compare the anomaly score with the set anomaly threshold. When the anomaly score is greater than the set anomaly threshold, it indicates that the tablet has a defect. When the anomaly score is less than or equal to the set anomaly threshold, it indicates that the tablet has no defect. The anomaly threshold in this embodiment is set according to actual needs.
[0064] To verify the effectiveness of the unsupervised learning pill defect detection method based on the diffusion model proposed in the present invention, in this embodiment, the MVTEC dataset (MVTec Anomaly Detection Dataset, anomaly detection dataset) combined with the pill defect dataset augmented by actual pictures collected on the industrial front line (specifically including 1000 normal pills of various shapes and 2000 abnormal photos of various types of defects) is used for verification.
[0065] The operating system adopted in this embodiment is ubuntu20.04, the CPU (Central Processing Unit) model is 14 vCPU Intel(R) Xeon(R) Platinum 8362 CPU @ 2.80GHz, the GPU (Graphics Processing Unit) model is RTX3090, the video memory size is 48GB, and the memory size is 64GB.
[0066] All models adopted in this embodiment are implemented based on Pytorch, and the cuda version is 11.3.
[0067] In this embodiment, the F1 score, AP (Average Precision), PRO (per-region–overlap), and AUROC (Area Under the Receiver Operating Characteristic Curve) are used to evaluate the pill defect detection results. The following is a detailed description of the F1 score, AP, PRO, and AUROC: The F1 score combines the harmonic mean of precision (the proportion of correctly predicted positive examples among the predicted positive examples) and recall (the proportion of correctly predicted positive examples among the actual positive examples). The F1 score will be high only when both precision and recall are high. The F1 score is often used for performance evaluation in binary classification and multi-classification tasks to help evaluate the overall performance of the model in terms of precision and recall.
[0068] AP calculates the average precision of the model at different recall levels, reflecting the accuracy of the model's ranking of relevant samples. The higher the AP value, the better the performance of the model in retrieving or detecting relevant objects.
[0069] PRO is the degree of overlap between different regions in the image. It is used to evaluate the accuracy of the model's detection results. By calculating the overlap ratio, the matching degree between the predicted region and the actual target region is judged, and then the performance of the model in locating the target is measured.
[0070] The ROC curve is plotted with the false positive rate on the horizontal axis and the true positive rate on the vertical axis. AUROC is the area under the ROC curve, and its value ranges from 0 to 1. The larger the AUROC value, the stronger the ability of the model to distinguish positive and negative examples, that is, the better the classification performance of the model.
[0071] In this embodiment, the method proposed in the present invention is compared with four algorithms, namely LDM (Latent Diffusion Model), PaDiM (Patch Distribution Modeling), SimpleNet (a simple and easy-to-apply neural network for detecting and locating anomalies), and DRAEM (Discriminatively Trained Reconstruction Anomaly Embedding Model) on the tablet defect dataset expanded by combining the MVTEC dataset with actual pictures collected in the industrial front line. The comparison results are shown in Table 1 and Figure 8 as shown. In Table 1, ↑ indicates that the higher the evaluation index, the better the performance of the algorithm model.
[0072] Table 1 Comparison Results
[0073] The data in Table 1 show that the method proposed in the present invention has the best detection effect, and all four evaluation indexes are the highest. Therefore, the method proposed in the present invention performs better than the four algorithms of LDM, PaDiM, SimpleNet, and DRAEM, and can complete the tablet defect detection task more excellently.
[0074] To verify the effectiveness of the feature fusion extraction module, different feature extraction networks (Resnet50, Resnet101, VGG16, VGG19, Efficient net, and Inception net) are compared with the feature fusion extraction module. In this embodiment, the feature fusion extraction module of the method proposed in the present invention is replaced with Resnet50, Resnet101, VGG16, VGG19, Efficient net, and Inception net respectively. The ablation experiment results are shown in Table 2. In Table 2, ↑ indicates that the higher the evaluation index, the better the performance of the algorithm model.
[0075] Table 2 Ablation Experiment Results
[0076] The data in Table 2 show that the feature fusion extraction module in the method proposed in the present invention is superior to the other six feature extraction network combination algorithms in all indicators, which shows that the method proposed in the present invention can achieve a more ideal effect in tablet defect detection. The average accuracy of the method proposed in the present invention is 2.5% higher than the best data in the other six algorithms, the F1 score is 2.1% higher than the best data in other algorithms, the area under the receiver operating characteristic curve is 0.9% higher than the best data in other algorithms, and the area overlap is 1.1% higher than the best data in other algorithms, which further proves the advantages of the method proposed in the present invention in the task of tablet defect detection.
[0077] The reconstruction effect diagram and tablet defect detection result diagram of the tablet defect dataset expanded by combining the MVTEC dataset with the actual images collected from the industrial front line in this embodiment are shown in the figure below. Figure 9 As shown, Figure 9 The left side shows the reconstruction effect of 4 types of contaminated tablets and the detection effect of abnormal areas (also called defect areas), and the right side shows the reconstruction effect of 4 types of defective tablets and the detection effect of abnormal areas. Figure 9 It can be seen that for tablets of different types and with different defect types, the reconstructed images of the method proposed in the present invention are similar to normal tablets, and can achieve good reconstruction effects. In terms of tablet defect detection effects, whether it is a pure color tablet with simple texture or a multicolored tablet with more complex texture, the method proposed in the present invention can perform good distinction and detection, and can also perform abnormal heat map ( Figure 9 The anomaly heat map is used for visual display.
[0078] Embodiment 4: See also Figure 10 As shown, the present invention also provides an electronic device 100 for an unsupervised learning tablet defect detection method based on a diffusion model; the electronic device 100 includes a memory 101, at least one processor 102, a computer program 103 stored in the memory 101 and executable on the at least one processor 102, and at least one communication bus 104.
[0079] The memory 101 can be used to store the computer program 103. By running or executing the computer program stored in the memory 101 and invoking the data stored in the memory 101, the processor 102 implements the steps of the unsupervised learning pill defect detection method based on the diffusion model described in Embodiment 1. The memory 101 may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the electronic device 100 (such as audio data, etc.). In addition, the memory 101 may include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices.
[0080] The at least one processor 102 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 102 may be a microprocessor or the processor 102 may also be any conventional processor, etc. The processor 102 is the control center of the electronic device 100, and connects various parts of the entire electronic device 100 through various interfaces and lines.
[0081] The memory 101 in the electronic device 100 stores multiple instructions to implement the unsupervised learning pill defect detection method based on the diffusion model. The processor 102 can execute the multiple instructions to implement: Obtain the original pill image; Perform encoding processing on the original pill image to obtain a latent space vector; Construct a diffusion model denoising network and an auxiliary reconstruction network. First, use the auxiliary reconstruction network to extract the semantic information of the original pill image, and then use the extracted semantic information of the original pill image to assist the diffusion model denoising network to reconstruct the obtained latent space vector to obtain a reconstructed pill image; Feature extraction is respectively performed on the original tablet image and the reconstructed tablet image to obtain the original tablet image features and the reconstructed tablet image features; Using the original tablet image features and the reconstructed tablet image features, calculate the Euclidean distance and cosine distance at different scales, where the scale refers to the size of the original tablet image features and the reconstructed tablet image features; Weight the cosine distance and Euclidean distance at different scales to obtain an anomaly score; Detect tablet defects based on the anomaly score to obtain the tablet defect detection result.
[0082] Example 5: If the modules / units integrated in the electronic device 100 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above method embodiments of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, and read-only memory (ROM, Read-Only Memory).
[0083] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, system, or computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0084] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementing the processes Figure 1one or more processes and / or blocks Figure 1 a device for the functions specified in one or more blocks
[0085] These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device, and the instruction device implements the processes Figure 1 one or more processes and / or blocks Figure 1 the functions specified in one or more blocks
[0086] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the processes Figure 1 one or more processes and / or blocks Figure 1 the functions specified in one or more blocks
[0087] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: the specific implementation manners of the present invention can still be modified or equivalently replaced, and any modification or equivalent replacement without departing from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
Claims
1. An unsupervised learning tablet defect detection method based on a diffusion model, characterized in that: The following steps are involved: Get the original pill image; Encode the original pill image to obtain a latent space vector; Constructing a diffusion model denoising network and an auxiliary reconstruction network, first using the auxiliary reconstruction network to extract the semantic information of the original pill image, and then using the extracted semantic information of the original pill image to assist the diffusion model denoising network to reconstruct the obtained latent space vector to obtain a reconstructed pill image; Performing feature extraction on the original tablet image and the reconstructed tablet image respectively to obtain the original tablet image features and the reconstructed tablet image features; Using the original tablet image features and the reconstructed tablet image features, calculating the cosine distance and the Euclidean distance of different scales, wherein the scale refers to the size of the original tablet image features and the reconstructed tablet image features; The cosine distance and Euclidean distance of different scales are weighted to obtain the anomaly score; The tablet defects are detected according to the abnormal scores to obtain the tablet defect detection results.
2. The unsupervised learning tablet defect detection method based on diffusion model according to claim 1, characterized in that: The construction of the diffusion model denoising network and the auxiliary reconstruction network first uses the auxiliary reconstruction network to extract the semantic information of the original pill image, and then uses the extracted semantic information of the original pill image to assist the diffusion model denoising network to reconstruct the obtained latent space vector to obtain a reconstructed pill image, specifically including: Performing noise processing on the obtained latent space vector to obtain a noisy latent space vector; Construct diffusion model denoising network and auxiliary reconstruction network; The latent space vector after adding noise is input into the diffusion model denoising network for denoising to obtain the latent space vector after denoising; The original pill image and the latent space vector after adding noise are input into the auxiliary reconstruction network. The semantic information of the original pill image is first extracted by the auxiliary reconstruction network. Then, the semantic information of the original pill image is extracted to assist the diffusion model denoising network to reconstruct the denoised latent space vector, thereby obtaining a reconstructed latent space vector. The reconstructed latent space vector is decoded to obtain a reconstructed tablet image.
3. The unsupervised learning tablet defect detection method based on diffusion model according to claim 1, characterized in that: The feature extraction of the original tablet image and the reconstructed tablet image is respectively performed to obtain the original tablet image features and the reconstructed tablet image features, specifically including: The residual network is used to extract features of the original tablet image and the reconstructed tablet image respectively, and several original tablet image features and initial reconstructed tablet image features of different scales are obtained; The feature pyramid network is used to fuse several initial original tablet image features and initial reconstructed tablet image features of different scales, respectively, to obtain several final original tablet image features and final reconstructed tablet image features of different scales.
4. The unsupervised learning tablet defect detection method based on diffusion model according to claim 3 is characterized in that: In the step of using a residual network to extract features from the original tablet image and the reconstructed tablet image respectively to obtain several original tablet image features and initial reconstructed tablet image features of different scales, a convolutional attention module is also added to each residual block of the residual network.
5. The unsupervised learning tablet defect detection method based on diffusion model according to claim 1, characterized in that: In the step of calculating cosine distances and Euclidean distances of different scales by using the original tablet image features and the reconstructed tablet image features, the scale refers to the size of the original tablet image features and the reconstructed tablet image features. The calculation formula of cosine distances of different scales is: in, For the i The cosine distance of the scale, For the i The cosine similarity of the scale, is the original pill image feature, is the reconstructed pill image feature, i is the number of scales of the original pill image features and the reconstructed pill image features; The calculation formula of Euclidean distance at different scales is: in, For the i The Euclidean distance of the scale, is the original pill image feature No. j element values, Reconstructed pill image features No. j element values, i is the number of scales of the original pill image features and the reconstructed pill image features, n is the original pill image feature The number of elements contained, j is the index variable used to traverse the original pill image features and reconstructed pill image features The element value in j The value range is [1, n ].
6. The unsupervised learning tablet defect detection method based on diffusion model according to claim 5, characterized in that: In the step of weighting the cosine distances and Euclidean distances of different scales to obtain anomaly scores, the calculation formula of the anomaly scores is: in, represents the anomaly score, For the i The upsampling factor of the scale, For the i The Euclidean distance of the scale, For the i The cosine distance of the scale, i is the number of scales of the original pill image features and the reconstructed pill image features, n is the original pill image feature The number of elements contained, is the weight coefficient, The value is between 0-1.
7. The unsupervised learning tablet defect detection method based on diffusion model according to claim 1, characterized in that: The tablet defects are detected according to the abnormal score to obtain the tablet defect detection result, which is specifically: The anomaly score is compared with the set anomaly threshold. When the anomaly score is greater than the set anomaly threshold, it indicates that the tablet is defective; when the anomaly score is less than or equal to the set anomaly threshold, it indicates that the tablet is not defective.
8. An unsupervised learning tablet defect detection system based on a diffusion model, characterized in that: It includes a data acquisition module, a latent space vector acquisition module, an image reconstruction module, a feature extraction module, a distance determination module, an abnormal score determination module and a tablet defect detection module; The data acquisition module is used to acquire the original tablet image; The latent space vector acquisition module is used to encode the original tablet image to obtain the latent space vector; The image reconstruction module is used to construct a diffusion model denoising network and an auxiliary reconstruction network. The auxiliary reconstruction network is first used to extract semantic information of the original pill image, and then the semantic information of the extracted original pill image is used to assist the diffusion model denoising network to reconstruct the obtained latent space vector to obtain a reconstructed pill image. The feature extraction module is used to extract features from the original tablet image and the reconstructed tablet image respectively to obtain features of the original tablet image and features of the reconstructed tablet image; The distance determination module is used to calculate cosine distances and Euclidean distances of different scales using original tablet image features and reconstructed tablet image features, wherein the scale refers to the size of the original tablet image features and the reconstructed tablet image features; The anomaly score determination module is used to weight the cosine distance and Euclidean distance of different scales to obtain an anomaly score; The tablet defect detection module is used to detect tablet defects according to the abnormality score to obtain a tablet defect detection result.
9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the unsupervised learning tablet defect detection method based on the diffusion model as claimed in any one of claims 1 to 7 are implemented.
10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the unsupervised learning tablet defect detection method based on a diffusion model according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Image defect detection method and device and related equipment
CN117788408A
Defect detection method, device, system and equipment based on image processing, medium and product
CN118297884A
Multi-class defect detection method and device based on variational auto-encoder and denoising network
CN119516249A
Cited By
Tablet disintegration detection method and system
CN121121658A