Defect detection method, device and equipment for medical consumables and medium

By employing unsupervised learning and multimodal fusion, pseudo-defect samples are generated using multimodal data for feature extraction and alignment. This addresses the shortcomings of traditional detection methods in terms of real-time performance and generalization ability in medical consumables testing, achieving efficient and low-cost defect detection.

CN121982007APending Publication Date: 2026-05-05SUZHOU HUANQIU MEDICAL TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SUZHOU HUANQIU MEDICAL TECHNOLOGY CO LTD
Filing Date
2026-01-29
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Traditional detection methods are difficult to meet the real-time and consistency requirements of mass production of medical consumables, and single-modal detection has limited ability to detect occlusion, reflection and material differences. Existing deep learning methods rely on large-scale labeled data and have insufficient generalization ability.

Method used

By combining unsupervised learning and multimodal fusion, pseudo-defect samples are generated by acquiring multimodal data (visible light images, infrared images, and 3D point cloud information), and feature extraction and alignment are performed. Defect detection is then performed using a self-supervised training model.

Benefits of technology

It improves the accuracy, robustness, and applicability of defect detection, reduces reliance on labeled data, adapts to the lightweight and low-cost requirements of industrial production environments, and expands the detection coverage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121982007A_ABST
    Figure CN121982007A_ABST
Patent Text Reader

Abstract

The invention discloses a defect detection method and device for medical consumables, equipment and a medium. The method comprises the following steps: acquiring multi-modal data of the defect-free medical consumables as normal sample data; the multi-modal data comprises a visible light image, an infrared image and three-dimensional point cloud information; interference is added to the normal sample data to generate pseudo-defect sample data, and a multi-modal training data pair is determined according to the normal sample data and the pseudo-defect sample data; performing feature extraction on the multi-modal training data to obtain multi-modal feature information, and performing feature alignment on the multi-modal feature information through comparative learning to obtain target feature information; inputting the target feature information into a defect detection model for self-supervised training, and performing defect detection on the to-be-detected medical consumables by using the trained defect detection model; and the defect detection model performs defect detection based on the difference between the target feature information and the reconstruction result thereof. According to the scheme, the defect detection coverage range can be expanded while the defect labeling dependence is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, and in particular to a method, apparatus, equipment and medium for defect detection of medical consumables. Background Technology

[0002] Medical consumables are typically characterized by mass production, relatively low unit value but significant safety responsibility, diverse defect types, and a low percentage of defective samples. During production, packaging, transportation, and storage, medical consumables may develop defects such as cracks, gaps, dents, peeling, wear, corrosion, deformation, and abnormal surface textures. Failure to detect these defects in a timely manner may pose risks to clinical use and quality traceability.

[0003] Traditional inspection methods often rely on manual visual inspection or single-modal sensor detection. Manual visual inspection is affected by subjective experience, lighting conditions, and fatigue, and it is difficult to meet the real-time and consistency requirements of high-speed production lines. Single-modal detection (such as RGB mode only) is only sensitive to certain defect types, but its ability to detect occlusion, reflection, and abnormal material differences is limited.

[0004] With the development of computer vision algorithms, deep learning-based visual detection methods can achieve more accurate detection and classification through data learning and model training, and have good robustness and versatility. However, supervised defect detection often relies heavily on large-scale, accurately labeled data in practical applications. The acquisition and labeling of abnormal samples are costly and time-consuming, and the generalization ability when facing new or complex defects still needs to be improved. Summary of the Invention

[0005] This invention provides a method, apparatus, equipment, and medium for defect detection of medical consumables. By combining unsupervised learning with multimodal fusion for defect detection of medical consumables, it can expand the coverage of defect detection while reducing the dependence on defect labeling, which helps to improve the accuracy, robustness, and applicability of defect detection.

[0006] According to one aspect of the present invention, a defect detection method for medical consumables is provided, the method comprising: Multimodal data of defect-free medical consumables are acquired as normal sample data; wherein, the multimodal data includes visible light images, infrared images, and three-dimensional point cloud information; Add interference to the normal sample data to generate pseudo-defect sample data, and determine multimodal training data pairs based on the normal sample data and the pseudo-defect sample data; Multimodal feature information is obtained by extracting features from the multimodal training data pairs, and target feature information is obtained by aligning the multimodal feature information through contrastive learning. The target feature information is input into the defect detection model for self-supervised training, so as to use the trained defect detection model to detect defects in the medical consumables under test; wherein, the defect detection model performs defect detection based on the difference between the target feature information and its reconstruction result.

[0007] According to another aspect of the present invention, a defect detection device for medical consumables is provided, the device comprising: The normal sample acquisition module is used to acquire multimodal data of defect-free medical consumables as normal sample data; wherein, the multimodal data includes visible light images, infrared images and three-dimensional point cloud information; A pseudo-defect sample construction module is used to add interference to the normal sample data to generate pseudo-defect sample data, and to determine multimodal training data pairs based on the normal sample data and the pseudo-defect sample data; The feature alignment module is used to extract features from the multimodal training data pairs to obtain multimodal feature information, and to align the multimodal feature information to obtain target feature information through contrastive learning. The model training module is used to input the target feature information into the defect detection model for self-supervised training, so as to use the trained defect detection model to detect defects in the medical consumables under test; wherein, the defect detection model performs defect detection based on the difference between the target feature information and its reconstruction result.

[0008] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the defect detection method for medical consumables according to any embodiment of the present invention.

[0009] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the defect detection method for medical consumables according to any embodiment of the present invention.

[0010] The technical solution of this invention first acquires multimodal data of defect-free medical consumables as normal sample data; wherein, the multimodal data includes visible light images, infrared images, and three-dimensional point cloud information; then, interference is added to the normal sample data to generate pseudo-defect sample data, and multimodal training data pairs are determined based on the normal sample data and the pseudo-defect sample data; next, feature extraction is performed on the multimodal training data pairs to obtain multimodal feature information, and the multimodal feature information is aligned through comparative learning to obtain target feature information; then, the target feature information is input into a defect detection model for self-supervised training, so as to use the trained defect detection model to detect defects in the medical consumables under test; wherein, the defect detection model performs defect detection based on the difference between the target feature information and its reconstruction result. This technical solution combines unsupervised learning with multimodal fusion for defect detection in medical consumables. By complementing multimodal information, it addresses the problem that single-modal methods cannot cover multiple defect types. Defect detection based on unsupervised learning can learn the distribution of normal samples and identify anomalies that deviate from the distribution. It is more suitable for the needs of lightweight and low labeling costs in actual industrial production environments. It can expand the coverage of defect detection while reducing the dependence on defect labeling, which helps to improve the accuracy, robustness and applicability of defect detection.

[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 This is a flowchart of a defect detection method for medical consumables according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of a defect detection model provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the overall process of a defect detection method for medical consumables provided in an embodiment of the present invention; Figure 4 This is a flowchart of another defect detection method for medical consumables provided according to an embodiment of the present invention; Figure 5 This is a schematic diagram of a defect detection device for medical consumables according to an embodiment of the present invention; Figure 6 This is a schematic diagram of the structure of an electronic device that implements a defect detection method for medical consumables according to an embodiment of the present invention. Detailed Implementation

[0014] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0015] It should be noted that the terms "first," "second," "target," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0016] Example 1 Figure 1 This is a flowchart of a defect detection method for medical consumables provided in Embodiment 1 of the present invention. This embodiment is applicable to unsupervised multimodal defect detection of medical consumables. The method can be executed by a defect detection device for medical consumables, which can be implemented in hardware and / or software and can be configured in an electronic device with data processing capabilities. Figure 1 As shown, the method includes: S110, acquire multimodal data of defect-free medical consumables as normal sample data; wherein, the multimodal data includes visible light images, infrared images and three-dimensional point cloud information.

[0017] In this embodiment, a multimodal data acquisition platform integrating a visible light camera, an infrared camera, and a 3D point cloud sensor can be pre-built. For example, the visible light camera can be an industrial camera, and the 3D point cloud sensor can be a structured light depth camera. By writing a microcontroller-based synchronous triggering program, the multimodal data acquisition platform can be used to acquire multimodal data from defect-free medical consumables, achieving simultaneous acquisition of visible light images, infrared images, and 3D point cloud information, and using the acquired multimodal data as normal sample data. For example, the visible light image is an RGB image. It should be noted that by introducing the infrared modality, the types of defects existing inside the medical consumables can be better identified, broadening the range of defect types that can be detected.

[0018] S120, add interference to normal sample data to generate pseudo-defect sample data, and determine multimodal training data pairs based on normal sample data and pseudo-defect sample data.

[0019] In this embodiment, after determining the normal sample data, interference can be added to the normal sample data to generate corresponding pseudo-defect sample data, while ensuring semantic consistency of the three-modal anomalies. Optionally, adding interference to the normal sample data to generate pseudo-defect sample data includes: generating a defect mask and implanting texture perturbation in a visible light image to obtain pseudo-defect sample data corresponding to the visible light image; generating radiation intensity interference in an infrared image to obtain pseudo-defect sample data corresponding to the infrared image; and adding deformation noise to a local area of ​​the point cloud in the three-dimensional point cloud information to obtain pseudo-defect sample data corresponding to the three-dimensional point cloud information.

[0020] Specifically, for visible light images, defect masks can be generated using Perlin noise or random cropping and pasting, and texture perturbations can be implanted to construct pseudo-defect sample data corresponding to the visible light image. For infrared images, radiation intensity interference can be generated at the same location (i.e., the location of the defect mask) to construct pseudo-defect sample data corresponding to the infrared image. The amplitude of the interference or perturbation can be adaptively sampled within a preset range. For 3D point cloud information, given the original point cloud P, Represents the first in P For each point, pseudo-defect sample data can be constructed through the following process: (1) Calculate the FPFH (Fast Point Feature Histograms) features of each point cloud: ,in, express (2) Select the point with the highest FPFH feature as the center point of the local region: (3) Based on the center point c, randomly select a total point cloud r∈[1%,10%] around c as a local region: Where d is the distance threshold, ensuring The approximation r includes the midpoint of P. (4) Randomly select a variance of The value, using Generate a normally distributed random noise with a mean of 0. ,in, (5) Add noise to each dimension of a point cloud in a local area to generate anomaly regions: Repeat this process for all point clouds in the local area to obtain the set of outliers: (6) Merge A and P to obtain pseudo-defect sample data: .

[0021] After generating pseudo-defect sample data in multimodal modes, each set of normal sample data and its corresponding pseudo-defect sample data can be used as a training data pair, ultimately obtaining multimodal training data pairs (including multiple training data pairs) for subsequent defect detection model training based on these multimodal training data pairs. It should be noted that this invention synthesizes pseudo-defect sample data (including RGB texture perturbation, infrared radiation interference, and point cloud deformation noise) with consistent trimodal semantics based on normal sample data, thereby constructing normal-pseudo-defect training data pairs for subsequent defect detection model training. This enhances the coverage of scarce defects, provides stable anomalous signals for unsupervised learning, improves the detection capability of rare defects, and reduces the dependence on scarce samples and the cost of manual annotation.

[0022] S130: Feature extraction is performed on the multimodal training data pairs to obtain multimodal feature information. The multimodal feature information is then aligned through contrastive learning to obtain the target feature information.

[0023] In this embodiment, after determining the multimodal training data pairs, a pre-trained multimodal feature extractor can be used to extract features from the multimodal training data pairs to obtain multimodal feature information. To eliminate differences in imaging mechanisms of different sensors and ensure semantic consistency of corresponding positions across modalities, this invention further introduces a multimodal feature alignment process. Contrastive learning is used to map the features of each modality to a unified feature space and complete feature alignment, thereby improving the cross-modal generalization ability of the defect detection model. Optionally, the multimodal feature information is aligned to obtain target feature information through contrastive learning, including: mapping the multimodal feature information to the target feature space through nonlinear mapping; and performing self-supervised contrastive learning on the multimodal feature information in the target feature space based on a contrastive loss function to obtain the target feature information.

[0024] Specifically, assuming that the original features obtained from the RGB image, IR image, and 3D point cloud information through the multimodal feature extractor are as follows: , and By using a nonlinear mapping, these original features are mapped to a unified target feature space, and the mapped features can be represented as follows: , , .in, , and These are the mapping functions corresponding to the three modes, , and denoted by , where is the number of samples in the three modal datasets, and d is the dimension of the mapped features. For example, the nonlinear mapping of the three modalities can be implemented using an MLP (Multilayer Perceptron), where 3D point cloud features are sampled, grouped, interpolated, projected, and then mapped to the target feature space using an MLP.

[0025] After mapping multimodal feature information to the target feature space through nonlinear mapping, self-supervised contrastive learning can be performed on the multimodal feature information in the target feature space based on a contrastive loss function, thereby obtaining the feature-aligned target feature information. This invention employs a contrastive learning strategy to achieve effective alignment between multimodal features. To this end, a contrastive loss function is designed to maximize the feature similarity of the same defect in different modalities and minimize the similarity between different defects. Optionally, the contrastive loss function includes a local contrastive loss function and a global contrastive loss function. The local contrastive loss function measures the feature alignment effect between the two modal feature information, while the global contrastive loss function measures the feature alignment effect of the multimodal feature information. The global contrastive loss function is constructed based on the local contrastive loss function.

[0026] For example, the local contrastive loss function is constructed based on the InfoNCE (Information Noise-Contrastive Estimation) contrastive loss function to achieve intra-modal feature alignment. Specifically, assuming... and If we represent the features of two modalities in the target feature space, then the local contrastive loss function can be expressed as follows: ; in, This is a temperature parameter used to adjust the discriminative power between features. By optimizing the local contrast loss function, similar defects from different modalities can be mapped to a unified feature space, while dissimilar defects are mapped to distant regions. Furthermore, to enhance the global feature alignment effect of RGB images, IR images, and 3D point cloud information, this invention introduces a global contrast loss function. Feature alignment between modalities is achieved. The global contrastive loss function integrates contrastive learning from three modalities: IR-RGB, IR-3D, and RGB-3D, ensuring consistency of multimodal data in the global feature space. Specifically, it can be represented as follows: ; in, These are the weighting coefficients for each pair of modes. This is a local contrastive loss function for each modality pair. In this way, not only are the features between each modality pair aligned, but global alignment is also achieved across the entire modality combination, thereby improving the effectiveness of multimodal data alignment. Based on the contrastive loss function, self-supervised contrastive learning is performed by combining data augmentation with rotation, pruning, and scaling transformations to improve the generalization ability across different batches of medical consumables.

[0027] S140, Input the target feature information into the defect detection model for self-supervised training, so as to use the trained defect detection model to detect defects in the medical consumables under test; wherein, the defect detection model performs defect detection based on the difference between the target feature information and its reconstruction result.

[0028] The defect detection model can refer to a neural network model capable of detecting defects in medical consumables. Specifically, it can include a reconstruction network and a multimodal discrimination module. The reconstruction network can be used for data reconstruction, and the multimodal discrimination module can be used for defect localization. Optionally, the reconstruction network adopts a U-shaped network structure and includes an encoder and a decoder. The encoder is constructed based on depthwise separable convolution and adaptive discrete wavelet transform. The multimodal discrimination module includes an attention mechanism fusion module and a discriminator.

[0029] like Figure 2 As shown, the defect detection model includes a reconstruction network and a multimodal discrimination module. The reconstruction network adopts a UNET structure, reconstructing data through four downsampling passes and four upsampling passes. Specifically, the reconstruction network includes an encoder (corresponding to the downsampling path) and a decoder (corresponding to the upsampling path). The encoder's backbone network uses depthwise separable convolutions to reduce the number of network parameters, making it easier to deploy in actual industrial production. The first three layers of the encoder introduce an adaptive discrete wavelet transform module, which enhances the transmission of effective frequency information while suppressing noise propagation. The decoder introduces skip connections, concatenating the encoder features after center clipping with the decoder features after upsampling to enhance information sharing. For example, the attention mechanism fusion module is built on a fully connected layer and dynamically allocates fusion weights for the three modalities based on the attention mechanism. Even if a single modality is affected by reflection, occlusion, or material differences, it can still maintain stable output, improving robustness and reliability. This allows the discriminator to fully utilize the complementary information of the three modalities to achieve accurate detection of defect areas.

[0030] In this embodiment, after determining the target feature information, it can be input into the defect detection model for self-supervised training. The purpose of model training is to enable the reconstruction network to learn the normal appearance, spectrum, and geometric distribution of medical consumables, so as to use the trained defect detection model to detect defects in the medical consumables under test. The medical consumables under test can refer to those requiring defect detection. It should be noted that defect detection based on the reconstruction network is based on a prior assumption that the error difference between the reconstructed defect area and the normal area is significant. Therefore, the defect area can be detected and located by comparing the error magnitude before and after reconstruction. The error magnitude before and after reconstruction can be calculated using the difference between the target feature information and its reconstruction result.

[0031] In this embodiment, optionally, the target feature information is input into the defect detection model for self-supervised training, including: inputting the target feature information into the reconstruction network in the defect detection model, reconstructing the target feature information through the reconstruction network to obtain a reconstruction result, and determining the multimodal reconstruction error based on the difference between the target feature information and the reconstruction result; inputting the multimodal reconstruction error into the multimodal discrimination module in the defect detection model, adaptively weighting and fusing the multimodal reconstruction error through the multimodal discrimination module to obtain a fusion anomaly result, and determining the defect segmentation result and defect score based on the fusion anomaly result.

[0032] For example, with Figure 2 Taking the defect detection model in the example, the features of the three modalities after feature alignment (i.e., target feature information, see...) Figure 2 The infrared features, visible light features, and point cloud features in the image are reconstructed using the same reconstruction network to reconstruct images of three modalities. The process is as follows: For each of the first three layers of the encoder, the target feature information U is first decomposed into four frequency components in the frequency domain using discrete wavelet transform, which can be represented as: Where DWT represents Discrete Wavelet Transform. These four frequency components are then concatenated, passed through a 1×1 convolutional layer, and a batch normalization and ReLU (Rectified Linear Unit) activation function are applied to the convolutional layer output. This maps the decomposed features to an embedding space for better subsequent filtering. This process can be represented as: .in, Indicates batch normalization. This indicates a convolution operation with a 1×1 kernel. Next, we will... Average pooling and max pooling operations are performed along the channel, and the results of the two pooling operations are concatenated, which can be represented as follows: Furthermore, A 4×4 convolutional layer is fed into the array to learn the spatial mask, resulting in filtered frequency features, which can be represented as follows: .in, For the Sigmoid function, This indicates a convolution operation with a 4×4 kernel. Finally, The four channels are decomposed into four frequency components, and then restored to their initial size through an inverse discrete wavelet transform, which can be specifically expressed as: Where IWT stands for Inverse Discrete Wavelet Transform. By performing four downsampling operations on the encoder, the output features of each layer of the encoder can be obtained. For the decoder, by performing four upsampling operations on the decoder, the output features of each layer of the decoder can be obtained. Then, by using skip connections, the output features of each layer of the encoder after center clipping are concatenated with the corresponding upsampled output features of each layer of the decoder. Finally, the encoder outputs the reconstruction results of each mode.

[0033] In this embodiment, self-supervised training is performed based on the loss function of the reconstruction network. For example, the loss function of the reconstruction network can be expressed as: ,in, and These represent the input (i.e., target feature information) and output (i.e., reconstruction result) of the reconstruction network, respectively. This represents the mean square error between the input and output of the reconstructed network. This represents the structural similarity between the input and output of the reconstructed network. The loss function represents the wavelet transform. and They are respectively and The weighting coefficients. For example, , For example, the loss function of wavelet transform can be expressed as: ,in, It represents 4 frequency components.

[0034] After obtaining the reconstruction results of each modality, the target feature information can be subtracted from the corresponding reconstruction result, and the difference can be taken as the reconstruction error under the corresponding modality, thus obtaining the multimodal reconstruction error, which is used as the discrimination criterion for defect detection. Next, the multimodal reconstruction error can be input into the multimodal discrimination module in the defect detection model, and the multimodal reconstruction error can be adaptively weighted and fused through the attention mechanism fusion module, thereby obtaining the fused anomaly result. It should be noted that, in order to achieve the fusion of multimodal detection results and improve robustness, this invention designs an attention mechanism to dynamically allocate the weights of each modality reconstruction error. For example, assuming that the reconstruction errors corresponding to RGB image, IR image and 3D point cloud information are respectively represented as... , and The attention of the three modalities can then be calculated using the following formula: , where parameters This can be learned. Then, based on the attention of the three modalities, the reconstruction errors of the three modalities are weighted and fused, which can be specifically expressed as: .in, This is to resolve abnormal results from fusion.

[0035] After obtaining the fusion anomaly results, they can be input into the discriminator, which outputs the defect segmentation result and defect score. Specifically, the maximum value is found from the fusion anomaly results, the mean value of the neighborhood containing the maximum value is calculated, and then the mean value is compared with multiple pre-set reference thresholds. Based on the comparison results, the defect segmentation result and defect score are determined. The reference thresholds serve as the basis for determining the existence of defects; each reference threshold corresponds to a different defect confidence level. The higher the defect confidence level, the greater the probability of a defect. If the mean value of a region is greater than a certain reference threshold, a defect mask is generated for the corresponding region as the defect segmentation result, thereby achieving defect detection and localization. The reference threshold closest to the mean value is used as the defect score for the corresponding region. Furthermore, a warning level can be pre-set for each reference threshold (the larger the reference threshold, the higher the corresponding warning level). After obtaining the defect score, the corresponding warning level is determined based on the defect score, and a warning is issued.

[0036] like Figure 3 As shown, firstly, multimodal data of defect-free medical consumables (including visible light images, infrared images, and 3D point cloud information) are acquired as normal sample data. During the data preparation stage, the normal sample data undergoes denoising, enhancement, and downsampling. Then, the prepared normal sample data is input into a pseudo-anomaly generation module. This module adds interference to the normal sample data to generate pseudo-defect sample data, and determines multimodal training data pairs based on the normal sample data and pseudo-defect sample data. Next, multimodal feature extraction and alignment are performed on the multimodal training data to obtain target feature information. Finally, the aligned target feature information is input into the defect detection model (i.e.,...). Figure 3 The detection model is self-supervised to enable the trained defect detection model to detect defects in the medical consumables under test.

[0037] In this invention, a defect detection model is constructed with a lightweight reconstruction network as its core, and the reconstruction error is used as the basis for defect discrimination to achieve defect detection and localization. During the model training phase, self-supervised training is performed using normal samples and simulated abnormal samples (i.e., pseudo-defect samples), enabling the reconstruction network to accurately reconstruct normal regions while generating significant reconstruction errors for abnormal regions. During the inference phase, only the multimodal data of the medical consumables to be tested is input. The reconstruction network outputs the reconstruction results and reconstruction errors for each modality, and the multimodal discrimination module adaptively weights and fuses the reconstruction errors of each modality to obtain the final defect segmentation result and defect score. Since the training data required by the reconstruction network can be obtained solely from normal samples, the dependence on scarce samples and the cost of manual annotation are reduced. The reconstruction network itself employs a lightweight design such as depthwise separable convolution and skip connections, balancing reconstruction accuracy and computational efficiency, making it suitable for real-time deployment on medical consumable production lines.

[0038] The technical solution of this invention first acquires multimodal data of defect-free medical consumables as normal sample data; wherein, the multimodal data includes visible light images, infrared images, and three-dimensional point cloud information; then, interference is added to the normal sample data to generate pseudo-defect sample data, and multimodal training data pairs are determined based on the normal sample data and the pseudo-defect sample data; next, feature extraction is performed on the multimodal training data pairs to obtain multimodal feature information, and the multimodal feature information is aligned through comparative learning to obtain target feature information; then, the target feature information is input into a defect detection model for self-supervised training, so as to use the trained defect detection model to detect defects in the medical consumables under test; wherein, the defect detection model performs defect detection based on the difference between the target feature information and its reconstruction result. This technical solution combines unsupervised learning with multimodal fusion for defect detection in medical consumables. By complementing multimodal information, it addresses the problem that single-modal methods cannot cover multiple defect types. Defect detection based on unsupervised learning can learn the distribution of normal samples and identify anomalies that deviate from the distribution. It is more suitable for the needs of lightweight and low labeling costs in actual industrial production environments. It can expand the coverage of defect detection while reducing the dependence on defect labeling, which helps to improve the accuracy, robustness and applicability of defect detection.

[0039] Example 2 Figure 4 This is a flowchart of a defect detection method for medical consumables provided in Embodiment 2 of the present invention. This embodiment is based on the above embodiment and optimized. Specifically, the optimization includes: after obtaining multimodal data of defect-free medical consumables as normal sample data, the method further includes: preprocessing the normal sample data; wherein, the preprocessing includes noise reduction and enhancement, spatial registration, size alignment, and data partitioning.

[0040] like Figure 4 As shown, the method in this embodiment specifically includes the following steps: S210, acquire multimodal data of defect-free medical consumables as normal sample data; wherein, the multimodal data includes visible light images, infrared images and three-dimensional point cloud information.

[0041] S220 preprocesses the normal sample data; the preprocessing includes noise reduction and enhancement, spatial registration, size alignment, and data partitioning.

[0042] In this embodiment, to ensure the quality and validity of the collected data, after acquiring normal sample data, it is necessary to preprocess the normal sample data, specifically including denoising enhancement, spatial registration, size alignment, and data partitioning. Data partitioning can include the partitioning of image blocks and point cloud blocks. For example, for the denoising enhancement operation, the visible light image is denoised, the infrared image is denoised and super-resolution image enhanced, and the 3D point cloud information is denoised and downsampled. During spatial registration, the intrinsic and extrinsic parameters of each modality can be obtained through multi-sensor calibration, and the depth point cloud is projected onto the visible light plane to establish a pixel-level correspondence. A random sampling consensus algorithm is used to eliminate mismatched points to optimize the registration result.

[0043] For example, taking the matching of RGB and IR images as an example, Gaussian blur is used to separate the RGB image and the depth image. The RGB and IR images were adjusted to the same resolution to reduce the impact of resolution differences on registration accuracy. Intrinsic parameter matrices for both the RGB and IR cameras were obtained through camera intrinsic parameter calibration. , and distortion coefficient , And calculate the pixel coordinates of the undistorted image. , Acquiring depth images using a depth camera The depth information of each pixel is integrated into a 3D array, representing the pixel's position in the 3D world coordinate system. This is achieved through the intrinsic parameter matrix of the RGB camera. and depth value Pixel coordinates in an RGB image Points mapped to the 3D world coordinate system : 3D points are represented using homogeneous coordinates. And through the relative coordinate transformation matrix between the IR camera and the RGB camera. Projecting 3D points onto an IR image: .in, Including rotation matrix Translation matrix , is used to represent the coordinate transformation between RGB cameras and IR cameras.

[0044] The following formulas can be used to calculate the relationship between the RGB camera and the IR camera and the world coordinate system. Relationship: , .in, This represents the rotation matrix from the RGB camera coordinate system to the world coordinate system. This represents the rotation matrix from the IR camera coordinate system to the world coordinate system. and These represent the RGB camera coordinates and the IR camera coordinates, respectively. Since the RGB and IR cameras are located close to each other, their camera coordinate systems can be considered to be equidistant, thus yielding the relative coordinate transformation matrix between the two cameras: .

[0045] S230, add interference to normal sample data to generate pseudo-defect sample data, and determine multimodal training data pairs based on normal sample data and pseudo-defect sample data.

[0046] S240: Feature extraction is performed on the multimodal training data pairs to obtain multimodal feature information. The multimodal feature information is then aligned through contrastive learning to obtain the target feature information.

[0047] S250, the target feature information is input into the defect detection model for self-supervised training, so as to use the trained defect detection model to detect defects in the medical consumables under test; wherein, the defect detection model performs defect detection based on the difference between the target feature information and its reconstruction result.

[0048] The technical solution of this invention combines unsupervised learning with multimodal fusion for defect detection in medical consumables. By leveraging complementary multimodal information, it addresses the limitation of single-modal methods in covering multiple defect types. Unsupervised learning-based defect detection learns the distribution of normal samples and identifies anomalies that deviate from this distribution, making it more suitable for the lightweight and low-annotation-cost requirements of actual industrial production environments. It expands the defect detection coverage while reducing reliance on defect annotation, thus improving the accuracy, robustness, and applicability of defect detection. Furthermore, after acquiring multimodal data of defect-free medical consumables as normal sample data, this technical solution performs preprocessing operations on the normal sample data, including noise reduction and enhancement, spatial registration, size alignment, and data segmentation. This ensures the quality and validity of the collected data, further improving the accuracy of defect detection.

[0049] Example 3 Figure 5This is a schematic diagram of a defect detection device for medical consumables provided in Embodiment 3 of the present invention. This device can execute the defect detection method for medical consumables provided in any embodiment of the present invention, and possesses the corresponding functional modules and beneficial effects for executing the method. For example... Figure 5 As shown, the device includes: The normal sample acquisition module 310 is used to acquire multimodal data of defect-free medical consumables as normal sample data; wherein, the multimodal data includes visible light images, infrared images and three-dimensional point cloud information; The pseudo-defect sample construction module 320 is used to add interference to the normal sample data to generate pseudo-defect sample data, and to determine multimodal training data pairs based on the normal sample data and the pseudo-defect sample data. The feature alignment module 330 is used to extract features from the multimodal training data pairs to obtain multimodal feature information, and to align the multimodal feature information to obtain target feature information through contrastive learning. The model training module 340 is used to input the target feature information into the defect detection model for self-supervised training, so as to use the trained defect detection model to detect defects in the medical consumables under test; wherein, the defect detection model performs defect detection based on the difference between the target feature information and its reconstruction result.

[0050] Optionally, the pseudo-defect sample construction module 320 is used for: A defect mask is generated in a visible light image and texture perturbation is implanted to obtain pseudo-defect sample data corresponding to the visible light image. Radiation intensity interference is generated in the infrared image to obtain pseudo-defect sample data corresponding to the infrared image; Deformation noise is added to local area point clouds in 3D point cloud information to obtain pseudo-defect sample data corresponding to 3D point cloud information.

[0051] Optionally, the feature alignment module 330 is used for: The multimodal feature information is mapped to the target feature space through nonlinear mapping; The target feature information is obtained by performing self-supervised contrastive learning on the multimodal feature information of the target feature space based on the contrastive loss function.

[0052] Optionally, the contrast loss function includes a local contrast loss function and a global contrast loss function. The local contrast loss function is used to measure the feature alignment effect of two modal feature information, and the global contrast loss function is used to measure the feature alignment effect of multimodal feature information. The global contrast loss function is constructed based on the local contrast loss function.

[0053] Optionally, the model training module 340 is used for: The target feature information is input into the reconstruction network in the defect detection model. The reconstruction network is used to reconstruct the target feature information to obtain the reconstruction result. The multimodal reconstruction error is determined based on the difference between the target feature information and the reconstruction result. The multimodal reconstruction error is input into the multimodal discrimination module in the defect detection model. The multimodal discrimination module performs adaptive weighted fusion on the multimodal reconstruction error to obtain a fusion anomaly result. Based on the fusion anomaly result, the defect segmentation result and defect score are determined.

[0054] Optionally, the reconstruction network adopts a U-shaped network structure, and the reconstruction network includes an encoder and a decoder. The encoder is constructed based on depthwise separable convolution and adaptive discrete wavelet transform. The multimodal discrimination module includes an attention mechanism fusion module and a discriminator.

[0055] Optionally, the apparatus further includes: a data preprocessing module, used for: After acquiring multimodal data of defect-free medical consumables as normal sample data, the normal sample data is preprocessed; wherein, the preprocessing includes noise reduction and enhancement, spatial registration, size alignment and data partitioning.

[0056] The defect detection device for medical consumables provided in this embodiment of the invention can execute the defect detection method for medical consumables provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method.

[0057] Example 4 Figure 6 A schematic diagram of an electronic device 10, which can be used to implement embodiments of the present invention, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0058] like Figure 6As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0059] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0060] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as defect detection methods for medical consumables.

[0061] In some embodiments, the defect detection method for medical consumables may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the defect detection method for medical consumables described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the defect detection method for medical consumables by any other suitable means (e.g., by means of firmware).

[0062] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0063] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0064] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0065] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0066] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0067] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0068] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0069] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A defect detection method for medical consumables, characterized in that, The method includes: Multimodal data of defect-free medical consumables are acquired as normal sample data; wherein, the multimodal data includes visible light images, infrared images, and three-dimensional point cloud information; Add interference to the normal sample data to generate pseudo-defect sample data, and determine multimodal training data pairs based on the normal sample data and the pseudo-defect sample data; Multimodal feature information is obtained by extracting features from the multimodal training data pairs, and target feature information is obtained by aligning the multimodal feature information through contrastive learning. The target feature information is input into the defect detection model for self-supervised training, so as to use the trained defect detection model to detect defects in the medical consumables under test; wherein, the defect detection model performs defect detection based on the difference between the target feature information and its reconstruction result.

2. The method according to claim 1, characterized in that, Adding interference to the normal sample data to generate pseudo-defect sample data includes: A defect mask is generated in a visible light image and texture perturbation is implanted to obtain pseudo-defect sample data corresponding to the visible light image. Radiation intensity interference is generated in the infrared image to obtain pseudo-defect sample data corresponding to the infrared image; Deformation noise is added to local area point clouds in 3D point cloud information to obtain pseudo-defect sample data corresponding to 3D point cloud information.

3. The method according to claim 1, characterized in that, The target feature information is obtained by aligning the multimodal feature information through contrastive learning, including: The multimodal feature information is mapped to the target feature space through nonlinear mapping; The target feature information is obtained by performing self-supervised contrastive learning on the multimodal feature information of the target feature space based on the contrastive loss function.

4. The method according to claim 3, characterized in that, The contrast loss function includes a local contrast loss function and a global contrast loss function. The local contrast loss function is used to measure the feature alignment effect of two modal feature information, and the global contrast loss function is used to measure the feature alignment effect of multimodal feature information. The global contrast loss function is constructed based on the local contrast loss function.

5. The method according to claim 1, characterized in that, The target feature information is input into the defect detection model for self-supervised training, including: The target feature information is input into the reconstruction network in the defect detection model. The reconstruction network is used to reconstruct the target feature information to obtain the reconstruction result. The multimodal reconstruction error is determined based on the difference between the target feature information and the reconstruction result. The multimodal reconstruction error is input into the multimodal discrimination module in the defect detection model. The multimodal discrimination module performs adaptive weighted fusion on the multimodal reconstruction error to obtain a fusion anomaly result. Based on the fusion anomaly result, the defect segmentation result and defect score are determined.

6. The method according to claim 5, characterized in that, The reconstruction network adopts a U-shaped network structure and includes an encoder and a decoder. The encoder is constructed based on depthwise separable convolution and adaptive discrete wavelet transform. The multimodal discrimination module includes an attention mechanism fusion module and a discriminator.

7. The method according to any one of claims 1-6, characterized in that, After obtaining multimodal data of defect-free medical consumables as normal sample data, the following is also included: The normal sample data is preprocessed; wherein the preprocessing includes noise reduction and enhancement, spatial registration, size alignment and data partitioning.

8. A defect detection device for medical consumables, characterized in that, The device includes: The normal sample acquisition module is used to acquire multimodal data of defect-free medical consumables as normal sample data; wherein, the multimodal data includes visible light images, infrared images and three-dimensional point cloud information; A pseudo-defect sample construction module is used to add interference to the normal sample data to generate pseudo-defect sample data, and to determine multimodal training data pairs based on the normal sample data and the pseudo-defect sample data; The feature alignment module is used to extract features from the multimodal training data pairs to obtain multimodal feature information, and to align the multimodal feature information to obtain target feature information through contrastive learning. The model training module is used to input the target feature information into the defect detection model for self-supervised training, so as to use the trained defect detection model to detect defects in the medical consumables under test; wherein, the defect detection model performs defect detection based on the difference between the target feature information and its reconstruction result.

9. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the defect detection method for medical consumables according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the defect detection method for medical consumables as described in any one of claims 1-7.

Citation Information

Cited By

  • A mask substrate defect image recognition method based on deep learning

    CN122244050A