A method for detecting defects in rubber stopper production based on image processing

By using image processing methods and combining images and point cloud data under multiple pressure conditions, simultaneous detection of rubber stopper appearance and airtightness defects was achieved, solving the problem of fragmented detection functions in existing technologies and improving detection efficiency and accuracy.

CN122490445APending Publication Date: 2026-07-31SHANDONG JIKANG PHARM PACKAGING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANDONG JIKANG PHARM PACKAGING CO LTD
Filing Date
2026-06-17
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In current rubber stopper production, the detection of appearance defects and the detection of airtightness defects are independent of each other, making it impossible to achieve online non-destructive full inspection, and the defect location information is incomplete.

Method used

An image processing-based approach is adopted to acquire image data and 3D point cloud data of rubber stoppers under multiple preset pressure states. Multi-scale features are extracted using a dual-stream coding network and cross-modal fusion is performed by combining pressure feature vectors. A dual-task decoding network is used to simultaneously detect appearance defects and airtightness defects, and the defect segmentation mask and 3D heat map are output.

Benefits of technology

It enables simultaneous detection of appearance and airtightness defects in rubber stoppers, improving the accuracy and efficiency of detection. It does not require damaging the rubber stopper and is suitable for large-scale industrial production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122490445A_ABST
    Figure CN122490445A_ABST
Patent Text Reader

Abstract

This application provides a method for detecting defects in rubber stoppers based on image processing, relating to the fields of artificial intelligence and machine vision. It solves the technical problems of existing technologies, such as difficulty in simultaneously detecting appearance defects and airtightness defects, low detection efficiency, and poor accuracy. The method includes: acquiring multi-view RGB images, 3D point cloud data, and pressure time-series parameters of the rubber stopper under multiple preset pressure states; preprocessing the images and point clouds; extracting image features and point cloud features using a dual-stream coding network; encoding the pressure time-series parameters into pressure feature vectors; performing hierarchical fusion of image features, point cloud features, and pressure feature vectors based on a cross-modal attention fusion mechanism to obtain multi-level cross-modal fusion features; synchronously outputting pixel-level segmentation masks for appearance defects and 3D heatmaps for airtightness defects through a dual-task decoding network; and locating the defect position based on the segmentation mask and heatmap.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence and machine vision technology, specifically a method for detecting defects in rubber stopper production based on image processing. Background Technology

[0002] In the production of rubber stoppers, defect detection is mainly divided into two tasks: appearance defect detection and airtightness defect detection. Current technologies typically treat these two as independent processes: appearance defect detection often employs manual visual inspection or automated equipment based on machine vision. High-resolution cameras capture images of the rubber stopper surface, and image processing algorithms identify surface defects such as scratches, burrs, missing material, and bubbles. This method enables non-contact online detection, but it only acquires two-dimensional texture information and cannot perceive the internal structural integrity and sealing performance of the rubber stopper. Airtightness defect detection commonly uses methods such as colorimetric methods, bubble methods, vacuum attenuation methods, or helium mass spectrometry leak detection. These methods require assembling the rubber stopper onto containers such as vials or pre-filled syringes, and then observing the penetration of dye, bubble generation, or pressure changes after pressurization or vacuuming to determine the seal. Not only is the detection process destructive or contact-based, but it also typically uses offline sampling inspection, resulting in low efficiency and long cycles, failing to meet the 100% inspection requirements of high-speed production lines. More importantly, in existing technologies, appearance inspection and airtightness inspection operate independently, and there is a lack of data correlation between the inspection equipment. This makes it impossible to simultaneously obtain the external appearance quality and internal sealing performance of the rubber stopper at the same inspection station, and it is even more difficult to jointly locate and comprehensively analyze the location information of the two types of defects.

[0003] Therefore, existing testing methods suffer from technical problems such as fragmented testing functions, destructive sampling for airtightness testing, inability to achieve online non-destructive full inspection, and incomplete defect location information, which restrict the intelligent development of rubber stopper production quality control. Summary of the Invention

[0004] This application provides a method for detecting defects in rubber stopper production based on image processing, which solves the technical problems in the prior art where appearance defect detection and airtightness defect detection are separated, different equipment is required for step-by-step detection, resulting in low detection efficiency. Furthermore, airtightness detection is mostly destructive sampling inspection, which cannot achieve non-destructive online full inspection, and it is difficult to simultaneously locate the two-dimensional image position of appearance defects and the three-dimensional spatial position of airtightness defects.

[0005] To achieve the above objectives, this application adopts the following technical solution:

[0006] Firstly, an image processing-based method for detecting defects in rubber stopper production is provided, including:

[0007] Image data, three-dimensional point cloud data, and corresponding pressure time parameters of the rubber stopper are acquired under multiple preset pressure states, and the multi-view image data and three-dimensional point cloud data are preprocessed; wherein, the three-dimensional point cloud data includes the three-dimensional spatial coordinates of each point on the surface of the rubber stopper and the deformation relative to the pressureless state;

[0008] A two-stream coding network is used to extract multi-scale image features from preprocessed image data and point cloud features from preprocessed point cloud data, respectively.

[0009] The pressure time-series parameters are encoded into pressure feature vectors, and the image features, point cloud features, and pressure feature vectors are fused based on a cross-modal attention fusion mechanism to obtain cross-modal fused features.

[0010] Using a dual-task decoding network, appearance defect detection and airtightness defect detection tasks are performed based on the cross-modal fusion features. The segmentation mask of appearance defects and the three-dimensional heat map of airtightness defects are output, and the defect location is located based on the segmentation mask and the three-dimensional heat map.

[0011] Based on the above technical solution, the image processing-based method for detecting defects in rubber stoppers provided in this application acquires image data, 3D point cloud data, and pressure time-series parameters of the rubber stopper under multiple preset pressure states. It then performs multimodal fusion of appearance texture information, 3D geometric deformation information, and dynamic pressure conditions to achieve simultaneous detection of external defects and internal airtightness defects. This method utilizes a dual-stream coding network to extract image features and point cloud features separately, and combines them with pressure feature vectors for cross-modal attention fusion. This allows the network to adaptively adjust the contribution weights of image features and geometric features according to different pressure states, thereby improving the accuracy and robustness of defect detection. Simultaneously, a dual-task decoding network outputs pixel-level segmentation masks for appearance defects and 3D heatmaps for airtightness defects. This not only enables defect type discrimination but also accurately locates the 2D image position and 3D spatial position of the defects, overcoming the problems of separation and low efficiency in traditional detection methods. Furthermore, this method can complete the detection without damaging the rubber stopper, possessing advantages such as non-contact operation, high efficiency, and online deployment, making it suitable for large-scale industrial production scenarios.

[0012] Furthermore, the process of acquiring the image data, 3D point cloud data, and corresponding pressure time-series parameters includes:

[0013] Multi-view image data of the rubber stopper is obtained by simultaneously capturing multi-view images of the stopper from different angles under various preset pressure conditions using multiple industrial cameras.

[0014] A coded grating pattern is projected onto the surface of the rubber stopper using structured light projection. A binocular stereo vision system is used to collect three-dimensional point cloud data of the rubber stopper surface under no pressure and at multiple incremental characteristic pressure threshold points. The pressure value at each characteristic pressure threshold point is recorded by a pressure sensor to form pressure time series parameters. The multiple incremental characteristic pressure threshold points are the preset pressure states.

[0015] For each characteristic pressure threshold point, the point cloud data collected is obtained by subtracting the point cloud coordinates under no-pressure conditions from the point cloud coordinates under pressurized conditions, and the deformation of each point is calculated. The deformation is then added to the feature vector of the point as an attribute value of the three-dimensional point cloud data.

[0016] Furthermore, the preprocessing includes:

[0017] The distortion correction of image data is performed using Zhang Zhengyou calibration method, and the corrected image is then subjected to bilateral filtering for noise reduction, histogram equalization, and background removal based on threshold segmentation and morphological operations to obtain preprocessed image data.

[0018] The three-dimensional point cloud data is subjected to voxel downsampling and discrete points are removed by statistical filtering to obtain preprocessed three-dimensional point cloud data. The feature vector of each point in the preprocessed three-dimensional point cloud data is represented as [x,y,z,Δz], where (x,y,z) are the spatial coordinates after coordinate centering and normalization, and Δz is the deformation of each point.

[0019] Furthermore, the dual-stream coding network includes an image coding stream and a point cloud coding stream, wherein,

[0020] The image encoding stream uses a CNN-Transformer hybrid backbone network. It uses a convolutional neural network (CNN) to extract the first feature of the image data, and uses a Transformer module to perform global self-attention processing on the first feature to obtain the second feature. Then, it uses a dilated convolutional spatial pyramid pooling module to perform multi-scale context aggregation on the second feature to obtain the third feature. Finally, it uses global average pooling and convolutional dimensionality reduction layers to process the third feature and output multi-scale image features.

[0021] The point cloud encoding stream uses an improved PointNet++ network. The improved PointNet++ network improves upon the PointNet++ network by introducing sub-manifold sparse convolution to enhance local feature aggregation and adding a weighted position encoding module to encode the coordinates and deformation of points. Furthermore, through a hierarchical feature aggregation module, the multi-scale point cloud features output from each layer are cascaded through skip connections to output point cloud features.

[0022] Furthermore, the improved PointNet++ network includes a first set abstraction layer, a second set abstraction layer, and a third set abstraction layer;

[0023] The first set abstraction layer downsamples the input point cloud from N points to M first key points through farthest point sampling. Then, within the first neighborhood radius of each first key point, local feature aggregation is performed using a sub-manifold sparse convolution kernel. Before aggregation, the coordinates and deformation of the points are sinusoidally encoded and concatenated into the point features through a weighted position encoding module, outputting the first multidimensional feature vector. Here, M is taken as one-quarter of N, and the first neighborhood radius is set to 0.2.

[0024] The second set abstraction layer downsamples the M key points into P second key points. Then, within the second neighborhood radius of each second key point, local feature aggregation is performed using a submanifold sparse convolution kernel. Weighted position encoding is performed before aggregation to output the second multidimensional feature vector. Here, P is one-quarter of M, and the second neighborhood radius is 0.4.

[0025] The third set abstraction layer downsamples P second key points into Q third key points. Then, within the third neighborhood radius of each third key point, local feature aggregation is performed using a submanifold sparse convolution kernel. Weighted position encoding is performed before aggregation to output a third multidimensional feature vector. Here, Q is one-quarter of P, and the third neighborhood radius is 0.8.

[0026] Finally, the hierarchical aggregation feature module concatenates the first multidimensional feature vector, the second multidimensional feature vector, and the third multidimensional feature vector through skip connections to output point cloud features.

[0027] Wherein, N>M>P>Q, and the dimension of the first multidimensional feature vector < the dimension of the second multidimensional feature vector < the dimension of the third multidimensional feature vector.

[0028] Further, encoding the pressure time-series parameters into a pressure feature vector includes:

[0029] The stress time-series parameters are input into a hybrid coding network. The first layer of the hybrid coding is a first multilayer perceptron with multiple neurons, which maps each stress value to a multidimensional vector to obtain an intermediate representation. The second layer is a long short-term memory network layer, which outputs the hidden states of all time steps based on the intermediate representation, and takes the hidden state of the last time step as the temporal dependency feature. The third layer is a second multilayer perceptron with multiple neurons, which maps the temporal dependency feature to a multidimensional stress feature vector.

[0030] Furthermore, the cross-modal attention fusion mechanism specifically includes a hierarchical fusion step, which performs cross-modal attention fusion on the features of each level output from the image encoding stream and the corresponding level features output from the point cloud encoding stream, respectively, to obtain multi-level cross-modal fused features; wherein:

[0031] The first feature in the image encoding stream and the first multidimensional feature vector in the point cloud encoding stream are processed by the first-level cross-modal attention fusion module to obtain the first-level cross-modal fusion feature;

[0032] The second feature in the image encoding stream and the second multidimensional feature vector in the point cloud encoding stream are processed by the second-level cross-modal attention fusion module to obtain the second-level cross-modal fusion feature;

[0033] The third feature in the image encoding stream and the third multidimensional feature vector in the point cloud encoding stream are processed by the third-level cross-modal attention fusion module to obtain the third-level cross-modal fusion feature;

[0034] The multi-scale image features and the point cloud features are processed by the fourth-level cross-modal attention fusion module to obtain the fourth-level cross-modal fusion features.

[0035] Furthermore, the specific workflow of the cross-modal attention fusion module is as follows:

[0036] Using pre-calibrated camera intrinsic parameter matrices, rotation matrices, and translation vectors, a bidirectional projection mapping between 3D point clouds and image pixels is established.

[0037] Based on the bidirectional projection mapping, the image feature map and the depth map are back-projected to generate point cloud features derived from the image. Each point in the point cloud features derived from the image has three-dimensional spatial coordinates and carries corresponding image semantic features. The depth map is a two-dimensional matrix composed of the depth values ​​of each pixel calculated from the acquired raster image through structured light projection and binocular stereo matching algorithm. The image feature map is a two-dimensional feature map output by the image encoding stream at the current level.

[0038] Based on the bidirectional projection mapping, the point cloud feature map is forward-projected onto the image plane to generate image features derived from the point cloud. The image features derived from the point cloud are two-dimensional feature maps, and each non-empty pixel position carries the geometric and deformation features of the corresponding point cloud. The point cloud feature map is a set of point cloud feature vectors output by the point cloud encoding stream at the current level.

[0039] First, cross-attention from image to point cloud is performed, including: using the point cloud feature map as the query vector, using the point cloud features exported from the image as the key vector and value vector after linear transformation, using the scaling dot product attention formula to calculate the attention weight between each point cloud point and each point in the point cloud exported from the image, and then weighting and summing the attention weights with the point cloud features exported from the image to obtain the point cloud features after image enhancement.

[0040] Then, cross-attention from point cloud to image is performed, including: using the image feature map as the query vector, using the image features exported from the point cloud as the key vector and value vector after linear transformation, using the scaling dot product attention formula to calculate the attention weight of each image pixel and each pixel in the image features exported from the point cloud, and then weighting and summing the attention weights with the image features exported from the point cloud to obtain the image features after point cloud enhancement.

[0041] The pressure feature vector, along with the image-enhanced point cloud features and the point cloud-enhanced features, are input into a gating network to output cross-modal fusion features.

[0042] Furthermore, the appearance defect detection task is performed by an appearance defect decoder, which is built based on the U-Net architecture. The specific workflow of the appearance defect decoder includes:

[0043] The input to the appearance defect decoder is the image features enhanced by point clouds at each level, which are organized into the encoder-side feature pyramid of U-Net according to resolution from high to low.

[0044] The cross-modal fusion features and stress feature vectors at each level are injected as conditional modulation signals into each upsampling stage of the U-Net decoder: In each upsampling stage, the cross-modal fusion features are adaptively pooled to adjust to the spatial size of the current level feature map, and then channel-concatenated with the current level feature map. The stress feature vector is passed through a multilayer perceptron to generate affine transformation parameters. The affine transformation parameters are then used to perform FiLM modulation with the channel-concatenated features to obtain the modulation features.

[0045] The U-Net decoder employs a symmetrical upsampling path and receives point cloud-enhanced image features from the corresponding level of the encoder-side feature pyramid via skip connections.

[0046] A Triplet Attention module is embedded in each skip connection to generate three-dimensional attention weights along the channel dimension, height dimension, and width dimension respectively, and adaptively weights the feature maps passed by the skip connections.

[0047] The upsampling path ultimately restores the feature map to the same spatial resolution as the input image, and the number of output channels is equal to the number of defect categories, which include no defects, scratches, missing material, burrs, and bubbles.

[0048] Finally, a pixel-level segmentation probability map is output through 1×1 convolution and the Softmax activation function, which serves as a segmentation mask for appearance defects.

[0049] Furthermore, the airtightness defect detection task is performed by an airtightness defect decoder, which is built based on a differential point Transformer architecture. The specific workflow includes:

[0050] The input to the airtightness defect decoder is the point cloud features enhanced by the image at each level, namely the first-level point cloud features, the second-level point cloud features, the third-level point cloud features, and the fourth-level point cloud features.

[0051] The point cloud features of each layer are input into a corresponding parallel branch. Each branch independently performs the following operations: First, the cross-modal fusion features of the same layer are back-projected onto each point cloud point using camera parameters to obtain the cross-modal feature vector of the point cloud point. Then, the pressure feature vector is used to generate the affine modulation parameters of the current branch through a multilayer perceptron. The affine modulation parameters are used to perform FiLM modulation on the point cloud features of the current layer. Finally, the modulated features are concatenated with the cross-modal feature vector to obtain the enhanced point cloud features. Then, differential set abstraction is performed, including farthest point sampling on the enhanced point cloud features. To reduce the number of points, the deformation difference between the neighboring points and the keypoint is calculated in the ball query neighborhood of each sampled keypoint. The deformation difference is then concatenated with the spatial coordinate offset of the neighboring points and the enhanced features, and input into the submanifold sparse convolution and shared multilayer perceptron. The output is the local aggregated features of the keypoint. The features of all keypoints constitute the downsampled feature map. Finally, feature propagation is performed, including restoring the number of points of the original input of the branch to the downsampled feature map through inverse distance weighted interpolation, and concatenating the interpolation result with the original enhanced features and then fusing them through the multilayer perceptron to output the propagated features of the branch.

[0052] The four branches output different numbers of propagation feature points, which are denoted as the first propagation feature, the second propagation feature, the third propagation feature, and the fourth propagation feature, respectively. The number of points in the first propagation feature is equal to the number of points in the original point cloud. The second, third, and fourth propagation features are upsampled to the number of points in the original point cloud through inverse distance weighted interpolation to obtain upsampled features with the same number of points as the first propagation feature. The first propagation feature and the three upsampled features are concatenated along the channel dimension to obtain the multi-scale fused feature.

[0053] Multi-scale fused features are input into a multilayer perceptron, which outputs the dimensionality-reduced features of each point cloud point. The dimensionality-reduced features of each point are then input into a regression head, which outputs the probability that each point cloud point belongs to an airtightness defect. The probabilities of all points constitute a three-dimensional heatmap. The regression head consists of two fully connected layers and a sigmoid activation function.

[0054] Finally, the post-processing module performs threshold filtering on the 3D heat map, marking points with probability values ​​greater than a preset threshold as candidate defect points. The DBSCAN density clustering algorithm is used to cluster the candidate defect points to obtain several point clusters. The minimum axis-aligned bounding box of each point cluster is calculated, and the vertex coordinates of each bounding box are output as the 3D spatial location of the airtightness defect.

[0055] Furthermore, the step of locating the defect position based on the segmentation mask and the three-dimensional heatmap includes:

[0056] For appearance defects, if there is a continuous pixel region corresponding to any defect category in the pixel-level segmentation mask and the area of ​​the region is greater than the preset minimum area threshold, then the appearance is determined to be unqualified, and the contour point coordinate sequence of the continuous pixel region is used as the two-dimensional image position of the appearance defect.

[0057] For airtightness defects, if there is a cluster of points in the three-dimensional thermal map with a probability value greater than the preset probability threshold and a number of consecutive points greater than the preset value, then the airtightness is determined to be unqualified, and the three-dimensional axis of the cluster is aligned with the bounding box coordinates as the three-dimensional spatial location of the airtightness defect.

[0058] During the comprehensive judgment, only when both appearance and airtightness are judged to be qualified will the final qualified judgment and comprehensive confidence score be output; otherwise, the unqualified judgment will be output, and the two-dimensional image position mask of all detected appearance defects and the three-dimensional spatial bounding box coordinates of all detected airtightness defects will be output at the same time.

[0059] In addition, by using a pre-calibrated camera and point cloud coordinate system transformation matrix, the three-dimensional bounding box coordinates of the airtightness defect are projected inversely to the two-dimensional image coordinate system. The appearance defect mask and the airtightness defect projection area are superimposed on the same RGB image to form a unified defect visualization result.

[0060] Compared with the prior art, the beneficial effects of this application are:

[0061] First, in the data acquisition and preprocessing stage, this application uses multiple industrial cameras to synchronously acquire multi-view images in a surround-like manner. Combined with structured light projection and a binocular stereo vision system, it acquires 3D point clouds and deformation under multiple increasing pressure states, while simultaneously recording pressure time-series parameters. This multi-modal, multi-pressure state acquisition method not only preserves the texture details of the rubber stopper surface but also dynamically captures the 3D deformation response during pressurization, providing rich spatiotemporal information for subsequent defect detection. Distortion correction, bilateral filtering, histogram equalization, and background removal are applied to the images. Voxel downsampling and statistical filtering are performed on the point clouds, effectively eliminating noise and outliers and improving data quality. In particular, explicitly encoding the deformation as the fourth dimension attribute of the point cloud allows the network to directly perceive geometric changes under pressure, laying a data foundation for the quantitative analysis of airtightness defects.

[0062] Secondly, regarding feature encoding and fusion, this application designs a dual-stream encoding network: the image encoding stream employs a CNN-Transformer hybrid backbone network, which can capture local texture details and aggregate multi-scale contextual information through global self-attention and dilated convolutional pyramids; the point cloud encoding stream makes several improvements to PointNet++, introducing sub-manifold sparse convolution to enhance local geometric feature extraction, adding weighted positional encoding to explicitly fuse coordinates and deformation, and using a hierarchical feature aggregation module to skip-connect point cloud features at different scales, effectively preserving geometric information from fine-grained to coarse-grained. Furthermore, this application proposes mapping stress temporal parameters to stress feature vectors through a hybrid encoding network and employing a hierarchical cross-modal attention fusion mechanism, performing bidirectional cross-attention interaction between multiple corresponding layers of the image encoding stream and the point cloud encoding stream, while using the stress feature vector as a condition to guide gated adaptive fusion. This deep, multi-scale fusion strategy enables information from different modalities to enhance each other: image features supplement the point cloud with texture semantics, point cloud features supplement the image with geometric deformation information, and pressure conditions allow the fusion weights to be dynamically adjusted as the pressure is applied, thereby improving the joint perception capability of minor external defects and internal sealing defects.

[0063] Finally, in terms of defect decoding and localization output, this application adopts a dual-task decoding network: the appearance defect decoder is based on the U-Net architecture, utilizing the image branch in the multi-level fusion features and enhancing the sensitivity to minute defects through the TripletAttention module in the skip connections, outputting a pixel-level segmentation mask that can accurately delineate the contours of defects such as scratches, missing material, burrs, and bubbles; the airtightness defect decoder is based on the differential point Transformer architecture, amplifying abnormal deformation signals by calculating the difference in local deformation, regressing the probability of airtightness defects point by point and generating a three-dimensional heat map, and then obtaining the three-dimensional bounding box of the defect through density clustering. This design not only simultaneously completes the detection of both appearance and airtightness defects, but also provides the two-dimensional image position and three-dimensional spatial position respectively. Finally, the three-dimensional defect position is back-projected onto the two-dimensional image through coordinate transformation to achieve a unified visualization. The overall method does not require destruction of the rubber stopper, has high detection efficiency and accurate localization, and can be deployed online, providing a reliable technical solution for fully automated and full-function inspection of rubber stopper production quality. Attached Figure Description

[0064] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0065] Figure 1 A system architecture diagram of an image processing-based rubber stopper manufacturing defect detection system provided in this application embodiment;

[0066] Figure 2 A schematic flowchart of an image processing-based method for detecting defects in rubber stopper production is provided in an embodiment of this application.

[0067] Figure 3 A schematic flowchart of another image processing-based method for detecting defects in rubber stopper production provided in this application embodiment;

[0068] Figure 4 This is a schematic flowchart of another image processing-based method for detecting defects in rubber stopper production, provided in an embodiment of this application. Detailed Implementation

[0069] In the description of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B. The "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Furthermore, "at least one" means one or more, and "multiple" means two or more. The terms "first," "second," etc., do not limit the quantity or order of execution, and "first," "second," etc., do not necessarily imply differences.

[0070] It should be noted that, in this application, the terms "exemplary" or "for example" are used to indicate that something is being described as an example, illustration, or illustration. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0071] To address the technical problems in existing technologies, such as the separation of appearance defect detection and airtightness defect detection, the fact that airtightness detection is mostly destructive sampling and cannot be fully inspected online, and the incomplete defect location information, this application provides an image processing-based method for detecting defects in rubber stopper production. The method includes:

[0072] Image data, 3D point cloud data, and corresponding pressure time parameters of the rubber stopper are acquired under multiple preset pressure conditions, and the multi-view image data and 3D point cloud data are preprocessed; wherein, the 3D point cloud data includes the 3D spatial coordinates of each point on the surface of the rubber stopper and the deformation relative to the pressureless state;

[0073] A two-stream coding network is used to extract multi-scale image features from preprocessed image data and point cloud features from preprocessed point cloud data, respectively.

[0074] The stress time series parameters are encoded into stress feature vectors, and image features, point cloud features, and stress feature vectors are fused based on a cross-modal attention fusion mechanism to obtain cross-modal fused features;

[0075] Using a dual-task decoding network, appearance defect detection and airtightness defect detection tasks are performed based on cross-modal fusion features. The output is a segmentation mask for appearance defects and a 3D heat map for airtightness defects. The defect location is then located based on the segmentation mask and the 3D heat map.

[0076] Based on this, this application achieves simultaneous and accurate detection of rubber stopper appearance and airtightness defects by simulating the pressure conditions of the rubber stopper and integrating multimodal data, which significantly improves the comprehensiveness and reliability of the detection.

[0077] like Figure 1 As shown in the figure, an image processing-based method for detecting defects in rubber stoppers provided in this application includes the following steps:

[0078] S1. Acquire image data, 3D point cloud data and corresponding pressure timing parameters of the rubber stopper under multiple preset pressure states, and preprocess the multi-view image data and 3D point cloud data.

[0079] The 3D point cloud data includes the 3D spatial coordinates of each point on the rubber stopper surface and its deformation relative to the unpressurized state. Pressure time-series parameters characterize the pressure changes experienced by the rubber stopper during different pressurization stages. These parameters serve as conditional information for subsequent fusion processes, enabling the network to dynamically adapt to the differences in defect manifestations under different pressure states. By collecting data under multiple preset pressure states, the geometric deformation evolution of the rubber stopper during pressurization can be captured, thus providing key dynamic deformation characteristics for the identification of airtightness defects.

[0080] In some implementations, the specific methods of data acquisition can be diverse. For example, image data can be simultaneously acquired by multiple industrial cameras arranged around the rubber stopper station to achieve multi-view coverage; 3D point cloud data can be acquired through non-contact scanning using structured light sensors or laser profilometers; and pressure timing parameters are recorded in real time by high-precision pressure sensors installed on the pressurization device. Preset pressure states can be set with multiple gradients according to the actual usage scenario of the rubber stopper, such as pre-pressurization, half-pressurization, and full-pressurization states.

[0081] It is important to note that preprocessing steps are crucial for the quality of subsequent feature extraction. Preprocessing includes at least: denoising, geometric correction, brightness normalization, and viewpoint alignment for multi-view image data; filtering and denoising, downsampling to balance computational efficiency and accuracy, and hole repair for 3D point cloud data. In particular, the preprocessing of image data and point cloud data must maintain the consistency of the spatial coordinate system, which is usually achieved through calibration to accurately map the image pixel coordinates to the point cloud spatial coordinates.

[0082] S2. Using a dual-stream coding network, multi-scale image features are extracted from the preprocessed image data, and point cloud features are extracted from the preprocessed point cloud data.

[0083] Two-stream coding networks (DBCs) are deep learning networks that consist of two independent but potentially shared shallow weights in parallel coding structures. Their core function is to encode image data and point cloud data (two-dimensional grids and three-dimensional unordered point sets) using the most suitable feature extractors, thereby preserving the original information and unique representations of each modality to the greatest extent possible. DBCs typically take the following forms: image streams are often based on convolutional neural network architectures, such as ResNet or EfficientNet; point cloud streams are based on point-set-oriented deep learning architectures, such as PointNet++ or Point Transformer.

[0084] In some implementations, image streams can employ classic 2D convolutional neural networks such as the Visual Geometric Group Network (VGGNet), Residual Network (ResNet), Densely Connected Convolutional Network (DenseNet), or EfficientNet. Lightweight convolutional neural networks such as MobileNet or ShuffleNet can also be used, or hybrid architectures combining convolution and self-attention mechanisms can be employed. The image features output by the image stream can be feature maps of one or more scales, for example, multi-scale features can be output through feature layers with different downsampling factors.

[0085] Point cloud streams can employ the PointNet family of point networks, where PointNet uses a point-wise multilayer perceptron and global pooling, and PointNet++ further introduces hierarchical sampling and local neighborhood aggregation. Point cloud streams can also utilize sparse convolution-based 3D point cloud networks such as the Submanifold Sparse Convolutional Network or Transformer-based point cloud networks such as PointTransformer. The point cloud features output by a point cloud stream can be point-wise feature vectors or sets of downsampled keypoint features.

[0086] It should be noted that two-stream networks are not completely independent and can be designed to interact at specific layers. For example, cross-stream knowledge distillation or feature alignment loss can be introduced into the middle layer of the encoder to enable image features to learn geometric information consistent with point cloud features, thereby improving the expressive power of point cloud features in sparse texture regions.

[0087] S3. Encode the stress time series parameters into stress feature vectors, and fuse the image features, point cloud features and stress feature vectors based on the cross-modal attention fusion mechanism to obtain cross-modal fused features.

[0088] The cross-modal attention fusion mechanism is a computational module that draws inspiration from the principles of human visual attention. Instead of simply concatenating features from different modalities, it learns a weighting matrix to dynamically calculate the importance of a particular modality (e.g., pressure features) to different regions or channels in another modality (e.g., image features), guiding the weighted integration of information accordingly. Simultaneously, introducing pressure feature vectors as conditional information allows the fusion process to dynamically depend on the current pressure state, as deformation features under different pressures have varying indicative effects on airtightness defects. The cross-modal attention fusion mechanism allows the texture, edge, and color information in image features to complement the geometric shape, spatial location, and deformation information in point cloud features, forming a more comprehensive defect representation than a single modality.

[0089] In some implementations, the fusion process can be as follows: First, the stress temporal parameters are converted into a fixed-length stress feature vector through a temporal coding network (such as a Long Short-Term Memory network or a Gated Recurrent Unit network). Then, this stress feature vector is used as a query vector and cross-attention is performed with image features and point cloud features respectively to obtain stress-guided image enhancement features and stress-guided point cloud enhancement features. Finally, these enhancement features, along with the original features, are input into a multi-head self-attention module to complete the final deep fusion and generate cross-modal fused features.

[0090] It is worth noting that the introduction of pressure time series parameters is a key innovation. It enables the model to "understand" the evolution of defects under different pressures. For example, when the pressure of an airtight defect reaches a critical value, the corresponding point cloud deformation region will undergo a sudden change, and the pressure feature vector will guide the model to focus on this deformation region at this moment.

[0091] S4. Using a dual-task decoding network, perform appearance defect detection and airtightness defect detection tasks based on cross-modal fusion features. Output the segmentation mask of appearance defects and the three-dimensional heat map of airtightness defects, and locate the defect location based on the segmentation mask and the three-dimensional heat map.

[0092] The dual-task decoding network is a network structure capable of simultaneously handling two related but distinct prediction tasks. The two tasks share the same fused feature representation but each has its own independent decoding branch. The appearance defect detection task aims to identify visible flaws on the rubber stopper surface, such as scratches, missing material, burrs, and bubbles, and outputs the probability of each pixel belonging to each type of defect, i.e., a segmentation mask. The airtightness defect detection task aims to identify whether the rubber stopper has a risk of leakage or sealing failure under pressure, and outputs the probability of each point cloud point belonging to an airtightness defect, i.e., a 3D heatmap. This method achieves good detection results because: firstly, the cross-modal fusion features simultaneously include appearance texture information and geometric deformation information, providing complementary inputs for the two tasks; secondly, the two tasks are jointly optimized in the same network, and the identification results of appearance defects can indirectly assist in the discrimination of airtightness defects. For example, cracks often lead to both appearance and airtightness problems, and vice versa, thereby improving the overall detection performance.

[0093] In some implementations, the appearance defect decoder can employ a semantic segmentation network based on an encoder-decoder structure, such as a fully convolutional network (FCN), a U-shaped network (U-Net), a deep Lab series network, or a pyramid scene parsing network (PSPNet). The output resolution of the segmentation mask can be the same as or different from the input image; if different, it can be restored to the original resolution through upsampling. Defect categories can be defined according to actual production needs and can include no defects as well as multiple defect types.

[0094] The airtightness defect decoder can employ point cloud-based semantic segmentation networks or anomaly detection networks, such as a segmentation version of PointNet++, a point attention network, or a graph convolution-based point cloud segmentation network. The 3D heatmap output consists of the same number of probability values ​​as the input point cloud, with each value representing the confidence that the corresponding point belongs to an airtightness defect region. To pinpoint the specific 3D location of the defect, the 3D heatmap can be post-processed. For example, threshold segmentation can be used to obtain a set of points with probabilities higher than a certain value, and then connected component analysis or clustering algorithms can be used to extract each independent defect region, calculating the minimum 3D bounding box for each region.

[0095] It should be noted that the two tasks can be further coordinated. For example, a task-coordinated loss function can be designed such that if the appearance defect segmentation mask shows damage in a certain area, then the predicted value of that area on the airtightness heatmap should also be penalized to be higher, because physical damage inevitably leads to a decrease in airtightness. Conversely, for stain defects with only color abnormalities and no structural damage, the airtightness heatmap should not be penalized. This relationship can be achieved through a coordinated regularization term. Furthermore, the two decoding branches can share some parameters or remain completely independent. To enhance the coordination between the two tasks, information exchange connections can be added between the decoding branches, such as projecting the intermediate feature map of appearance segmentation onto the point cloud space and fusing it with the airtightness defect features.

[0096] Based on the above technical solutions, this application provides an image processing-based method for detecting defects in rubber stoppers. By simultaneously acquiring images, point clouds, and pressure time-series data under multiple pressure conditions, and designing a dedicated network structure with dual-stream coding, cross-modal fusion, and dual-task decoding, it achieves integrated, high-precision detection of both external and internal airtightness defects in rubber stoppers. This method fully utilizes the complementarity and coupling between multimodal data, overcomes the limitations of traditional single static visual inspection, significantly improves the defect detection rate, reduces the false detection rate, and provides strong technical support for quality control in rubber stopper production.

[0097] In one possible implementation of the embodiments of this application, as shown, the above-mentioned S1 can be specifically implemented by the following S101, S102 and S103, which are described in detail below:

[0098] S101. Using multiple industrial cameras, multi-view images of the rubber stopper are simultaneously captured in a surround manner under various preset pressure conditions to obtain multi-view image data.

[0099] To comprehensively capture the texture details of the rubber stopper surface, multiple industrial cameras are needed to simultaneously capture images from different angles. The number of cameras is typically four to six, arranged at equal angles around the rubber stopper to ensure that each area of ​​the stopper surface is covered by at least two cameras, thus eliminating potential occlusion or glare issues from a single viewpoint. Preset pressure states include a no-pressure state and at least two incremental pressure thresholds, such as normal pressure, low pressure, medium pressure, and high pressure. The cameras synchronously trigger image acquisition at each pressure state to ensure that the multi-view images acquired at the same time strictly correspond to the current pressure state.

[0100] In some implementations, the camera employs a charge-coupled device (CCD) or complementary metal-oxide-semiconductor (CMOS) sensor with a resolution of at least 5 megapixels, and is used in conjunction with a multi-angle ring light source for strobe illumination. The brightness and color temperature of the ring light source can be adaptively adjusted according to the material of the rubber stopper, such as halogenated butyl rubber or silicone rubber, to eliminate surface reflections and enhance defect contrast. The camera and light source are controlled by a programmable logic controller (PLC) that sends a synchronous trigger signal to ensure that image acquisition is completed instantly after the pressure reaches a preset threshold and stabilizes. The acquired multi-view image data is stored according to the pressure state and viewpoint number, for example, P_i_viewpoint j, where i represents the i-th pressure state and j represents the j-th camera viewpoint.

[0101] S102. A coded grating pattern is projected onto the surface of the rubber stopper using structured light projection. A binocular stereo vision system is used to collect three-dimensional point cloud data of the rubber stopper surface at no pressure and at multiple incremental characteristic pressure threshold points. The pressure value of each characteristic pressure threshold point is recorded by a pressure sensor to form pressure time series parameters.

[0102] Among them, the pressure time series parameter records the pressure value sequence from no pressure to each pressure threshold point. This sequence not only contains the pressure value, but also the order information of pressure loading, providing a time dependency for subsequent encoding of pressure feature vectors.

[0103] In some implementations, the 3D point cloud data and pressure time-series parameters are obtained as follows:

[0104] First, the rubber stopper is fixed to a special carrier inside the transparent pressure chamber. The chamber is connected to a precision pressure controller, which can gradually increase the internal air pressure according to a preset program. The pressure controller integrates a piezoresistive pressure sensor with a range of 0~200 kPa and a sampling frequency of 100 Hz to record the pressure value inside the chamber in real time.

[0105] Next, the structured light projection binocular stereo vision system is activated. This system consists of a digital light processing (DLP) projector and two industrial cameras (Basler acA2440-35uc, 5-megapixel resolution). Both the projector and cameras are rigidly fixed outside the cavity, and the light path illuminates the surface of the rubber stopper through the transparent cavity wall. Then, a four-step phase-shifting method combined with Gray code encoding is used to generate a grating pattern.

[0106] The data acquisition process proceeded sequentially according to the following pressure threshold points: no pressure (0 kPa), first pressurized state (50 kPa), second pressurized state (100 kPa), and third pressurized state (150 kPa). The pressure controller started at 0 kPa and increased the pressure at a constant rate of 10 kPa / s. Once the pressure sensor detected that the pressure value stabilized within the target threshold ±0.5 kPa range for 100 ms, it triggered the camera to synchronously acquire the deformed grating image of the current state. Two cameras simultaneously captured the grating pattern projected onto the rubber stopper surface, acquiring a total of 9 pairs of images (9 images per camera) consisting of four phase-shift steps and five Gray code frames. The pressure sensor recorded the pressure value at the stable state, forming a pressure time series array s=[0,50.2,99.8,150.1] (unit: kPa), which was stored as a label corresponding to the point cloud data.

[0107] After image acquisition, the 3D point cloud is reconstructed through the following steps: First, epipolar correction is performed on the raster images acquired by each pair of cameras to eliminate lens distortion; then, the wrapping phase is calculated using the four-step phase-shifting method, and the absolute phase is obtained by unwrapping using Gray code; phase-based stereo matching is performed between the left and right camera images to obtain the disparity value of each pixel; finally, based on the principle of binocular triangulation, the 3D spatial coordinates (x, y, z) corresponding to each pixel are recovered from the disparity, generating a dense point cloud. The original point cloud obtained under each pressure state contains approximately 100,000 to 150,000 points, stored in binary PLY format, with each point recording the x, y, and z coordinates in millimeters, rounded to three decimal places.

[0108] To ensure accurate calculation of the deformation variable Δz from point clouds under different pressure conditions, after acquiring the point cloud under no-pressure conditions, the position of the rubber stopper and all camera and projector parameters were kept absolutely unchanged, and point clouds under each pressure condition were acquired sequentially. After acquisition, each pressurized point cloud was rigidly registered with the unpressurized point cloud, and the rotation matrix and translation vector were calculated using the Iterative Closest Point (ICP) algorithm to align the two point clouds in space. After registration, for each point in the pressurized point cloud, the nearest neighbor point was searched in the unpressurized point cloud, and the difference between their z-coordinates was calculated as the deformation variable Δz for that point. Δz was then written as an additional attribute into the point cloud file. Finally, the feature vector of each point cloud point is [x, y, z, Δz].

[0109] S103. Preprocess the multi-view image data and the 3D point cloud data respectively to obtain preprocessed image data and preprocessed 3D point cloud data. The feature vector of each point in the preprocessed 3D point cloud data is represented as [x,y,z,Δz]. The deformation Δz is calculated by subtracting the point cloud coordinates under pressure from the point cloud coordinates under no-pressure conditions.

[0110] Image preprocessing mainly includes distortion correction, noise reduction filtering, contrast enhancement, and background segmentation; point cloud preprocessing mainly includes downsampling, outlier filtering, coordinate normalization, and deformation calculation. The deformation variable Δz is the core physical quantity connecting pressure state and geometric change, and its calculation formula is as follows:

[0111] Let the coordinates of a point in the point cloud under no-pressure conditions be... Under a certain pressurized state (pressure value) Under this condition, the corresponding point in the same spatial location is: Because the rubber stopper may undergo slight horizontal displacement under pressure, the coordinates of the same index point cannot be simply subtracted. Instead, nearest neighbor matching or iterative nearest point ICP algorithms are needed for point cloud registration, followed by calculating the vertical deformation of each point. The deformation is defined as follows:

[0112] Where i represents the i-th point in the pressurized point cloud, and i' represents the corresponding point found in the reference point cloud through nearest neighbor search. After registration, each point cloud point obtains a unique deformation value Δz, which can be positive (surface bulging outward) or negative (surface concave inward).

[0113] In some implementations, image preprocessing includes the following steps: First, using the pre-calibrated camera intrinsic parameters and distortion coefficients in S101, the original image is corrected pixel-by-pixel using the distortion correction formula proposed by Zhang Zhengyou's calibration method to eliminate radial and tangential distortion. Then, a bilateral filtering algorithm is used for noise reduction. This algorithm can preserve edge details while filtering out Gaussian noise, and the spatial domain standard deviation of the filtering parameters... Set to 2.0, pixel range standard deviation The threshold is set to 50.0. Next, adaptive histogram equalization is performed, dividing the image into 8×8 local blocks. Histogram equalization is applied to each block, and then bilinear interpolation is used to eliminate block artifacts, thereby enhancing the visibility of defects in low-contrast areas. Finally, background removal is performed based on threshold segmentation and morphological operations: the Otsu method is used to calculate a global threshold, separating the plug region from the background; then, an opening operation (erosion followed by dilation) is used to remove isolated small noise points, resulting in a foreground image containing only the plug region, and the image size is uniformly scaled to 512×512 pixels.

[0114] In some implementations, point cloud preprocessing includes the following steps:

[0115] First, the original point cloud under each pressure state is downsampled using voxels, with the voxel size set to 0.1 mm, reducing the number of points from approximately 100,000 to approximately 20,000 while preserving geometric features.

[0116] Then, statistical filtering is used to remove outliers: for each point, the average distance to its K nearest surrounding points (K is 30) is calculated, assuming that the distance follows a Gaussian distribution, and points whose average distance exceeds the global mean plus one standard deviation are removed.

[0117] Next, the coordinates are centered and normalized: the overall centroid coordinates of the rubber stopper point cloud are calculated. Subtract the centroid coordinates from the coordinates of all points so that the center of the point cloud is located at the origin. Then divide the coordinates by its maximum range (e.g., maximum radius) to scale it to the interval [-1, 1].

[0118] Finally, the deformation variable Δz is calculated using a point cloud registration algorithm: taking the point cloud under no-pressure conditions as a reference, the iterative nearest neighbor (ICP) algorithm is used to perform rigid registration for each point cloud under pressure conditions. Then, for each point in the registered pressure point cloud, the nearest neighbor point is searched in the reference point cloud, and the difference between the z coordinates of the two points is calculated as the deformation variable Δz. Δz is then added as an additional feature dimension to the feature vector of each point. Finally, the feature vector of each point is represented as a four-dimensional vector [x, y, z, Δz].

[0119] Based on the above technical solution, step S1 acquires multi-view images simultaneously using multiple cameras, acquires three-dimensional point clouds of multiple pressure states using structured light binocular stereo vision, and records pressure time-series parameters simultaneously. After specialized preprocessing of the images and point clouds, especially by using the deformation variable Δz as the fourth dimension feature of the point cloud, high-quality and information-rich input data is provided for subsequent dual-stream coding networks and cross-modal attention fusion, ensuring the consistency of appearance texture and geometric deformation information in time and space.

[0120] In one possible implementation of the embodiments of this application, combined with Figure 1 ,like Figure 2 As shown, the above S2 can be implemented through the following S201, S202 and S203, which are explained in detail below:

[0121] S201. Input the preprocessed image data into the image encoding stream, which uses a convolutional neural network-Transformer hybrid backbone network to extract multi-scale image features.

[0122] The role of the image coding stream is to extract discriminative texture, edge, and color features from a two-dimensional image and capture appearance information under different receptive fields through a multi-scale structure.

[0123] In some implementations, the specific structure of the image encoded stream is as follows:

[0124] The first step involves inputting a preprocessed RGB image of size 512×512×3 into the initial convolutional block. This convolutional block consists of a 7×7 convolutional layer (stride 2), a batch normalization layer, a ReLU activation function, and a 2×2 max pooling layer concatenated together, resulting in an output feature map of size 128×128×64. This step rapidly reduces the spatial resolution while expanding the number of channels to 64, extracting low-level texture and edge information, i.e., the first feature.

[0125] The second step involves passing the first feature through two residual blocks. Each residual block contains two 3×3 convolutional layers, a batch normalization layer, a ReLU activation function, and skip connections, resulting in an output feature map size of 64×64×128. Then, it passes through a lightweight Transformer block, which consists of a 3×3 depthwise separable convolutional layer, a windowed multi-head self-attention (W-MSA) layer, and a channeled multilayer perceptron. The W-MSA layer divides the feature map into 4×4 windows, calculating self-attention independently within each window, thus reducing computational complexity. After passing through the Transformer block, the feature map size remains 64×64×128, but the features at each location incorporate global contextual information, forming the second feature.

[0126] The third step involves inputting the second feature into a dilated convolutional spatial pyramid pooling module. This module contains four parallel dilated convolutional layers with dilation rates of 1, 2, 4, and 8, each using 256 3×3 convolutional kernels. The outputs of the four branches are concatenated along the channel dimension and then dimensionality-reduced using 1×1 convolutions, resulting in an output feature map of size 32×32×256, which is the third feature. This module expands the receptive field and captures multi-scale contextual information without increasing the number of parameters by using convolutional kernels with different dilation rates.

[0127] The fourth step involves global average pooling of the third feature to obtain a 1×1×256 global feature vector. This is then reduced to 128 channels via 1×1 convolution, and finally restored to 32×32×128 via bilinear interpolation upsampling, serving as the final multi-scale image feature output. This feature contains both global semantic information and preserves spatial details, and can be represented as follows: .

[0128] It should be noted that all convolutional and Transformer layers in the image encoding stream employ batch normalization and ReLU activation to accelerate training and prevent gradient vanishing. The window self-attention in the lightweight Transformer block uses relative position encoding, enabling the model to perceive the relative spatial relationships between pixels within the window. In the dilated convolutional spatial pyramid pooling module, the receptive fields of the convolutional kernels with different dilation rates are 3×3, 7×7, 15×15, and 31×31, respectively, thus enabling the simultaneous detection of minute scratches and larger-area material defects.

[0129] S202. Input the preprocessed 3D point cloud data into the point cloud encoding stream. The point cloud encoding stream adopts an improved PointNet++ network, which enhances local feature aggregation through sub-manifold sparse convolution and weighted position encoding, and outputs point cloud features using a hierarchical feature aggregation module.

[0130] The point cloud encoding stream extracts geometric and deformation features from the unstructured 3D point cloud. Each point in the preprocessed point cloud contains a feature vector [x, y, z, Δz], with N points in total. The improved PointNet++ network progressively downsamples the point cloud through three ensemble abstraction layers. In each ensemble abstraction layer, submanifold sparse convolutional kernels are used for local feature aggregation, and a weighted positional encoding module is introduced to sinusoidally encode the point coordinates and deformation before concatenating them into the point features. Finally, a hierarchical feature aggregation module concatenates the multi-scale features output from the three ensemble abstraction layers through skip connections, outputting the final feature vector for each original point cloud point.

[0131] In some implementations, the specific structure of the point cloud encoding stream is as follows:

[0132] The first abstraction layer: The input point cloud is downsampled from N points to M first keypoints (M=N / 4) using farthest point sampling. For each first keypoint, local feature aggregation is performed within a neighborhood of radius 0.2 using a sub-manifold sparse convolution kernel. Before aggregation, the coordinates and deformation of the points are sinusoidally encoded using a weighted position encoding module and then concatenated into the point features.

[0133] The sinusoidal position encoding formula is: Where L is the encoding dimension, typically 8. The four dimensions x, y, z, and Δz are encoded separately and then concatenated to obtain a 4×2L=64-dimensional position encoding vector. This encoded vector is concatenated with the original point features, then input into a submanifold sparse convolution kernel (kernel size 3×3×3, dilation rate 1) and a PointNet multilayer perceptron, outputting a first multidimensional feature vector with a dimension of 128. This feature vector contains local geometric structure and positional information.

[0134] The second set abstraction layer: The M keypoints output from the first set abstraction layer are downsampled into P second keypoints by sampling the farthest point, where P=M / 4, and the neighborhood radius is expanded to 0.4. Submanifold sparse convolution kernels are also used for local feature aggregation, and weighted position encoding is performed before aggregation to output a second multidimensional feature vector with a dimension of 256.

[0135] The third set abstraction layer: The P keypoints output from the second set abstraction layer are downsampled into Q third keypoints, where Q = P / 4, and the neighborhood radius is expanded to 0.8. The same process is applied to output a third multi-dimensional feature vector with a dimension of 512.

[0136] The hierarchical feature aggregation module concatenates the first, second, and third multi-dimensional feature vectors through skip connections. Specifically, it performs feature propagation (based on distance interpolation) on the second feature vector to recover the number of points in the first feature vector, performs two feature propagation operations on the third feature vector to recover the number of points in the first feature vector, then concatenates the three feature vectors along the channel dimension, and finally reduces the number of channels to 256 through a 1×1 convolution, obtaining the final point cloud features for each original point cloud point. .

[0137] It should be noted that farthest-point sampling ensures that the downsampled keypoints are evenly distributed on the rubber stopper surface, avoiding excessively dense local point clusters. Compared to ordinary sparse convolution, submanifold sparse convolution kernels only perform convolution operations on non-empty locations, significantly reducing computational cost. Weighted position encoding enables the network to perceive the absolute position and deformation amplitude of points, as the deformation patterns differ in different regions of the rubber stopper (e.g., the crown and tail regions show significant differences in deformation under pressure). Hierarchical feature aggregation, by fusing multi-scale features, preserves both local details (such as minor deformation anomalies) and global structural information (such as overall compression deformation).

[0138] S203. The multi-scale image features output from the image coding stream and the point cloud features output from the point cloud coding stream are used as the output of the dual-stream coding network for subsequent cross-modal attention fusion.

[0139] In this system, the two branches of the dual-stream coding network are computed independently and in parallel, without sharing parameters. The image coding stream outputs multi-scale image features. The point cloud features output by the point cloud encoding stream are 32×32×128 pixels in size. Size is The two features retain the core information of their respective modalities: image features contain rich texture and semantic information, while point cloud features contain precise 3D spatial coordinates, deformations, and local geometric structures. In the subsequent S3 step, these features will be deeply fused through a cross-modal attention mechanism.

[0140] In some implementations, to facilitate subsequent fusion, two independent adaptation layers can be added after the two-stream coding network: one adaptation layer flattens the spatial dimensions of the image features and projects them to the same number of channels as the point cloud features through a linear transformation; the other adaptation layer aggregates the point cloud features through max pooling and performs preliminary alignment with the image features. However, this application preferably preserves the original structure of the features and establishes a correspondence through bidirectional projection mapping during cross-modal attention fusion, thereby retaining their respective spatial or geometric structures.

[0141] It should be noted that the dual-stream coding network can be jointly trained end-to-end, or it can be pre-trained separately and then fine-tuned. The image coding stream can be initialized using weights pre-trained on ImageNet to speed up convergence; the point cloud coding stream can be initialized using PointNet++ weights pre-trained on ShapeNet or ModelNet. Due to the specific nature of the rubber stopper detection task, it is recommended to fine-tune it on a self-built rubber stopper dataset.

[0142] Based on the above technical solution, step S2 extracts multi-scale image features using a CNN-Transformer hybrid backbone network through an image encoding stream, capturing appearance information from local texture to global context. A point cloud encoding stream extracts point cloud features fused with coordinates and deformation using an improved PointNet++, and multi-scale geometric information is preserved through hierarchical feature aggregation. The two encoding streams work in parallel and independently, providing high-quality heterogeneous feature representations for subsequent cross-modal fusion, thus laying the feature foundation for simultaneously detecting appearance defects and airtightness defects.

[0143] In one possible implementation of the embodiments of this application, combined with Figure 1 ,like Figure 3 As shown, the above S3 can be implemented through the following S301, S302 and S303, which are explained in detail below:

[0144] S301. Input the pressure time series parameters into the hybrid coding network and encode them into pressure feature vectors.

[0145] The pressure time-series parameter is a sequence of pressure values ​​recorded by a pressure sensor under multiple preset pressure states in S1. This sequence not only contains numerical values ​​but also implicitly contains information about the order in which the pressure increases over time, reflecting the complete loading process of the rubber stopper from no pressure to high pressure. The role of the hybrid coding network is to map this one-dimensional time-series sequence into a fixed-length feature vector, enabling subsequent cross-modal attention fusion to dynamically adjust the fusion weights of image features and point cloud features according to the current pressure state.

[0146] In some implementations, the hybrid coding network consists of three layers connected in series. The first layer is a first multilayer perceptron with 64 neurons, which independently maps each stress value to a 64-dimensional vector. Let the stress temporal parameters be... Given a total of T=4 time steps, the output shape of the first multilayer perceptron is: The middle representation is H. The second layer is a Long Short-Term Memory (LSTM) network with a hidden layer dimension of 64. The input is each time step of the middle representation H. The LSTM captures the temporal dependencies between different stress values ​​through its internal state and gating mechanism. Its update formula is as follows:

[0147] ;

[0148] ;

[0149] ;

[0150] ;

[0151] ;

[0152] ;

[0153] in, For the input at time step t, In hidden state, The cell state is taken. The hidden state at the last time step t=3 is retrieved. As a temporally dependent feature, it is a 64-dimensional vector. The third layer is a second multilayer perceptron with 64 neurons, which... Further mapping to a 64-dimensional pressure feature vector .

[0154] It should be noted that LSTM layers can effectively learn the nonlinearity and memory effects during the stress loading process. For example, some defects only exhibit obvious deformation anomalies under specific stress thresholds, and LSTM can capture this stress sensitivity. Without temporal information, the fusion network will be unable to distinguish under which stress state the current feature was extracted, thus losing its ability to dynamically conditionalize.

[0155] S302. Using the pre-calibrated camera intrinsic parameter matrix, rotation matrix, and translation vector, establish a bidirectional projection mapping between the 3D point cloud and image pixels. Based on this mapping, generate the point cloud features derived from the image and the image features derived from the point cloud.

[0156] Cross-modal attention fusion presupposes establishing a spatial correspondence between image pixels and 3D point cloud points. Since the image encoding stream and the point cloud encoding stream output 2D feature maps and 3D point cloud features respectively, which reside in different coordinate spaces, spatial alignment must be achieved through projection transformation. Bidirectional projection mapping includes two directions: back projection from the image to the point cloud and forward projection from the point cloud to the image.

[0157] In some implementations, generating point cloud features and point cloud-derived image features based on projection mapping specifically includes the following steps:

[0158] First, using the pre-calibrated camera intrinsic parameter matrix in S101 Rotation matrix Translation vector This establishes the forward projection relationship from the 3D point cloud space to the 2D image plane. For the point cloud feature set output at a certain level of the point cloud encoding stream... ,in The feature vector of this point (dimension depends on the layer), projected onto the image pixel coordinates. The calculation formula is: Where K is a 3×3 intrinsic parameter matrix, It is a 3×4 extrinsic parameter matrix. After calculation, the pixel coordinates (which can be non-integers) corresponding to each point are obtained.

[0159] Then, construct a blank two-dimensional feature map with the same size as the current layer's image feature map. Initially, all positions are set to zero. For each projection point... The corresponding point cloud features Write the data to the pixel location. When multiple points are projected onto the same pixel, an average pooling strategy is used: record the features of all points falling into the pixel, sum them, and divide by the number of points. This results in a sparse two-dimensional feature map, called the "image features derived from the point cloud." In this feature map, each non-empty pixel carries the three-dimensional spatial coordinates, deformation, and local geometric features of the corresponding region's point cloud.

[0160] Secondly, a back projection relationship is established from the 2D image plane to the 3D point cloud space to generate "image-derived point cloud features". Back projection requires the use of depth maps. The depth value D(u,v) of each pixel is calculated in S1 using a structured light binocular stereo matching algorithm. This is for the image feature map output from the current layer of the image encoding stream. For each pixel (u,v), if the depth value D(u,v) is valid (non-zero), the coordinates of the corresponding 3D point are calculated using the following formula: In the formula, The inverse of the intrinsic parameter matrix. This is the transpose of the rotation matrix.

[0161] The calculated 3D point coordinates (x, y, z) along with the image feature vector of that pixel This constitutes a point. After transforming all effective pixels, a set of three-dimensional points is obtained. Each point has three-dimensional spatial coordinates and corresponding image semantic features, called "image-derived point cloud features," denoted as... The number of dots is the number of effective pixels (usually about 30% to 50% of the original image resolution).

[0162] Finally, the generated point cloud is exported as image features. Point cloud features derived from images As input to the cross-modal attention fusion module at this level, it is used for bidirectional cross-attention calculation in subsequent S303.

[0163] S303. Perform bidirectional cross-attention from image to point cloud and from point cloud to image in sequence, and use the pressure feature vector as a conditional input to the gating network to adaptively fuse and obtain multi-level cross-modal fusion features.

[0164] The cross-modal attention fusion mechanism is executed at multiple levels. Specifically, it fuses the corresponding outputs of each layer of the image encoding stream and the point cloud encoding stream to obtain fusion features at the first to fourth levels. Finally, these hierarchical features are aggregated into the final multi-level cross-modal fusion features. The specific workflow of each level's cross-modal attention fusion module includes three sub-steps: cross-attention from image to point cloud, cross-attention from point cloud to image, and pressure-gated adaptive fusion.

[0165] In some implementations, the internal operation of the cross-modal attention fusion module includes:

[0166] The first step is to calculate the cross-attention from the image to the point cloud. Let the point cloud features output by the point cloud encoding stream at the current level be... Where N is the number of points in the point cloud. The feature dimension is denoted as . The image features output by the image encoding stream are . Extracting point cloud features from the generated image. , where M is the number of effective depth pixels (i.e. the number of points in the point cloud derived from the image), and each point has three-dimensional spatial coordinates and corresponding image semantic features.

[0167] Point cloud features As a query, the point cloud features exported from the image are linearly transformed and used as the key and value. Specifically, three learnable linear transformation matrices are defined. , , ,calculate: , , ;in and These are the dimensions for keys and values, typically taken as... .

[0168] Then, the attention weights are calculated and summed using a weighted average. This result represents the enhanced information obtained from the semantics of the image for each point cloud point.

[0169] The original point cloud features are fused with this enhanced feature, for example, through residual connections and linear transformations: ; Obtain the point cloud features after image enhancement Dimensions remain unchanged .

[0170] The second step is to calculate the cross-attention from the point cloud to the image. Image features are then derived using the generated point cloud. This feature is a sparse two-dimensional feature map, where non-empty pixels carry geometric and deformation information of the point cloud. Image features... As a query, the point cloud image features are exported and linearly transformed to serve as the key and value.

[0171] Similarly, define three more linear transformation matrices. , , First, flatten the image features into : , , ;in It collects non-empty pixels from the sparse feature map as The matrix (L is the number of non-empty pixels).

[0172] Then calculate the attention. : Rearrange the results as H×W× Then, the number of channels is restored through linear transformation and concatenated with the feature residuals of the original image: ; Obtain image features after point cloud enhancement Dimensional preservation .

[0173] The third step is pressure-gated adaptive fusion. This involves combining the pressure feature vectors... (generally =64) and the enhanced point cloud features and image features are input into a gating network, and the fusion weights are dynamically calculated. First, the image features of the enhanced point cloud are... Perform spatial upsampling or feature recalibration to make it consistent with the feature scale of the original image (if it is already consistent, leave it unchanged). Then apply this to the point cloud features after image enhancement. The point cloud features are aggregated to the same spatial size as the image features using max pooling or attention pooling (optionally, different strategies can be used to maintain the point cloud structure at different levels). For simplicity, in higher-level fusion, the point cloud features can be mapped to the image feature size via spherical projection or feature propagation.

[0174] The structure of the gating network is as follows: The pressure feature vector is... The pressure feature map is obtained by copying and expanding it to the same size as the image feature space. Then and splicing along the channel dimension, then with... The network is concatenated, taking a network consisting of two 1×1 convolutional (or fully connected) layers as input, and finally outputting the fusion weights for each spatial location through a sigmoid activation function. :

[0175] ;in, It is necessary to first align its spatial scale to H×W through a learnable projection transformation (e.g., using point-based feature propagation or inverse distance weighted interpolation).

[0176] The final fusion features are: Output fusion features Where C is the preset number of fusion channels (usually 256).

[0177] The above operations are performed independently at the four levels in sequence, resulting in... , , , Then, the final multi-level cross-modal fusion features are formed by cascading the channel dimensions or adding them element by element.

[0178] It should be noted that hierarchical fusion can capture modal complementary information at different scales: low-level fusion preserves detailed textures and local geometry, making it suitable for detecting minor scratches and local deformation anomalies; high-level fusion includes global semantics and the overall deformation field, making it suitable for detecting large-area material shortages or overall sealing failures. The pressure gating mechanism causes the fusion weights to change with pressure. For example, in the unpressurized state, the deformation Δz=0, and the point cloud geometric features may not provide effective information. In this case, the gating network automatically reduces the weight of the point cloud features to avoid introducing noise.

[0179] Based on the above technical solution, step S3 encodes the pressure temporal parameters into pressure feature vectors and utilizes a bidirectional cross-attention and pressure-gated adaptive fusion mechanism to achieve deep interaction between image features and point cloud features at multiple levels. This design enables the network to dynamically fuse texture information and geometric deformation information according to the current pressure state, thereby simultaneously enhancing the perception of appearance defects and airtightness defects, and providing information-rich and pressure-adaptive cross-modal fusion features for the dual-task decoding of step S4.

[0180] In one possible implementation of the embodiments of this application, combined with Figure 1 ,like Figure 4 As shown, the above S4 specifically includes the following S401 to S403:

[0181] S401. Input the enhanced image features of the point cloud at each level into the appearance defect decoder. This decoder is based on the U-Net architecture and enhanced by Triplet Attention. At the same time, it uses the cross-modal fusion features and pressure feature vectors at each level as conditional modulation signals to output the pixel-level segmentation mask of the appearance defect.

[0182] Among them, the appearance defect decoder is used to classify and locate defects on the surface of the rubber stopper at the pixel level, and outputs a segmentation probability map with the same resolution as the input image. Each pixel is classified into one of the following categories: no defects, scratches, missing material, burrs, or bubbles.

[0183] The inputs to the appearance defect decoder include:

[0184] Main input: Image features enhanced from point clouds at four levels. , , , These correspond to 1 / 2, 1 / 4, 1 / 8, and 1 / 16 of the original image resolution, respectively, with the following number of channels: =64、 =128、 =256、 =512.

[0185] Conditional modulation input: cross-modal fusion features at each level ( =1, 2, 3, 4 (one for each level) and pressure feature vector Among them, cross-modal fusion features It is obtained through bidirectional cross-attention and stress-gated adaptive weighted fusion, and it simultaneously contains image texture information, point cloud geometric information and stress state information.

[0186] In some implementations, the appearance defect decoder uses the U-Net architecture, with the following specific structure:

[0187] (1) Encoder-side feature pyramid construction.

[0188] The decoder directly reuses the four-level point cloud-enhanced image features output from the cross-modal fusion stage as the feature pyramid on the encoder side, without additional encoding. These features have already incorporated the geometric information of the point cloud through cross-attention from the image to the point cloud, thus containing richer structural cues than the original image features.

[0189] (2) Decoder upsampling path.

[0190] The decoder employs a symmetrical upsampling path, comprising four upsampling stages. Each stage progressively doubles the spatial resolution of the feature map while halving the number of channels. Let the input feature of the k-th upsampling stage (k=1,2,3,4, where k=1 corresponds to the lowest resolution top layer) be denoted as . Its spatial dimensions are The number of channels is .

[0191] Phase 1 (Top level, k=1):

[0192] enter: Dimensions: H / 16×W / 16×512.

[0193] First, cross-modal feature modulation is performed: cross-modal features at the same level are fused. Adjusted to the same level through adaptive average pooling For the same space dimensions, , and then with Stitching along the channel dimension: ;in Compress the number of channels after splicing back to =512, keeping the feature dimension stable.

[0194] Next, pressure FiLM modulation is performed: utilizing pressure eigenvectors. Generate affine transformation parameters: ;in It is a two-layer fully connected network with an output dimension of , respectively corresponding to scaling factors and offset factor Channel-by-channel modulation of the feature map: ;in This indicates element-wise multiplication (broadcast to spatial dimensions).

[0195] Then through a transposed convolutional layer Upsampling is performed with a step size of 2, resulting in an output feature map size of H / 8 × W / 8, and the number of channels is halved to 256. ;

[0196] Phase 2 (k=2):

[0197] Input: Upsampling results from stage 1 (Dimensions H / 8×W / 8×256) and skip connection features from the corresponding level on the encoder side (Dimensions: H / 8 × W / 8 × 256).

[0198] First, the two are concatenated along the channel dimension and then fused using a 3×3 convolutional layer: The number of output channels remains at 256.

[0199] We apply the Triplet Attention module, which simultaneously captures interaction information across three dimensions: channel, height, and width. Let the input features be... The calculation process for Triplet Attention is as follows:

[0200] ;

[0201] Abbreviated as: ;in For the Sigmoid function, , , These represent global average pooling along the height, width, and channel dimensions, respectively. The output of this module has the same size as the input.

[0202] In this stage: ;

[0203] Then repeat the cross-modal eigenmodulation and pressure FiLM modulation: Pooling to the current size and with The segments are then concatenated and compressed back to 256 channels using a 1×1 convolution; Generate new , Perform FiLM modulation: ;

[0204] Transposed convolution upsampling to H / 4×W / 4 reduces the number of channels to 128. ;

[0205] Phase 3 (k=3):

[0206] enter: (H / 4×W / 4×128) and jump connection (H / 4×W / 4×128);

[0207] Perform splicing and fusion using 3×3 convolution: ;

[0208] Applying the Triplet Attention module: ;

[0209] Implementing cross-modal feature modulation: fusion FiLM modulation is performed;

[0210] After transposing and upsampling to H / 2×W / 2, the number of channels is reduced to 64. ;

[0211] Phase 4 (bottom layer, k=4):

[0212] enter: (H / 2×W / 2×64) and jump connection (H / 2×W / 2×64);

[0213] Perform splicing and fusion using 3×3 convolution: ;

[0214] Applying the Triplet Attention module: ;

[0215] Implementing cross-modal feature modulation: fusion FiLM modulation is then performed.

[0216] The final upsampling restores the feature map to the original input image resolution H×W, reducing the number of channels to 32. ;

[0217] (3) Output layer.

[0218] The feature map ultimately output by the upsampling path A 1×1 convolutional layer is used to map to the number of defect categories C=5 (no defects, scratches, missing material, burrs, bubbles), and then the Softmax activation function is applied along the channel dimension to obtain a pixel-level segmentation probability map: ;in This represents the probability that pixel (h, w) belongs to the c-th type of defect. Finally, the category with the highest probability for each pixel is taken as the defect label for that pixel, and the segmentation mask is output.

[0219] By combining cross-modal feature splicing with pressure FiLM modulation, pressure information and cross-modal fusion features are injected into each layer of the decoder, enabling the model to adaptively adjust the feature response according to the pressure conditions of the rubber stopper, thereby more accurately identifying pressure-related deformation and appearance defects.

[0220] It should be noted that by using cross-modal feature concatenation and pressure FiLM modulation, pressure information and cross-modal fusion features are injected into each layer of the decoder. This allows the model to adaptively adjust its feature response according to the pressure conditions applied to the rubber stopper, thereby more accurately identifying pressure-related deformations and appearance defects. Simultaneously, the combination of U-Net's skip connections and TripletAttention enables the decoder to utilize both high-resolution spatial details and low-resolution semantic information, thus accurately delineating defect contours. Since scratches and burrs on the rubber stopper are very small, ordinary U-Net easily loses details, while TripletAttention, through 3D attention recalibration, effectively preserves and amplifies these minute features.

[0221] S402. Input the enhanced point cloud features of the four levels of images into the airtightness defect decoder. The decoder is based on the differential point Transformer architecture. It amplifies the abnormal signal by calculating the difference in local deformation. At the same time, it uses the cross-modal fusion features and pressure feature vectors of each level as conditional modulation signals to output a three-dimensional thermal map of the airtightness defect.

[0222] Among them, the airtightness defect decoder is used to detect whether there is a risk of leakage or sealing failure of the rubber stopper during the pressure process, and outputs the probability that each point cloud point belongs to an airtightness defect, forming a three-dimensional heat map.

[0223] The inputs to the airtightness defect decoder include:

[0224] Main input: Four levels of image-enhanced point cloud features , , , These correspond to point cloud point counts from densest to sparsest: , , , The feature dimensions at each level are as follows: =128、 =256、 =512、 =1024.

[0225] Conditional modulation input: cross-modal fusion features at each level (l=1,2,3,4) and pressure eigenvectors Among them, cross-modal fusion features It is obtained through bidirectional cross-attention and stress-gated adaptive weighted fusion, and it simultaneously contains image texture information, point cloud geometric information and stress state information.

[0226] In some implementations, the airtightness defect decoder uses a parallel multi-branch differential point Transformer structure, as follows:

[0227] (1) Four parallel difference point Transformer branches.

[0228] Each branch corresponds to the input of a level, independently processes the point cloud features of that level, and extracts deformation anomaly features at that scale. All branches contain three sub-steps: stress-cross-modal modulation, differential set abstraction, and feature propagation.

[0229] Branch 1 (Denseest layer, number of points) ):

[0230] Sub-step 1.1: Pressure-cross-modal modulation. For the input point cloud of branch 1 Let the original feature of the i-th point be . Spatial coordinates Deformation .

[0231] First, cross-modal features at the same level are fused using pre-calibrated camera parameters. (Two-dimensional image features) are back-projected onto each point cloud point to obtain the cross-modal feature vector of that point. ,in This represents the number of channels for the cross-modal features at this level (e.g., 64).

[0232] Then, using the pressure feature vector The FiLM modulation parameters for this branch are generated using a lightweight multilayer perceptron: , ;

[0233] Perform a channel-wise affine transformation on the original point features: ;

[0234] Finally, the modulated features are concatenated with the cross-modal features to obtain the enhanced point features: .

[0235] Sub-step 1.2: Difference set abstraction. This involves abstracting the enhanced point features. Input the difference set abstraction layer along with its coordinates. This layer contains the following operations:

[0236] ① Sampling at the farthest point: from Sampled from each point The key points are denoted as set. The coordinates of each key point are .

[0237] ② Ball query: For each key point , with radius =0.3 Find the neighborhood point set .

[0238] ③ Calculation of deformation difference: For each point in the neighborhood Calculate the difference in deformation relative to the key points: This difference can amplify local deformation anomalies because leaks or internal defects can cause abrupt changes in the deformation gradient within the neighborhood.

[0239] ④ Local feature aggregation: Combines the original features (or modulated features) of each point in the neighborhood with... Spatial coordinate offset By splicing together, edge features are formed: ;

[0240] Then, submanifold sparse convolution is applied to all points in the neighborhood. Convolutional aggregation is performed, and then the local features of keypoint k_j are obtained through a shared PointNet multilayer perceptron (MLP): ;in, The output dimension is 256. The features of all keypoints constitute the downsampled feature map. .

[0241] Sub-step 1.3: Feature propagation. The downsampled feature map... The original point value was restored using inverse distance weighted interpolation. For each point i in the original point cloud, in the keypoint set... Find the three nearest keypoints and perform inverse distance weighted interpolation: , ; Obtain the coarse features of each point .

[0242] Then Compared with the original input features The pieces are stitched together and then fused using an MLP: Finally, the output of branch 1 is... .

[0243] Branch 2 (Medium level, number of points) ):

[0244] enter: , points Each point has a feature dimension of 256. Pressure-cross-modal modulation is also performed (using...). and ), to obtain enhanced features Then perform the difference set abstraction (radius). =0.4, downsampled to Output points. Finally, the features are propagated back. Point, get The specific process is the same as branch 1, only the parameters are different.

[0245] Branch 3 (sparser layer, number of points) ):

[0246] enter: , points Feature dimension 512. Post-modulation difference set abstraction (radius) =0.6, downsampled to Output points. Feature propagation back Point, get .

[0247] Branch 4 (most sparse layer, number of points) ):

[0248] enter: , points The feature dimension is 1024. After modulation, no further downsampling is performed because it is already at its most sparse; it is directly used as the feature of this layer, and then feature propagation is performed to obtain... .

[0249] (2) Multi-scale feature fusion.

[0250] Because the number of output points is different for the four branches ( They need to be unified to the original point number. The process of merging will be carried out. Details are as follows:

[0251] Output of Branch 2 Upsampled to using inverse distance interpolation Point, get .

[0252] Output of branch 3 First upsample to Point, then upsampled to Point, get .

[0253] Output of branch 4 Stepwise upsampling ( ),get .

[0254] It is important to note the output of branch 1. It is already at point N_1.

[0255] These four features are concatenated along the channel dimension to obtain multi-scale fused features: .

[0256] (3) Pointwise regression output layer.

[0257] Multi-scale fusion features Input a lightweight MLP for dimensionality reduction and regression: ;in It contains two fully connected layers: the first layer is 2816→256, the second layer is 256→128, and ReLU activation is used in between.

[0258] Then, the dimensionality reduction features of each point Use a regression head consisting of two fully connected layers and a sigmoid activation function: ;in , Output This represents the probability that the i-th point is an airtight defect.

[0259] Finally, the probabilities of all points form a three-dimensional heatmap: .

[0260] (4) Post-processing module: Defect location and clustering.

[0261] The probability value in the 3D heatmap is greater than a preset threshold. Points with a density coefficient of 0.5 are marked as candidate defect points. These points are then clustered using the DBSCAN density clustering algorithm.

[0262] Set cluster radius =0.5 (physical spatial unit, such as millimeter), minimum number of dots =10.

[0263] After clustering, several point clusters are obtained, and each point cluster represents a continuous airtightness defect region.

[0264] For each cluster of points, calculate its minimum axis-aligned bounding box (AABB) to obtain the coordinates of the eight vertices of the bounding box: ;

[0265] Output the three-dimensional spatial location and confidence level of each airtightness defect (the average probability within the cluster can be used).

[0266] It should be noted that the core of the Differential Point Transformer lies in the explicit calculation of deformation difference. It directly uses the local gradient of the deformation field as a feature, which is more sensitive to abnormal bulges or depressions caused by leakage than simply using the absolute value of deformation. Furthermore, through the above design, the airtightness defect decoder fully utilizes the image-enhanced point cloud features of all four levels, combines cross-modal fusion features and pressure feature vectors, and amplifies deformation anomalies through differential set abstraction, ultimately outputting a three-dimensional thermal map of airtightness defects that can be used for precise localization.

[0267] S403. Based on the segmentation mask of appearance defects and the three-dimensional thermal map of airtightness defects, the two-dimensional image position and three-dimensional spatial position of the defects are located respectively. The quality of the rubber stopper is comprehensively determined and a unified defect visualization result is output.

[0268] The positioning and judgment steps convert the outputs of S401 and S402 into quantifiable defect information, and project the three-dimensional spatial defects onto a two-dimensional image through coordinate transformation to achieve a visual overlay display of the defects.

[0269] In some implementations, the specific operations are as follows:

[0270] For appearance defects, a minimum area threshold is set. =5 pixels (can be adjusted according to actual detection accuracy). Traverse the segmentation mask. For each category c (c=1~4 corresponding to scratches, missing material, burrs, and bubbles), extract all connected regions with a probability greater than 0.5 and belonging to category c. For each connected region, if its pixel area is greater than... If the area is identified as a visual defect, its contour point coordinate sequence is recorded. .

[0271] For airtightness defects, a probability threshold is set. =0.5, minimum number of consecutive points =10. In the three-dimensional heat map The probability of selecting from the middle is greater than The point set is then used to cluster spatially adjacent points into clusters using the DBSCAN density clustering algorithm (neighborhood radius of 0.2, minimum number of points of 5). For each point cluster, if its number of points is greater than... If the condition is found to be airtight, the minimum axis-aligned bounding box of the cluster is calculated, and the coordinates of its eight vertices are output. The eight corner points.

[0272] During the comprehensive judgment, only when both the appearance defect list and the airtightness defect list are empty will the output "Pass" and a comprehensive confidence score (e.g., the average of the highest probabilities of all pixels) be output. Otherwise, "Fail" will be output, along with the contour coordinates of all appearance defects and the bounding box coordinates of all airtightness defects.

[0273] Furthermore, using the camera and point cloud coordinate system transformation matrix calibrated in S101, the eight vertices of the 3D bounding box of each airtight defect are back-projected onto the 2D image plane to obtain a 2D projected polygon. On the same RGB image, the outline of the appearance defect (e.g., red for scratches, blue for missing material) and the projected area of ​​the airtight defect (e.g., green for semi-transparent filling) are drawn with different colors to form a unified defect visualization result.

[0274] Based on the above technical solution, step S4 extracts appearance and airtightness defect information from image and point cloud features respectively through a dual-task decoding network, achieving pixel-level and point-level precise localization. Post-processing clustering and coordinate transformation provide intuitive defect visualization results. The entire process is non-destructive and fully automated, enabling online integrated inspection of the appearance and sealing performance of rubber stoppers.

[0275] Although this application has been described herein in conjunction with various embodiments, those skilled in the art, by reviewing the accompanying drawings, disclosure, and appended claims, will understand and implement other variations of the disclosed embodiments in carrying out the claimed application. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple instances. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce good results.

[0276] Although this application has been described in conjunction with specific features and embodiments, it is obvious that various modifications and combinations can be made thereto without departing from the spirit and scope of this application. Accordingly, this specification and drawings are merely illustrative descriptions of the application as defined by the appended claims, and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from the spirit and scope of this application. Thus, if such modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is also intended to include such modifications and variations.

Claims

1. A method for detecting defects in rubber stopper production based on image processing, characterized in that, include: Image data, three-dimensional point cloud data, and corresponding pressure time parameters of the rubber stopper are acquired under multiple preset pressure states, and the multi-view image data and three-dimensional point cloud data are preprocessed; wherein, the three-dimensional point cloud data includes the three-dimensional spatial coordinates of each point on the surface of the rubber stopper and the deformation relative to the pressureless state; A two-stream coding network is used to extract multi-scale image features from preprocessed image data and point cloud features from preprocessed point cloud data, respectively. The pressure time-series parameters are encoded into pressure feature vectors, and the image features, point cloud features, and pressure feature vectors are fused based on a cross-modal attention fusion mechanism to obtain cross-modal fused features. Using a dual-task decoding network, appearance defect detection and airtightness defect detection tasks are performed based on the cross-modal fusion features. The segmentation mask of appearance defects and the three-dimensional heat map of airtightness defects are output, and the defect location is located based on the segmentation mask and the three-dimensional heat map.

2. The method for detecting defects in rubber stopper production based on image processing according to claim 1, characterized in that, The process of acquiring the image data, 3D point cloud data, and corresponding pressure time-series parameters includes: Multi-view image data of the rubber stopper is obtained by simultaneously capturing multi-view images of the stopper from different angles under various preset pressure conditions using multiple industrial cameras. A coded grating pattern is projected onto the surface of the rubber stopper using structured light projection. A binocular stereo vision system is used to collect three-dimensional point cloud data of the rubber stopper surface under no pressure and at multiple incremental characteristic pressure threshold points. The pressure value at each characteristic pressure threshold point is recorded by a pressure sensor to form pressure time series parameters. The multiple incremental characteristic pressure threshold points are the preset pressure states. For each characteristic pressure threshold point, the point cloud data collected is obtained by subtracting the point cloud coordinates under no-pressure conditions from the point cloud coordinates under pressurized conditions, and the deformation of each point is calculated. The deformation is then added to the feature vector of the point as an attribute value of the three-dimensional point cloud data.

3. The method for detecting defects in rubber stopper production based on image processing according to claim 1, characterized in that, The preprocessing includes: The distortion correction of image data is performed using Zhang Zhengyou calibration method, and the corrected image is then subjected to bilateral filtering for noise reduction, histogram equalization, and background removal based on threshold segmentation and morphological operations to obtain preprocessed image data. The three-dimensional point cloud data is subjected to voxel downsampling and discrete points are removed by statistical filtering to obtain preprocessed three-dimensional point cloud data. The feature vector of each point in the preprocessed three-dimensional point cloud data is represented as [x,y,z,Δz], where (x,y,z) are the spatial coordinates after coordinate centering and normalization, and Δz is the deformation of each point.

4. The method for detecting defects in rubber stopper production based on image processing according to claim 1, characterized in that, The dual-stream coding network includes an image coding stream and a point cloud coding stream, wherein, The image encoding stream uses a CNN-Transformer hybrid backbone network. It uses a convolutional neural network (CNN) to extract the first feature of the image data, and uses a Transformer module to perform global self-attention processing on the first feature to obtain the second feature. Then, it uses a dilated convolutional spatial pyramid pooling module to perform multi-scale context aggregation on the second feature to obtain the third feature. Finally, it uses global average pooling and convolutional dimensionality reduction layers to process the third feature and output multi-scale image features. The point cloud encoding stream uses an improved PointNet++ network. The improved PointNet++ network improves upon the PointNet++ network by introducing sub-manifold sparse convolution to enhance local feature aggregation and adding a weighted position encoding module to encode the coordinates and deformation of points. Furthermore, through a hierarchical feature aggregation module, the multi-scale point cloud features output from each layer are cascaded through skip connections to output point cloud features.

5. The method for detecting defects in rubber stopper production based on image processing according to claim 4, characterized in that, The improved PointNet++ network includes a first set abstraction layer, a second set abstraction layer, and a third set abstraction layer; The first set abstraction layer downsamples the input point cloud from N points to M first key points through farthest point sampling. Then, within the first neighborhood radius of each first key point, local feature aggregation is performed using a sub-manifold sparse convolution kernel. Before aggregation, the coordinates and deformation of the points are sinusoidally encoded and concatenated into the point features through a weighted position encoding module, outputting the first multidimensional feature vector. Here, M is taken as one-quarter of N, and the first neighborhood radius is set to 0.

2. The second set abstraction layer downsamples the M key points into P second key points. Then, within the second neighborhood radius of each second key point, local feature aggregation is performed using a submanifold sparse convolution kernel. Weighted position encoding is performed before aggregation to output the second multidimensional feature vector. Here, P is one-quarter of M, and the second neighborhood radius is 0.

4. The third set abstraction layer downsamples P second key points into Q third key points. Then, within the third neighborhood radius of each third key point, local feature aggregation is performed using a submanifold sparse convolution kernel. Weighted position encoding is performed before aggregation to output a third multidimensional feature vector. Here, Q is one-quarter of P, and the third neighborhood radius is 0.

8. Finally, the hierarchical aggregation feature module concatenates the first multidimensional feature vector, the second multidimensional feature vector, and the third multidimensional feature vector through skip connections to output point cloud features. Wherein, N>M>P>Q, and the dimension of the first multidimensional feature vector < the dimension of the second multidimensional feature vector < the dimension of the third multidimensional feature vector.

6. The method for detecting defects in rubber stopper production based on image processing according to claim 1, characterized in that, Encoding the pressure time-series parameters into a pressure feature vector includes: The stress time-series parameters are input into a hybrid coding network. The first layer of the hybrid coding is a first multilayer perceptron with multiple neurons, which maps each stress value to a multidimensional vector to obtain an intermediate representation. The second layer is a long short-term memory network layer, which outputs the hidden states of all time steps based on the intermediate representation, and takes the hidden state of the last time step as the temporal dependency feature. The third layer is a second multilayer perceptron with multiple neurons, which maps the temporal dependency feature to a multidimensional stress feature vector.

7. The method for detecting defects in rubber stopper production based on image processing according to claim 1, characterized in that, The cross-modal attention fusion mechanism specifically includes a hierarchical fusion step, which performs cross-modal attention fusion on the features of each level output from the image encoding stream and the corresponding level features output from the point cloud encoding stream, respectively, to obtain multi-level cross-modal fused features; wherein: The first feature in the image encoding stream and the first multidimensional feature vector in the point cloud encoding stream are processed by the first-level cross-modal attention fusion module to obtain the first-level cross-modal fusion feature; The second feature in the image encoding stream and the second multidimensional feature vector in the point cloud encoding stream are processed by the second-level cross-modal attention fusion module to obtain the second-level cross-modal fusion feature; The third feature in the image encoding stream and the third multidimensional feature vector in the point cloud encoding stream are processed by the third-level cross-modal attention fusion module to obtain the third-level cross-modal fusion feature; The multi-scale image features and the point cloud features are processed by the fourth-level cross-modal attention fusion module to obtain the fourth-level cross-modal fusion features.

8. The method for detecting defects in rubber stopper production based on image processing according to claim 7, characterized in that, The specific workflow of the cross-modal attention fusion module is as follows: Using pre-calibrated camera intrinsic parameter matrices, rotation matrices, and translation vectors, a bidirectional projection mapping between 3D point clouds and image pixels is established. Based on the bidirectional projection mapping, the image feature map and the depth map are back-projected to generate point cloud features derived from the image. Each point in the point cloud features derived from the image has three-dimensional spatial coordinates and carries corresponding image semantic features. The depth map is a two-dimensional matrix composed of the depth values ​​of each pixel calculated from the acquired raster image through structured light projection and binocular stereo matching algorithm. The image feature map is a two-dimensional feature map output by the image encoding stream at the current level. Based on the bidirectional projection mapping, the point cloud feature map is forward-projected onto the image plane to generate image features derived from the point cloud. The image features derived from the point cloud are two-dimensional feature maps, and each non-empty pixel position carries the geometric and deformation features of the corresponding point cloud. The point cloud feature map is a set of point cloud feature vectors output by the point cloud encoding stream at the current level. First, cross-attention from image to point cloud is performed, including: using the point cloud feature map as the query vector, using the point cloud features exported from the image as the key vector and value vector after linear transformation, using the scaling dot product attention formula to calculate the attention weight between each point cloud point and each point in the point cloud exported from the image, and then weighting and summing the attention weights with the point cloud features exported from the image to obtain the point cloud features after image enhancement. Then, cross-attention from point cloud to image is performed, including: using the image feature map as the query vector, using the image features exported from the point cloud as the key vector and value vector after linear transformation, using the scaling dot product attention formula to calculate the attention weight of each image pixel and each pixel in the image features exported from the point cloud, and then weighting and summing the attention weights with the image features exported from the point cloud to obtain the image features after point cloud enhancement. The pressure feature vector, along with the image-enhanced point cloud features and the point cloud-enhanced features, are input into a gating network to output cross-modal fusion features.

9. The method for detecting defects in rubber stopper production based on image processing according to claim 7, characterized in that, The appearance defect detection task is performed by an appearance defect decoder, which is built on the U-Net architecture. The specific workflow includes: The input to the appearance defect decoder is the image features enhanced by point clouds at each level, which are organized into the encoder-side feature pyramid of U-Net according to resolution from high to low. The cross-modal fusion features and stress feature vectors at each level are injected as conditional modulation signals into each upsampling stage of the U-Net decoder: In each upsampling stage, the cross-modal fusion features are adaptively pooled to adjust to the spatial size of the current level feature map, and then channel-concatenated with the current level feature map. The stress feature vector is passed through a multilayer perceptron to generate affine transformation parameters. The affine transformation parameters are then used to perform FiLM modulation with the channel-concatenated features to obtain the modulation features. The U-Net decoder employs a symmetrical upsampling path and receives point cloud-enhanced image features from the corresponding level of the encoder-side feature pyramid via skip connections. A Triplet Attention module is embedded in each skip connection to generate three-dimensional attention weights along the channel dimension, height dimension, and width dimension respectively, and adaptively weights the feature maps passed by the skip connections. The upsampling path ultimately restores the feature map to the same spatial resolution as the input image, and the number of output channels is equal to the number of defect categories, which include no defects, scratches, missing material, burrs, and bubbles. Finally, a pixel-level segmentation probability map is output through 1×1 convolution and the Softmax activation function, which serves as a segmentation mask for appearance defects.

10. The method for detecting defects in rubber stopper production based on image processing according to claim 7, characterized in that, The airtightness defect detection task is performed by an airtightness defect decoder, which is built based on a differential point Transformer architecture. The specific workflow includes: The input to the airtightness defect decoder is the point cloud features enhanced by the image at each level, namely the first-level point cloud features, the second-level point cloud features, the third-level point cloud features, and the fourth-level point cloud features. The point cloud features of each layer are input into a corresponding parallel branch. Each branch independently performs the following operations: First, the cross-modal fusion features of the same layer are back-projected onto each point cloud point using camera parameters to obtain the cross-modal feature vector of the point cloud point. Then, the pressure feature vector is used to generate the affine modulation parameters of the current branch through a multilayer perceptron. The affine modulation parameters are used to perform FiLM modulation on the point cloud features of the current layer. Finally, the modulated features are concatenated with the cross-modal feature vector to obtain the enhanced point cloud features. Then, differential set abstraction is performed, including farthest point sampling on the enhanced point cloud features. To reduce the number of points, the deformation difference between the neighboring points and the keypoint is calculated in the ball query neighborhood of each sampled keypoint. The deformation difference is then concatenated with the spatial coordinate offset of the neighboring points and the enhanced features, and input into the submanifold sparse convolution and shared multilayer perceptron. The output is the local aggregated features of the keypoint. The features of all keypoints constitute the downsampled feature map. Finally, feature propagation is performed, including restoring the number of points of the original input of the branch to the downsampled feature map through inverse distance weighted interpolation, and concatenating the interpolation result with the original enhanced features and then fusing them through the multilayer perceptron to output the propagated features of the branch. The four branches output different numbers of propagation feature points, which are denoted as the first propagation feature, the second propagation feature, the third propagation feature, and the fourth propagation feature, respectively. The number of points in the first propagation feature is equal to the number of points in the original point cloud. The second, third, and fourth propagation features are upsampled to the number of points in the original point cloud through inverse distance weighted interpolation to obtain upsampled features with the same number of points as the first propagation feature. The first propagation feature and the three upsampled features are concatenated along the channel dimension to obtain the multi-scale fused feature. Multi-scale fused features are input into a multilayer perceptron, which outputs the dimensionality-reduced features of each point cloud point. The dimensionality-reduced features of each point are then input into a regression head, which outputs the probability that each point cloud point belongs to an airtightness defect. The probabilities of all points constitute a three-dimensional heatmap. The regression head consists of two fully connected layers and a sigmoid activation function. Finally, the post-processing module performs threshold filtering on the 3D heat map, marking points with probability values ​​greater than a preset threshold as candidate defect points. The DBSCAN density clustering algorithm is used to cluster the candidate defect points to obtain several point clusters. The minimum axis-aligned bounding box of each point cluster is calculated, and the vertex coordinates of each bounding box are output as the 3D spatial location of the airtightness defect.