A deep learning-based food packaging defect detection method

CN122737579APending Publication Date: 2026-09-11盐城市质量技术监督综合检验检测中心
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610802802.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-04
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

[0007]为解决上述技术问题,提供一种基于深度学习的食品包装缺陷检测方法,本技术方案解决了上述的不能利用不同模态进行互补的问题

Benefits of technology

本发明通过可见光、近红外、3D结构光与声学四种异构传感器的协同感知,能够覆盖表面缺陷、内部异物、透明包装密封完整性以及材质异常等各类缺陷的检测需求,四模态的互补性确保了单一模态的感知盲区可由其他模态弥补,检测覆盖率得到了显著提升;三级分组协同注意力网络将模态交互限定在组内与组间两个层级,避免了全连接注意力的指数级计算开销。几何组聚焦于空间形貌的一致性校验,成分组聚焦于材质与内部状态的协同判断,组间的协同注意力负责全局一致性校准,在保证特征交互充分性的同时大幅降低了计算复杂度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122737579A_ABST
    Figure CN122737579A_ABST
Patent Text Reader

Abstract

This invention discloses a deep learning-based method for detecting defects in food packaging, relating to the field of packaging inspection technology. The method includes: simultaneously acquiring multimodal data of food packaging using a visible light camera, a near-infrared (NIR) camera, a 3D structured light sensor, and an acoustic sensor, obtaining visible light images, NIR images, 3D point cloud data, and acoustic signals, respectively; synchronizing and aligning the timestamps of each modal data; and using the 3D point cloud data as a spatial reference, mapping the visible light image and NIR image to a unified coordinate system containing the 3D point cloud through perspective projection. This invention, through the collaborative sensing of four heterogeneous sensors—visible light, near-infrared, 3D structured light, and acoustic—can cover the detection needs of various defects such as surface defects, internal foreign objects, the integrity of transparent packaging seals, and material abnormalities. The complementarity of the four modalities ensures that the blind spots of a single modality can be compensated by other modalities, significantly improving the detection coverage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of packaging inspection technology, specifically to a method for detecting defects in food packaging based on deep learning. Background Technology

[0002] The food packaging industry is a typical continuous, high-volume production industry. Food production lines, such as those for beverages, dairy products, and condiments, can operate at speeds of 300 to 1,000 packaging units per minute. Traditional manual visual inspection methods face severe challenges, so automated inspection is gradually becoming the mainstream.

[0003] Patent CN119180793A discloses a wind turbine blade damage detection method based on four-modal fusion. It achieves comprehensive damage perception through data fusion of four sensors: visible light, infrared, thermal imaging, and ultrasonic waves. The fusion strategy of this patent adopts a multi-level feature interaction design. However, its cross-modal attention mechanism does not consider the functional correlation between modes. Various specific types of defects are applicable to different modes. The solution treats each mode equally and ignores the complementary value of different modes in specific defect types.

[0004] Patent CN119540241A proposes a PCB board defect detection method based on bimodal attention. Visible light images and X-ray images interact with each other through a cross-modal Transformer. This method introduces a query-key-value cross-computation mechanism in the cross-modal attention design. However, it uses fully connected attention computation, and the computational complexity increases sharply when the number of modalities increases, making it difficult to extend to more modalities.

[0005] Another paper, Crossmodal Feature Mapping, published at CVPR 2024, proposed a unified framework for cross-modal feature mapping. However, its validation scenarios mainly consist of natural images and it has not been optimized for the dual requirements of real-time performance and accuracy in industrial inspection. The MVTec3D-AD dataset provides a standard benchmark for 3D defect detection. The introduction of 3D point cloud data significantly improves the detection accuracy of complex surface defects, but it only focuses on the 3D visual modality and does not involve other physical modalities such as acoustics.

[0006] Existing technologies have failed to provide a multimodal fusion detection scheme that balances detection accuracy, computational efficiency, and industrial reliability. This application proposes a deep learning-based food packaging defect detection method to overcome the above-mentioned shortcomings. Summary of the Invention

[0007] To address the aforementioned technical problems, a deep learning-based method for detecting defects in food packaging is provided. This technical solution solves the problem of not being able to utilize different modalities for complementarity.

[0008] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A deep learning-based method for detecting defects in food packaging includes: S1. Multimodal data of food packaging are collected simultaneously using a visible light camera, a near-infrared (NIR) camera, a 3D structured light sensor, and an acoustic sensor to obtain visible light images, NIR images, 3D point cloud data, and acoustic signals, respectively. S2. The data of each modality are time-stamped and aligned. Using the 3D point cloud data as a spatial reference, the visible light image and NIR image are mapped to the unified coordinate system of the 3D point cloud through perspective projection. The acoustic signal is converted into a spatially distributed detection signal through spatial correlation with the 3D point cloud to obtain the spatiotemporally aligned four-modal data. S3. Input the spatiotemporally aligned four-modal data into the three-level grouped collaborative attention fusion network for feature fusion. The three-level grouped collaborative attention fusion network performs the following in sequence: first-level intramodal self-attention modeling, and each modality independently extracts intramodal long-range dependency features. The second level of intra-group cross-modal attention interaction involves grouping visible light modes and 3D structured light modes into a geometry group, and NIR modes and acoustic modes into a component group, and performing cross-modal attention interaction within each group. The third level of inter-group collaborative attention fusion performs cross-group attention calculation between the outputs of the geometry group and the component group to generate four-modal collaborative global features; S4. Based on the four-modal collaborative global features, calculate the deviation between each modality and the normal sample database as the absolute anomaly score, and calculate the weighted sum of the prediction errors of each cross-modality as the cross-modal inconsistency measure. The absolute anomaly score and the cross-modal inconsistency measure are weighted and fused to generate the final anomaly score. S5. Output the defect detection results based on the final anomaly score.

[0009] Preferably, in the second-level intra-group cross-modal attention interaction, the geometry group and the component group respectively perform bidirectional feature interaction through a cross-query-key-value mechanism: positive cross-modal attention is calculated by using the query vector of one modal feature and the key vector and value vector of the other modal feature, while the query end and key-value end are exchanged to calculate reverse cross-modal attention; In the third-level inter-group collaborative attention fusion, the geometric group fusion feature and the component group fusion feature are used as the query end and key value end respectively to perform bidirectional cross-group attention calculation. The attention outputs between the two groups are concatenated and then linearly projected to generate four-modal collaborative global features.

[0010] Preferably, the cross-modal prediction error in step S4 is calculated as follows: For the four modes—visible light, NIR, 3D structured light, and acoustic—are paired in an ordered manner, and cross-modal prediction networks are trained separately for each mode. Each cross-modal prediction network is trained using a modal... The features are used as inputs, and the mode is predicted by a multilayer perceptron. The features are used to calculate the mean square error between the predicted features and the actual extracted features as the cross-modal prediction error. ; Cross-modal inconsistency measurement The calculation formula is: ; in, For modality For modes The prediction weights are dynamically adjusted based on the signal-to-noise ratio and signal strength of each modality in the current detection scenario; the final anomaly score... The calculation formula is: ; in, For absolute anomaly scoring, and These are learnable fusion weight parameters.

[0011] Preferably, when the absolute anomaly score A is lower than the preset defect judgment threshold and the cross-modal inconsistency measurement... When the inconsistency exceeds a preset inconsistency threshold, a second, more detailed analysis is triggered to locate the discrepancy. The cross-modal prediction error term with the largest absolute contribution The high-resolution Region of Interest (ROI) corresponding to the modality pair is extracted from the original acquired data. The feature extraction and cross-modal prediction errors are recalculated on this ROI, and the absolute anomaly score is updated. and cross-modal inconsistency measurement The final anomaly score was then recalculated. .

[0012] Preferably, before step S3, a modal dynamic routing step is included: the spatiotemporally aligned four-modal data is extracted using lightweight features and then input into the defect type prediction head to obtain the defect type prediction probability distribution of the current sample; The probability distribution of defect type prediction is input into a learnable routing network. The routing network outputs the weight configuration vector of each modality. The weight configuration vector is multiplied with the feature map of each modality channel by channel to obtain the weighted modal features. The weighted modal features are then input into a three-level grouped collaborative attention fusion network.

[0013] Preferably, the routing network is trained in the following way: using the four-modal features of historical detection samples and labeled defect types as training data, the loss function of the routing network includes defect type prediction classification loss and detection accuracy reward of modality weight allocation, the detection accuracy reward of modality weight allocation is calculated based on the defect detection accuracy output by the fusion network after the weighted four-modal features are processed, and the defect type prediction head, the routing network and the three-level grouped collaborative attention fusion network are jointly trained end-to-end.

[0014] Preferably, an adaptive modal scheduling step is further included between step S1 and step S3: The expected defect type distribution vector for the current batch Real-time signal-to-noise ratio of each mode and system computing resource utilization Concatenate into a state vector Input a deep Q-network, and the deep Q-network outputs the Q-values ​​corresponding to four actions. Select the action with the largest Q-value as the modal activation mode. Modal activation modes include one of the following: all four modal activations, geometric group modal activations, component group modal activations, or single modal activations; In step S3, the corresponding fusion path is executed according to the activation mode: when the full modality is activated, the complete three-level fusion is executed; when the geometry group is activated, only the first-level intramodal self-attention and the second-level cross-modal attention within the geometry group are executed; when the component group is activated, only the first-level intramodal self-attention and the second-level cross-modal attention within the component group are executed; when the single modality is activated, only the first-level intramodal self-attention is executed. After each batch is completed, the reward value is calculated based on the defect detection rate, false alarm rate, and detection delay. ,in , , The weighting coefficients and The parameters of the deep Q-network are updated using reward values ​​through temporal difference learning; When the maximum component of the expected defect type distribution vector is lower than the confidence threshold, skip the deep Q-network decision and directly select the full-modal activation mode.

[0015] Preferably, the cross-modal inconsistency metric is calculated in real time during the detection process. ,when A modal fault is determined to exist when the dynamic diagnostic threshold is exceeded. The dynamic diagnostic threshold is the mean of the cross-modal inconsistency measure in historical normal operation data. Plus Double standard deviation ,Right now ,in This is a preset multiplier factor; The fault mode localization method is as follows: analyze the prediction error terms of each cross-modal mode. right The contribution, if it includes modality All as prediction sources or prediction targets The sum of the contributions of each item accounts for the total If the proportion exceeds a preset proportion threshold, then the mode is determined. It is a fault mode; When a modal fault is detected, the deep Q network scheduler automatically switches to a degraded modal activation mode that does not contain the faulty mode.

[0016] Preferably, a four-level degradation operation mode is defined: L0 level is full four-mode activation, L1 level is three-mode activation after removing the acoustic mode, L2 level is only visible light mode and NIR mode retained, and L3 level is only visible light mode retained; each degradation mode corresponds to a set of adjusted defect judgment threshold parameters in step S5. The lower the degradation level, the tighter the defect judgment threshold is; when the fault mode recovers, the parameters of the fusion network are updated online using the cross-modal corresponding data accumulated during normal operation.

[0017] Preferably, the specific method for using 3D point cloud data as a spatial reference in step S2 is as follows: a unified coordinate system is constructed using the 3D point cloud acquired by the 3D structured light sensor. The pixel coordinates of the visible light image and the NIR image are mapped to the 3D point cloud coordinate system through perspective projection using the pre-calibrated visible light camera extrinsic matrix and NIR camera extrinsic matrix to obtain visible light feature map and NIR feature map that are spatially aligned with the 3D point cloud. The specific method for converting acoustic signals into spatially distributed detection signals is as follows: an acoustic sensor array is arranged on the side of the conveyor belt to collect multi-channel time-domain acoustic signals. Combined with the position coordinates of the packaged object in the 3D point cloud, the acoustic signals of each channel are mapped to the corresponding spatial positions in the 3D point cloud coordinate system to generate spatially distributed acoustic detection feature vectors.

[0018] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention utilizes the collaborative sensing of four heterogeneous sensors—visible light, near-infrared, 3D structured light, and acoustic—to cover the detection needs of various defects, including surface defects, internal foreign objects, the integrity of transparent packaging seals, and material anomalies. The complementarity of the four modalities ensures that the blind spots of a single modality can be filled by the other modalities, significantly improving the detection coverage. The three-level grouped collaborative attention network confines modal interactions to two levels: intra-group and inter-group, avoiding the exponential computational overhead of fully connected attention. The geometry group focuses on the consistency verification of spatial morphology, the component group focuses on the collaborative judgment of material and internal state, and the inter-group collaborative attention is responsible for global consistency calibration, greatly reducing computational complexity while ensuring sufficient feature interaction.

[0019] A cross-modal prediction network is used to construct prediction relationships between modes, with prediction error serving as a measure of inconsistency. When a mode exhibits an abnormal response due to noise or local interference, the prediction values ​​of other modes can provide a reference. The inconsistency component in the anomaly score can distinguish between single-modal noise and actual defects, thereby effectively reducing the false detection rate. The adaptive modal scheduling mechanism can dynamically adjust the participation weights of each mode according to the expected defect type distribution, improving the detection sensitivity of specific defect types. The cross-modal consistency fault self-diagnosis module monitors the system health status in real time, promptly identifying and locating faulty modes. The four-level degradation configuration (L0 to L3) ensures that the system can still maintain basic detection functions when some sensors fail, avoiding production line downtime. Furthermore, this invention uses 3D point clouds as a spatial reference for coordinate transformation, achieving alignment of visible light images, near-infrared images, and acoustic signals in a unified spatial coordinate system. By leveraging the accuracy advantage of 3D structured light sensors in spatial positioning, it avoids detection deviations caused by spatial calibration errors in multi-sensor systems. Attached Figure Description

[0020] Figure 1 This is a schematic flowchart of the packaging defect detection method of the present invention; Figure 2 This is a schematic diagram of the three-level attention collaborative attention network structure of the present invention; Figure 3 A schematic diagram of the structure of a cross-modal prediction network and an inconsistency metric; Figure 4 A block diagram of an adaptive modal scheduling and fault degradation system; Figure 5 This is a schematic diagram illustrating the principle of multimodal spatial alignment. Detailed Implementation

[0021] The following description is intended to disclose the invention and enable those skilled in the art to implement it. The preferred embodiments described below are merely examples, and other obvious variations will occur to those skilled in the art.

[0022] Reference Figure 1 and 5 As shown, a deep learning-based method for detecting defects in food packaging includes: S1. Multimodal data of food packaging are collected simultaneously using a visible light camera, a near-infrared (NIR) camera, a 3D structured light sensor, and an acoustic sensor to obtain visible light images, NIR images, 3D point cloud data, and acoustic signals, respectively. S2. The data of each modality are time-stamped and aligned. Using the 3D point cloud data as a spatial reference, the visible light image and NIR image are mapped to the unified coordinate system of the 3D point cloud through perspective projection. The acoustic signal is converted into a spatially distributed detection signal through spatial correlation with the 3D point cloud to obtain the spatiotemporally aligned four-modal data. S3. Input the spatiotemporally aligned four-modal data into the three-level grouped collaborative attention fusion network for feature fusion. The three-level grouped collaborative attention fusion network performs the following in sequence: first-level intramodal self-attention modeling, and each modality independently extracts intramodal long-range dependency features. The second level of intra-group cross-modal attention interaction involves grouping visible light modes and 3D structured light modes into a geometry group, and NIR modes and acoustic modes into a component group, and performing cross-modal attention interaction within each group. The third level of inter-group collaborative attention fusion performs cross-group attention calculation between the outputs of the geometry group and the component group to generate four-modal collaborative global features; S4. Based on the four-modal collaborative global features, calculate the deviation between each modality and the normal sample database as the absolute anomaly score, and calculate the weighted sum of the prediction errors of each cross-modality as the cross-modal inconsistency measure. The absolute anomaly score and the cross-modal inconsistency measure are weighted and fused to generate the final anomaly score. S5. Output the defect detection results based on the final anomaly score.

[0023] Visible light cameras use resolutions no lower than For industrial cameras with a specific pixel count, the lens focal length is determined based on the working distance and field of view; a typical configuration is... A fixed-focus lens captures visible images of the food packaging surface, including visual information such as color, texture, printing, scratches, and stains.

[0024] The operating wavelength range of near-infrared cameras is to Using an InGaAs sensor to obtain a sensitive response in this band, near-infrared imaging can partially penetrate transparent packaging materials such as plastic film to obtain near-infrared reflection information inside the packaging, which has unique advantages for detecting the filling status of the contents and the presence of foreign matter.

[0025] A 3D structured light sensor projects an coded light pattern, which is then captured by an infrared camera to reconstruct the three-dimensional spatial coordinates using the principle of triangulation. This embodiment employs a planar structured light scheme, which can acquire the data in a single operation. For dense point clouds with high resolution, depth measurement accuracy is better than 3D point cloud data provides three-dimensional topographic information of packaging, which is crucial for detecting structural defects such as unevenness in the sealing area, lid tilting, and can deformation.

[0026] The acoustic sensor employs an air-coupled ultrasonic sensor array, containing at least eight MEMS microphone units arranged in a circular array. The transmitter of the acoustic sensor emits broadband ultrasonic pulses (frequency range...). to The receiver collects ultrasonic signals reflected from the packaging surface and transmitted through the packaging interior. The acoustic signals are highly sensitive to the integrity of the sealing interface. When there is a micro-leak or poor sealing, the propagation characteristics of the ultrasonic signals, such as attenuation, time delay, and spectral changes, will be significantly altered.

[0027] Four types of sensors achieve synchronous data acquisition under the control of a unified trigger signal. The trigger interval is set according to the production line speed and inspection requirements, ensuring that each packaging unit completes a full set of multimodal data acquisition when passing through the inspection station. The acquisition frequency and exposure parameters of each sensor can be configured independently according to the specific inspection task.

[0028] In S2, spatiotemporal data alignment is utilized in the time dimension, with data collected by each sensor carrying high-precision timestamps (time resolution better than...). In the data preprocessing stage, the data of each modality are first aligned to a unified sampling time based on the timestamp. If the sampling rate of a certain modality differs from that of other modalities, resampling is performed on a unified time reference using interpolation methods (such as cubic spline interpolation) to ensure that the four modalities are strictly synchronized in the time dimension.

[0029] In the spatial dimension, 3D point cloud data is used as the spatial reference for coordinate system one. The relationship between the optical center of the 3D structured light sensor and the coordinate origin is determined during the system calibration stage. Its output point cloud data is directly located in the sensor coordinate system. For visible light cameras and near-infrared cameras, it is necessary to establish the mapping relationship between pixel coordinates and spatial coordinates through a pinhole camera model. ; in, As a scale factor, For pixel coordinates, For the camera intrinsic parameter matrix, and These are the rotation matrix and translation vector of the camera relative to the 3D sensor coordinate system, respectively. The three-dimensional coordinates of the spatial points; after obtaining the intrinsic and extrinsic parameters of each camera in advance through the Zhang Zhengyou calibration method, each pixel in the visible light image and near-infrared image can be mapped to the spatial coordinate system of the 3D point cloud through perspective projection.

[0030] Acoustic signals themselves do not contain spatial location information and need to be spatialized through spatial association with 3D point clouds. This embodiment establishes a spatial mapping relationship between the acoustic signal and the spatial arrangement parameters of the acoustic sensor array and a sound wave propagation model. Specifically, the positions of each unit in the acoustic sensor array are precisely calibrated in the 3D sensor coordinate system, and the spatial distribution of the acoustic signal is estimated using beamforming technology. ; in, For spatial location Acoustic energy at the location, This refers to the number of microphone units. For the first Weighting coefficients for each microphone, For the corresponding acoustic signal, The wavenumber vector is used; by performing gridded calculations within the region of interest, the acoustic signal is converted into a spatial distribution map with the same resolution as the 3D point cloud. After the above processing, the visible light image, near-infrared image and acoustic signal are all mapped to a unified coordinate system based on the 3D point cloud, forming a spatiotemporally aligned four-modal data representation.

[0031] like Figure 2 As shown, the structure of the three-level group collaborative attention fusion network in S3 is as follows: Level 1: Intramodal self-attention modeling, where the spatiotemporally aligned four-modal data are... (Visible light characteristics) (Near-infrared characteristics) (3D point cloud features) (Acoustic characteristics), among which and For spatial resolution, This represents the number of feature channels for each modality. For each modality, a multi-head self-attention operation is performed independently: ; Among them, query vector Key vector AND value vector All are obtained from input features through linear transformation. Given the dimension of the key vector, the first-level self-attention modeling enables each modality to fully exploit its own long-range spatial dependencies, outputting enhanced feature representations for each modality. , , , .

[0032] Level 2: Intra-group cross-modal attention interaction. In S3, the visible light mode and the 3D structured light mode are grouped into the geometry group, and the near-infrared mode and the acoustic mode are grouped into the component group. Intra-group cross-modal attention adopts a cross-QKV bidirectional interaction mechanism.

[0033] Let the input of the geometry group be and The input for the component group is and Taking the geometry group as an example, the calculation process for cross-QKV bidirectional interaction is as follows: Cross-attention from visible light modes to 3D structured light modes: ; in, , , It is a learnable projection matrix.

[0034] Cross-attention from 3D structured light modes to visible light modes: ; The output of the geometry group is a fusion of bidirectional cross-attention: ; in, It is a feedforward neural network. For feature concatenation operations, the processing of component groups is the same, and the output is... .

[0035] Level 3: Inter-group collaborative attention fusion, in the output of the geometric group Output of component groups Cross-group collaborative attention computation is performed between them. This level of attention is responsible for the interaction and calibration of global information, enabling geometric features and component features to achieve consistency at a higher semantic level. ; Final output It is a four-modal collaborative global feature that integrates local details and global semantic information of each modality.

[0036] like Figure 3 As shown, in this embodiment, the cross-modal prediction network is implemented using a multilayer perceptron (MLP), comprising 12 prediction networks that model 12 cross-modal prediction relationships: visible light NIR, NIR Visible light, visible light 3D, 3D Visible light, visible light Acoustics, acoustics Visible light, NIR 3D, 3D NIR, NIR Acoustics, acoustics NIR, 3D Acoustics, acoustics 3D.

[0037] To reduce computational cost, these 12 cross-modal prediction networks can share a backbone network, first mapping the features of each modality to a unified space through a shared encoder: ; Then, each prediction head predicts features of other modalities based on the shared representation: ; in, For modality Predictive features, For modality Input features (after self-attention enhancement).

[0038] Prediction error is defined as the mean squared error between the predicted features and the true features: ; in, It is the Frobenius norm. , , Let be the number of feature channels in the spatial and modal j features.

[0039] Inconsistency measurement The calculation formula is: ; in, For cross-modal prediction error The weighting coefficients satisfy The weighting coefficients can be configured according to the detection requirements of different defect types. For example, for sealing defect detection, the weight of acoustic correlation prediction error can be appropriately increased.

[0040] Final anomaly score The calculation formula is: ; in, The absolute anomaly score is the deviation of each modality feature from the normal sample database. and For the weighting coefficients, satisfying , The calculation method is as follows: ; in, The database consists of normal samples (which can be modeled using a Gaussian distribution in the feature space or a nearest neighbor method based on a memory bank). Euclidean distance is used as the feature distance metric.

[0041] when If the threshold is exceeded, the sample is determined to have defects.

[0042] like Figure 4 As shown, when cross-modal inconsistency measures Exceeding the first preset threshold When this occurs, it indicates that the system has detected a high level of cross-modal inconsistency. To avoid false alarms caused by single-modal noise and to ensure that real defects are not missed, the system triggers a secondary fine analysis process.

[0043] Secondary refined analysis employs higher-resolution local feature extraction. Specifically, in In the event of an anomaly, the system performs sliding window sampling within the spatial neighborhood of the anomaly region, reducing the window size to a fraction of the original size. More fine-grained local features are extracted. These local features are then fed back into a three-level grouped collaborative attention network for inference, yielding anomaly scores at the local scale. .

[0044] Meanwhile, the secondary analysis introduces an attention visualization mechanism to calculate the contribution of each modality to the final decision: ; Gradient backpropagation is performed to the input layer to obtain the normalized contribution score of each modality feature. The modality with the highest contribution score is marked as the main perceptual modality of the anomaly region and used for subsequent defect type inference.

[0045] If the local anomaly score of the secondary fine analysis Still exceeds the second preset threshold ( If the defect is identified as a real defect, the defect location and type information will be output; if... If the noise is not detected, it is determined to be single-mode noise interference, and the detection result is corrected.

[0046] In this embodiment, a modal dynamic routing step is included before step S3. The dependence of different defect types on each mode varies significantly. For example, surface scratches mainly rely on texture analysis from visible light images, sealing defects mainly rely on the combined judgment of acoustic signals and near-infrared images, and internal foreign objects mainly rely on the penetration characteristics of near-infrared images and depth anomalies in 3D point clouds. Therefore, dynamically adjusting the modal participation weights according to the expected defect type distribution can improve the targeted sensitivity of the detection system.

[0047] The dynamic modal routing module is a learnable weight configuration vector. ,in Corresponding to four modes, The calculation process for the routing weights to determine the number of channels for feature fusion is as follows: For fusion features The first in For each channel, the channel statistics for each modality are first obtained through global average pooling: ; Then, each mode is output through a lightweight gating network at the 1st... Channel weights: ; in, It is the sigmoid activation function. and For the parameters of the gating network, Let c be the gate weight vector of the c-th channel. This is the bias term, initialized with a normal distribution before training.

[0048] Finally, the modal weights and fused features are adaptively weighted by multiplying them channel by channel: ; Weighted configuration vector The system stores the mapping relationship between defect types and modal contributions. When the system receives prior information about defect types, such as high-probability defect types determined based on production batches or product models, it can query this information. Obtain the corresponding modal weight configuration to achieve targeted detection enhancement.

[0049] The routing network, the three-level group collaborative attention fusion network, and the cross-modal prediction network together form a complete detection model, which is jointly trained in an end-to-end manner.

[0050] The training loss function consists of three components: 1) Defect classification loss The binary cross-entropy loss is used to measure the difference between the predicted defect labels and the true labels. ; in, The number of training samples, The labels are real (0 indicates normal, 1 indicates defect). To predict probabilities, the output of the classification head is: weighted fused features. Global average pooling and fully connected layer mapping are performed sequentially, followed by a sigmoid activation function for output. ,in and For classification header parameters, This is the sigmoid function.

[0051] 2) Consistency regularization loss Encourage routing weights to be differentiated across different defect types, while maintaining consistent weights for defects of the same type. ; in, For variance, This is the term representing the largest difference between the means. and The routing weights for two samples of the same defect type (used to measure consistency within the same category).

[0052] 3) Cross-modal prediction loss The sum of the mean square errors of each cross-modal prediction network: ; The total loss function is: ; in, and This is the balance coefficient.

[0053] During training, the Adam optimizer is used for gradient descent, and the normal sample database is updated every few epochs to adapt to changes in the production environment.

[0054] In continuous production, the distribution of defect types varies significantly across different batches of products, and there is a clear physical correspondence between the dependence of different defect types on each mode: surface scratches and printing defects mainly rely on texture analysis of visible light images and depth anomalies of 3D point clouds, corresponding to the geometric mode; seal integrity defects mainly rely on acoustic resonance signals and the penetration characteristics of near-infrared spectra, corresponding to the component mode; internal foreign object defects require joint judgment of near-infrared penetration detection and 3D point cloud depth, requiring the collaborative participation of both modes. When a batch of products is dominated by seal defects, the detection value of the component mode is far higher than that of the geometric mode. In this case, if fusion is still performed on all four modes, it not only wastes computational resources, but the characteristic noise of the geometric mode may also interfere with the interpretation results of the component mode.

[0055] In this embodiment, adaptive modal scheduling is introduced after step S1 and before step S3, forming a two-level scheduling system: the first level is adaptive modal scheduling (batch level), which determines which modalities in the current batch participate in subsequent processing based on the defect type distribution; the second level is dynamic modal routing (sample level), which performs fine-grained weight allocation on a channel-by-channel basis for activated modalities. The execution order of the two-level scheduling is as follows: the adaptive modal scheduler first outputs the modal activation mode; only data of activated modalities enters the subsequent feature extraction and routing network, while data of inactive modalities are bypassed in the current batch.

[0056] In the first-level adaptive modal scheduling, the correspondence between different defect types and the four modal activation modes is shown in the table below: Surface scratches, printing defects Visible light texture + 3D depth anomaly Geometric group activation Seal integrity defects Acoustic resonance + near-infrared penetration Ingredient group activation Internal foreign objects Near-infrared penetration + 3D depth Full-modal activation abnormal material composition Near-infrared spectroscopy Ingredient group activation ; When the proportion of a certain type of defect in the expected defect distribution exceeds a preset proportion threshold, the scheduler prioritizes the activation mode corresponding to that defect type; when the proportions of multiple defect types are similar and none of them are significantly dominant, the scheduler selects full-modal activation to retain complete perception capability.

[0057] The scheduler uses a deep Q-network (DQN) to map states to actions. The scheduler's state vector is... It is composed of the following three parts: ; in, This is the expected defect type distribution vector for the current batch, with four components corresponding to the probabilities of surface defects, sealing defects, internal foreign objects, and material anomalies, respectively, satisfying the following conditions: ; The real-time signal-to-noise ratio (SNR) estimate for the four modes is calculated by the ratio of the mean of the original signal of each mode to the standard deviation of the noise. This value represents the system's resource utilization rate, calculated as the ratio of the current GPU memory utilization rate to a preset upper limit. (State Dimension) .

[0058] The action space is a discrete selection of four modal activation modes: These correspond to full-modal activation, geometry group activation, component group activation, and single-modal activation, respectively.

[0059] The Q-network uses a three-layer fully connected network, with the input being a state vector. The output is the Q-values ​​corresponding to the four actions: ; in, Map the input dimension from 9 to 64. Mapping from 64 to 32, Mapping from 32 to 4 (corresponding to the Q values ​​of the four actions). To provide all trainable parameters for the Q-network, the scheduler sets the current state at the beginning of each batch. Input the Q-network and select the action with the largest Q-value as the modal activation mode for this batch: ; Calculate the reward value after each production batch is completed: ; in, This represents the defect detection rate for the current batch. False alarm rate This represents the average detection delay per sample. , , The weighting coefficients are determined based on the following principles: detection rate, being the core indicator of the detection system, receives the highest weight; false alarm rate, leading to unnecessary re-inspection costs in subsequent processes, receives a medium weight; and detection delay, with its relatively indirect impact on production line cycle time, receives the lowest weight. Typical values ​​are... , , ,satisfy In scenarios where production line cycle time is critical, the speed can be appropriately increased. The value of .

[0060] Before formal deployment, the scheduler is pre-trained through offline simulation. During the simulation, an experience pool is built using the defect distribution records of historical production batches and the corresponding detection effect data. The Q network is pre-trained using the standard DQN training process to enable it to have basic scheduling strategies.

[0061] During the online operation phase, the scheduler updates the policy in batches. After each batch is completed, the system obtains the detection performance feedback for that batch and constructs the transition tuple. And store it in the experience replay buffer. Each accumulated... After each batch of transferred tuples ( The typical value is 10). A batch of data is randomly sampled from the buffer, and the temporal difference objective is calculated: ; in, This is the discount factor (typically 0.95). For the parameters of the target Q network (every...) Each batch of software updates was released via the online Q network. , (The typical value of is 0.005), and the loss function of the Q-network is: ; in, To determine the sampling batch size, the loss function is minimized using gradient descent, and the Q-network parameters are updated. .

[0062] For samples detected as defective, subsequent manual re-inspection or automatic sorting equipment provides confirmation information. Samples confirmed to be defective are counted in the detection count, while those confirmed as false alarms are counted in the false alarm count. Samples judged as normal are sent for manual sampling at a preset sampling ratio (typically 5%). Samples missed during this sampling are counted in the missed detection count to correct the detection rate. The defect detection rate is calculated as follows: The false alarm rate is calculated as follows: The detection delay is obtained directly from the system log.

[0063] To prevent missed detections due to sudden policy changes during online learning, the scheduler's online updates employ the following constraint mechanism: 1) Gradient clipping: Each time the parameters are updated, the gradient norm is clipped to no more than a preset threshold. (Typical value is 1.0), limiting the scope of impact of a single update; 2) Safety lower limit of detection rate: When continuous batches ( The typical value is 3), and the detection rate is lower than the preset safety lower limit. At this time, the current Q network parameters are locked, and the system is forced to switch to full-modal activation mode. At the same time, a manual alarm is triggered, and online learning can only be resumed after manual confirmation. 3) Confidence constraint: When the expected defect type distribution vector maximum component ( When the typical value of is 0.4, it indicates that the defect distribution is uncertain. The scheduler skips the Q-network decision and directly selects the full-modal activation mode to avoid making high-risk modal bypass decisions when information is insufficient.

[0064] After the modal activation mode is determined, the three-level grouped collaborative attention fusion network in step S3 executes different fusion paths according to the activation mode: Full-modal activation: S3 normally performs a complete three-level fusion. The first level is independent self-attention for each modality. The second level is cross-modal attention within the geometry group and the component group. The third level is collaborative attention fusion between groups. Geometric group activation: Only visible light and 3D structured light modes participate in fusion. The first level performs intramodal self-attention between the two modes; the second level performs bidirectional cross-modal attention within the geometric group (skipping component groups); the third level is skipped (since there is only one group, there is no need for inter-group fusion), and the output of the geometric group is directly used as the final feature. Component group activation: Only near-infrared and acoustic modes participate in fusion. The first level performs intra-modal self-attention for both modes; the second level performs bidirectional cross-modal attention within the component group (skipping the geometry group); the third level is skipped, and the component group output is directly used as the final feature. Single-modal activation: Only the visible light mode is involved. The first level performs self-attention within the visible light mode; the second and third levels are skipped, and the self-attention output is directly used as the final feature.

[0065] The switching of the above-mentioned fusion path is achieved through conditional branching, without the need to modify the network structure or reload parameters.

[0066] In the transition between the second-level dynamic routing and adaptive modal scheduling, after the adaptive modal scheduler determines the modal activation mode, only the features of the activated modes enter the dynamic modal routing network of Example 6. The routing network performs channel-by-channel weight allocation for the activated modes, and its calculation method is consistent with Example 6, the only difference being that the input dimension of the routing network changes with the number of activated modes: when all four modes are activated, the routing weight vector dimension is... When only the geometry group or component group is activated, the dimension of the routing weight vector is reduced to [value missing]. When only the visible light mode is activated, the routing network is skipped, and the modal features are directly processed.

[0067] The overall data flow of the two-level scheduling is as follows: spatiotemporally aligned data output in step S2 → reinforcement learning scheduler selects modal activation mode according to defect distribution → activated modal data undergoes lightweight feature extraction → dynamic modal routing outputs channel-wise weighted features → step S3 is executed according to the fusion path corresponding to the activation mode.

[0068] When a significant change in the defect distribution is detected, such as when the JS divergence based on sliding window statistics exceeds a preset threshold, the reinforcement learning scheduler re-evaluates the modal activation patterns after the current batch ends, enabling the system to adapt to the new defect distribution.

[0069] During long-term operation, sensors may experience performance degradation or sudden failures due to lens contamination, light source attenuation, mechanical loosening, etc. Timely detection and location of sensor faults are crucial to ensuring system reliability. This embodiment designs a fault self-diagnosis module based on cross-modal consistency statistics.

[0070] The triggering conditions for fault diagnosis are based on the statistical properties of inconsistency measures, letting For the first A measure of cross-modal inconsistency for each test sample, in continuous Calculate on each sample mean with standard deviation When an abnormally high level of inconsistency is detected, it is determined that a sensor malfunction may exist. ; in, This is the fault diagnosis threshold. The confidence coefficient (typically taking the value of ) ).when Continue to exceed When there are consecutive samples, the fault diagnosis process is triggered.

[0071] Fault localization employs a contribution analysis method, calculating the inconsistency measure for each mode pair for each detection sample. Contributions: ; The mode with the highest contribution is marked as a suspected fault mode. To avoid interference from occasional noise, the system uses a sliding window to count the frequency of each mode being marked as suspected. The frequency of suspicious samples exceeded the threshold. When this occurs, output a fault alarm for that mode.

[0072] Fault alarm information includes: fault mode identifier, fault severity assessment (minor degradation / serious fault), and confidence score. The system can automatically trigger targeted sensor self-test processes (such as performing internal calibration verification) and decide whether to switch to a degraded configuration based on the fault level.

[0073] In this embodiment, a four-level degradation operation mode is defined, ranging from full functionality to the lowest functionality, to cope with sensor failure or extreme operating conditions: Level L0 (Full Functionality Mode): All four modalities are online, the three-level group collaborative attention network is working normally, detection performance is optimal, and the anomaly scoring threshold is the original configuration value. .

[0074] Level L1 (Mild Transition Mode): When a slight performance degradation in a single modality is detected but no fault alarm has been triggered, Level L1 configuration is enabled. In this mode, the system automatically reduces the weight of the affected modality while increasing the weight of other normal modalities. The anomaly scoring threshold is tightened. This is done to slightly improve sensitivity to compensate for the loss of single-mode information.

[0075] Level L2 (Moderate Degradation Mode): Level L2 configuration is activated when a single sensor fails or two sensors simultaneously exhibit minor anomalies. In this mode, the faulty mode is completely bypassed, and the fusion network makes decisions based solely on the outputs of the remaining two normal mode groups. To maintain detection sensitivity, the threshold is tightened. At the same time, the criteria for judging minor defects should be appropriately relaxed to reduce missed detections.

[0076] Level L3 (Minimum Functional Mode): Level L3 configuration is activated when two or more sensors experience serious failures. In this mode, the system retains only the visible light camera as the sole detection modality, switching to a lightweight single-modal detection model (such as a lightweight network based on MobileNet). The anomaly scoring threshold is tightened. This is to minimize missed detections and report only high-confidence defect samples.

[0077] The switching of downgraded configurations is coordinated by the fault self-diagnosis module. When the fault diagnosis module outputs an alarm, the system assesses the fault level and automatically selects the corresponding downgrade level. During the downgrade process, the system continuously monitors the recovery status of the faulty sensors. Once the fault is resolved, it will attempt to restore to a higher functional level, achieving flexible control.

[0078] In this embodiment, using 3D point cloud as the anchor point for spatial registration in step S2 can fully leverage the accuracy advantages of 3D structured light sensors in spatial measurement and provide a unified spatial reference framework for other modalities.

[0079] The core steps of space registration include: 1) Establish Region of Common Interest (ROI): Based on the coverage of the 3D point cloud, define the boundaries of the packaging region of interest for the detection task. The 3D point cloud within this region serves as a spatial reference, and data from other modalities are mapped to this region.

[0080] 2) Perspective projection parameter calibration: For visible light cameras and near-infrared cameras, the camera's intrinsic parameters are accurately calibrated using a calibration board (such as a checkerboard calibration board). distortion coefficients and external parameters relative to the 3D sensor coordinate system The calibration process is performed during system installation and is periodically rechecked to address any mechanical loosening.

[0081] 3) Spatial Correlation of Acoustic Signals: The spatial position of the acoustic sensor array is also calibrated in the 3D sensor coordinate system. The spatialization of the acoustic signals uses a beamforming algorithm to calculate the acoustic energy distribution at each spatial grid point. The resolution of the spatial grid is consistent with the resolution of the 3D point cloud (e.g., ...). ).

[0082] 4) Coordinate transformation and interpolation: After the data of each modality are transformed to a unified coordinate system through the above calibration parameters, interpolation is required to obtain spatially aligned feature maps due to different spatial sampling rates. This invention uses bilinear interpolation (for 2D image modalities) or trilinear interpolation (for 3D voxel modalities) to resample the features of each modality on a unified spatial grid.

[0083] Acoustic sensor array adopts The configuration (eight MEMS microphones arranged in a ring, plus one central reference sensor) enables this array configuration to achieve The directional sound beam scan completely covers the circumferentially sealed area of ​​the packaging, and the microphone unit's sensitivity is approximately... The signal-to-noise ratio is higher than Sampling rate This meets the accuracy requirements for sealing testing.

[0084] After the above spatial registration process, the four-modal data achieved pixel-level spatial alignment in a unified three-dimensional spatial coordinate system, providing spatial consistency assurance for subsequent group collaborative attention fusion.

[0085] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention. The scope of protection claimed by the appended claims and their equivalents is defined.

Claims

1. A method for detecting defects in food packaging based on deep learning, characterized in that, include: S1. Multimodal data of food packaging are collected simultaneously using a visible light camera, a near-infrared (NIR) camera, a 3D structured light sensor, and an acoustic sensor to obtain visible light images, NIR images, 3D point cloud data, and acoustic signals, respectively. S2. The data of each modality are time-stamped and aligned. Using the 3D point cloud data as a spatial reference, the visible light image and NIR image are mapped to the unified coordinate system of the 3D point cloud through perspective projection. The acoustic signal is converted into a spatially distributed detection signal through spatial correlation with the 3D point cloud to obtain the spatiotemporally aligned four-modal data. S3. Input the spatiotemporally aligned four-modal data into a three-level grouped collaborative attention fusion network for feature fusion. The three-level grouped collaborative attention fusion network performs the following in sequence: first-level intramodal self-attention modeling, and each modality independently extracts intramodal long-range dependency features. The second level of intra-group cross-modal attention interaction involves grouping visible light modes and 3D structured light modes into a geometry group, and NIR modes and acoustic modes into a component group, and performing cross-modal attention interaction within each group. The third level of inter-group collaborative attention fusion performs cross-group attention calculation between the outputs of the geometry group and the component group to generate four-modal collaborative global features; S4. Based on the four-modal collaborative global features, calculate the deviation between each modality and the normal sample database as the absolute anomaly score, and calculate the weighted sum of the prediction errors of each cross-modality as the cross-modal inconsistency measure. The absolute anomaly score and the cross-modal inconsistency measure are weighted and fused to generate the final anomaly score. S5. Output the defect detection results based on the final anomaly score.

2. The food packaging defect detection method based on deep learning according to claim 1, characterized in that: In the second-level intra-group cross-modal attention interaction, the geometry group and the component group respectively conduct bidirectional feature interaction through the cross-query-key-value mechanism: the query vector of one modality feature and the key vector and value vector of the other modality feature are used to calculate the positive cross-modal attention, while the query end and the key-value end are exchanged to calculate the reverse cross-modal attention; In the third-level inter-group collaborative attention fusion, geometric group fusion features and component group fusion features are used as query ends, respectively. Bidirectional cross-group attention calculation is performed on the key value end, and the attention outputs between the two groups are concatenated and linearly projected to generate four-modal collaborative global features.

3. The food packaging defect detection method based on deep learning according to claim 1, characterized in that: The cross-modal prediction error in step S4 is calculated as follows: For each of the four modes—visible light, NIR, 3D structured light, and acoustic—in a pairwise ordered combination, a cross-modal prediction network is trained. Each cross-modal prediction network is trained using a modal... The features are used as inputs, and the mode is predicted by a multilayer perceptron. The mean squared error between the predicted features and the actual extracted features is calculated as the cross-modal prediction error. ; Cross-modal inconsistency measurement The calculation formula is: ; in, For modality For modes The prediction weights are dynamically adjusted based on the signal-to-noise ratio and signal strength of each modality in the current detection scenario; the final anomaly score... The calculation formula is: ; in, For absolute anomaly scoring, and These are learnable fusion weight parameters.

4. The food packaging defect detection method based on deep learning according to claim 3, characterized in that: When the absolute anomaly score A is lower than the preset defect judgment threshold and the cross-modal inconsistency measurement is... When the inconsistency exceeds a preset inconsistency threshold, a second, more detailed analysis is triggered. Positioning The cross-modal prediction error term with the largest absolute contribution The high-resolution Region of Interest (ROI) corresponding to the modality pair is extracted from the original acquired data. The feature extraction and cross-modal prediction errors are recalculated on this ROI, and the absolute anomaly score is updated. and cross-modal inconsistency measurement The final anomaly score was then recalculated. .

5. The food packaging defect detection method based on deep learning according to claim 1, characterized in that: Before step S3, there is also a modal dynamic routing step: the spatiotemporally aligned four-modal data is extracted with lightweight features and then input into the defect type prediction head to obtain the defect type prediction probability distribution of the current sample; The probability distribution of defect type prediction is input into a learnable routing network. The routing network outputs a weight configuration vector for each modality. The weight configuration vector is multiplied channel by channel with the feature map of each modality to obtain the weighted modal features. The weighted modal features are then input into a three-level grouped collaborative attention fusion network.

6. The food packaging defect detection method based on deep learning according to claim 5, characterized in that: The routing network is trained in the following way: using the four-modal features of historical detection samples and labeled defect types as training data, the loss function of the routing network includes defect type prediction classification loss and detection accuracy reward of modality weight allocation. The detection accuracy reward of modality weight allocation is calculated based on the defect detection accuracy output by the fusion network after the weighted four-modal features are processed. The defect type prediction head, the routing network and the three-level grouped collaborative attention fusion network are jointly trained end-to-end.

7. The food packaging defect detection method based on deep learning according to claim 1, characterized in that: wherein, An adaptive mode scheduling step is also included between step S1 and step S3: The expected defect type distribution vector for the current batch Real-time signal-to-noise ratio of each mode and system computing resource utilization Concatenate into a state vector Input a deep Q-network, which outputs Q-values ​​corresponding to four actions, and select the action with the largest Q-value as the modal activation mode. The modal activation mode includes one of the following: full four-modal activation, geometric group modal activation, component group modal activation, or single-modal activation; In step S3, the corresponding fusion path is executed according to the activation mode: when the full modality is activated, the complete three-level fusion is executed; when the geometry group is activated, only the first-level intramodal self-attention and the second-level cross-modal attention within the geometry group are executed; when the component group is activated, only the first-level intramodal self-attention and the second-level cross-modal attention within the component group are executed; when the single modality is activated, only the first-level intramodal self-attention is executed. After each batch is completed, the reward value is calculated based on the defect detection rate, false alarm rate, and detection delay. ,in , , The weighting coefficients and The parameters of the deep Q-network are updated using the reward value through temporal difference learning; When the maximum component of the expected defect type distribution vector is lower than the confidence threshold, skip the deep Q-network decision and directly select the full-modal activation mode.

8. A food packaging defect detection method based on deep learning according to claim 7, Its features are: Among them, cross-modal inconsistency metrics are calculated in real time during the detection process. ,when A modal fault is determined to exist when the dynamic diagnostic threshold is exceeded. The dynamic diagnostic threshold is the mean of cross-modal inconsistency measures in historical normal operation data. Plus Double standard deviation ,Right now ,in This is a preset multiplier factor; The fault mode localization method is as follows: analyze the prediction error terms of each cross-modal mode. right The contribution, if it includes modality All as prediction sources or prediction targets The sum of the contributions of each item accounts for the total If the proportion exceeds a preset proportion threshold, then the mode is determined. It is a fault mode; When a modal fault is detected, the deep Q network scheduler automatically switches to a degraded modal activation mode that does not contain the faulty mode.

9. The food packaging defect detection method based on deep learning according to claim 8, characterized in that: Four-level degradation operation modes are defined: Level L0 is all four modes activated, Level L1 is three modes activated after removing the acoustic mode, Level L2 is only the visible light mode and NIR mode retained, and Level L3 is only the visible light mode retained. Each degradation mode corresponds to a set of adjusted defect judgment threshold parameters in step S5. The lower the degradation level, the tighter the defect judgment threshold. Once the fault mode recovers, the parameters of the fusion network are updated online using the cross-modal corresponding data accumulated during normal operation.

10. The method for detecting defects in food packaging based on deep learning according to claim 1, characterized in that: The specific method for using 3D point cloud data as a spatial reference in step S2 is as follows: a unified coordinate system is constructed using the 3D point cloud collected by the 3D structured light sensor. The pixel coordinates of the visible light image and the NIR image are mapped to the 3D point cloud coordinate system through perspective projection using the pre-calibrated visible light camera extrinsic matrix and NIR camera extrinsic matrix to obtain visible light feature map and NIR feature map that are spatially aligned with the 3D point cloud. The specific method for converting the acoustic signal into a spatially distributed detection signal is as follows: an acoustic sensor array is arranged on the side of the conveyor belt to collect multi-channel time-domain acoustic signals. Combined with the position coordinates of the packaged object in the 3D point cloud, the acoustic signals of each channel are mapped to the corresponding spatial positions in the 3D point cloud coordinate system to generate a spatially distributed acoustic detection feature vector.

Citation Information

Patent Citations

  • Fan blade defect detection system and method based on multi-mode perception

    CN119180793A

  • Multi-mode printed circuit board defect detection method and device based on deep learning and storage medium

    CN119540241A