Packaging defect detection method based on machine vision
By combining deep probabilistic prediction networks and the finite element method, a global full-focus image is generated and a packaging physical simulation model is constructed, solving the problem of simultaneously acquiring high-resolution images and three-dimensional morphology in existing technologies, and realizing high-precision packaging defect detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANJING DEQIAN INFORMATION TECH CO LTD
- Filing Date
- 2026-04-20
- Publication Date
- 2026-05-15
AI Technical Summary
Existing machine vision-based packaging defect detection technologies cannot simultaneously acquire high-resolution, clear global images and accurate three-dimensional shapes, and are difficult to adapt to the dynamic deformation of flexible packaging during transportation, resulting in high false alarm and false negative rates, and failing to effectively distinguish between normal deformation and real defects.
A deep probabilistic prediction network is used to generate global full-focus images. A physical simulation model of packaging is constructed using the finite element method. Defect detection is performed by combining a multimodal feature interaction mechanism, and a visualized quality traceability report is generated.
It achieves high-quality synchronization of two-dimensional texture information and three-dimensional geometric information, effectively distinguishing between normal deformation and real defects, and generating high-precision defect detection results.
Smart Images

Figure CN122048952A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to a method for detecting packaging defects based on machine vision. Background Technology
[0002] In the wave of intelligent manufacturing and industrial automation, machine vision, as a core technology for achieving high-precision, non-contact online inspection, has been widely applied in the quality control of the packaging industry. Traditional packaging defect detection methods mainly rely on two-dimensional image analysis techniques, such as template matching, gray-level co-occurrence matrix texture analysis, edge detection, and traditional machine learning classifiers. In recent years, with breakthroughs in deep learning technology, algorithms represented by convolutional neural networks have improved the defect recognition rate in specific scenarios. Faced with the increasing demand in the packaging industry for high flexibility, high reflectivity, complex deformation, and minute defect detection, the existing technology system shows obvious limitations. Especially when dealing with flexible packaging in the food and daily chemical industries (such as puffed food bags and shampoo stand-up pouches), their inherent non-rigid deformation, surface high gloss reflection, local occlusion, and printing texture interference result in high false alarm and false negative rates for two-dimensional image-based detection methods, making it difficult to meet the stringent requirements of modern production lines for detection accuracy and stability.
[0003] Existing machine vision-based packaging defect detection technologies suffer from two main shortcomings: First, at the imaging level, most systems employ traditional area or line scan cameras, which cannot effectively overcome the problems of local defocusing, perspective distortion, and specular interference caused by the random changes in the posture of flexible packaging. Although some studies have introduced multi-view, multi-source, or structured light schemes to acquire 3D information, these systems are complex, costly, and difficult to simultaneously acquire high-resolution, globally clear images and accurate 3D topographic data on high-speed production lines, resulting in a weak foundation for subsequent processing. Second, at the defect discrimination level, existing methods generally adopt a static paradigm of "image acquisition - comparison with standard templates," which cannot adapt to the dynamic and nonlinear deformation of flexible packaging during transportation and cannot effectively distinguish the essential differences between image differences caused by normal physical deformation and real defects (such as scratches, stains, and damage). Summary of the Invention
[0004] In view of the aforementioned existing problems, the present invention is proposed.
[0005] Therefore, this invention provides a machine vision-based method for detecting packaging defects, which solves the problem that existing technologies cannot simultaneously acquire high-resolution, clear global images and accurate three-dimensional shapes.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: This invention provides a machine vision-based method for detecting packaging defects. The method includes: acquiring raw light field data of the flexible packaging to be inspected; generating a global full-focus image to eliminate local defocus using a depth probability prediction network; obtaining a high-resolution depth map characterizing the microscopic undulations of the packaging surface from the raw light field data; analyzing the high-resolution depth map and extracting feature parameter vectors describing the current three-dimensional deformation state of the packaging; constructing a packaging physical simulation model using the finite element method with the feature parameter vectors as boundary conditions; performing a fast static equilibrium solution on the packaging physical simulation model to generate a defect-free ideal dynamic template image; performing high-order continuous deformation registration between the global full-focus image and the ideal dynamic template image; performing pixel-by-pixel difference calculation and adaptive threshold filtering in the color space to generate a pure difference map that excludes normal deformation interference; performing pixel-level feature alignment and fusion between the pure difference map and the high-resolution depth map through a multimodal feature interaction mechanism to obtain the defect segmentation mask, category, and severity information; performing quality grading decisions and sorting control on the defect segmentation mask, category, and severity information; and associating the feature parameter vectors to generate a visualized quality traceability report.
[0007] As a preferred embodiment of the machine vision-based packaging defect detection method of the present invention, the original light field data includes spatial location information, angular direction information, and light intensity information.
[0008] As a preferred embodiment of the machine vision-based packaging defect detection method of the present invention, the steps of acquiring the original light field data of the flexible packaging to be detected and generating a global full-focus image that eliminates local defocusing using a depth probability prediction network are as follows: Raw light field data is acquired using a light field camera, and an initial depth probability distribution map and uncertainty heatmap are obtained through a depth probability prediction network. A set of non-uniformly sampled refocusing image patches is then generated. Interactively fuse the refocused image patch set with the initial depth probability distribution map to generate aggregated features and focus confidence for each pixel location; Using aggregation features and focus confidence as conditions, a global full-focus image with local defocusing is obtained by optimizing the composite loss function.
[0009] As a preferred embodiment of the machine vision-based packaging defect detection method of the present invention, the specific steps for obtaining a high-resolution depth map characterizing the microscopic undulations of the packaging surface from the original light field data are as follows: Based on the set of global full-focus images and refocused image patches, an initial matching cost volume is constructed, and the spatial propagation of the initial matching cost volume is guided by the focus confidence to obtain the initial depth estimate. Calculate the optimal sub-pixel level depth correction based on the initial depth estimate; Global optimization by multi-cue fusion of the optimal sub-pixel level depth correction, aggregated features, and focus confidence yields a high-resolution depth map characterizing the micro-undulations of the packaging surface.
[0010] As a preferred embodiment of the machine vision-based packaging defect detection method of the present invention, the specific steps for parsing the high-resolution depth map and extracting the feature parameter vector describing the current three-dimensional deformation state of the packaging are as follows: High-resolution depth maps are converted into 3D point clouds, and a physical perception map structure is constructed by combining geometric continuity constraints. The physical perception graph structure is input into a multi-layer deformable graph convolutional network, and the node deformation features are updated through a deformation perception attention mechanism. Adaptive graph pooling and physical mode decomposition are performed on the updated node deformation features to generate the feature parameter vector of the current packaging 3D deformation state.
[0011] As a preferred embodiment of the machine vision-based packaging defect detection method of the present invention, the specific steps for constructing a packaging physical simulation model using the finite element method with feature parameter vectors as boundary conditions are as follows: Physical parameter inversion is performed on the characteristic parameter vector to obtain the material property field and boundary load conditions; The material property field and boundary load conditions are assigned to the reference geometric model of the packaging, and the nonlinear constitutive relations and constraints of the reference geometric model are defined to form a parameterized finite element description. A physical simulation model of packaging is constructed based on parametric finite element description.
[0012] As a preferred embodiment of the machine vision-based packaging defect detection method of the present invention, the specific steps for rapidly solving the static equilibrium of the packaging physical simulation model to generate a defect-free ideal dynamic template image are as follows. A fast static equilibrium solution is performed on the packaging physical simulation model to obtain a high-fidelity deformation displacement field; The high-fidelity deformation displacement field is compressed using the nonlinear eigenorthogonal decomposition method to obtain low-dimensional eigencoordinates characterizing the current deformation state. Based on low-dimensional intrinsic coordinates and pre-calibrated detection viewpoint conditions, a defect-free ideal dynamic template image is generated through differentiable rendering.
[0013] As a preferred embodiment of the machine vision-based packaging defect detection method of the present invention, the specific steps for generating a clean difference map that excludes interference from normal deformation are as follows: High-order continuous deformation registration is performed on the global full-focus image and the ideal dynamic template image to generate a geometrically aligned deformable template image; Based on deformable template images, global full-focus images, high-resolution depth maps, and focus confidence, a consistency weighted threshold is calculated using a physical consistency verification function and then filtered to obtain a preliminary difference mask. A Markov random field optimization with multi-source information constraints is performed on the initial difference mask to generate a pure difference map that excludes interference from normal deformation.
[0014] As a preferred embodiment of the machine vision-based packaging defect detection method of the present invention, the specific steps for obtaining the segmentation mask, category, and severity information of the defect are as follows: The pure difference map and the high-resolution depth map are input into the multimodal feature interaction mechanism to perform pixel-level feature alignment and fusion, generating multi-scale fused features; Multi-scale fused features are input into a hierarchical decoder, and preliminary segmentation masks, categories, and severity information of defects are obtained through multi-task joint inference. Spatial context refinement is performed on the initial segmentation mask, category, and severity information to obtain the segmentation mask, category, and severity information of defects.
[0015] As a preferred embodiment of the machine vision-based packaging defect detection method of the present invention, the specific steps for generating a visual quality traceability report are as follows: Based on the segmentation mask, category, severity information and feature parameter vector of the defect, an intelligent decision state representation is constructed; The intelligent decision-making state representation is input into a deep reinforcement learning strategy to execute multi-objective optimization decisions and generate sorting control instructions. The system executes sorting control commands and collects real-time production line status data during the execution process. After fusing the real-time production line status data with feature parameter vectors, it performs causal inference and counterfactual analysis to generate a visualized quality traceability report.
[0016] The beneficial effects of this invention are as follows: by generating a global full-focus image through a depth probability prediction network, strict synchronization and pixel-level alignment of two-dimensional texture information and three-dimensional geometric information in time and space are ensured, realizing intelligent expansion of imaging depth of field and providing a high-quality two-dimensional information foundation; through deformation compensation and difference extraction by physical simulation and dynamic template generation, the abstraction from apparent geometry to intrinsic physical parameters is realized, providing accurate input conditions for physical simulation. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart of a machine vision-based packaging defect detection method.
[0019] Figure 2 A flowchart for generating a high-resolution depth map.
[0020] Figure 3 A flowchart for constructing a physical simulation model and generating an ideal dynamic template image.
[0021] Figure 4 This is a flowchart for multimodal feature fusion and defect information acquisition. Detailed Implementation
[0022] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0023] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0024] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0025] Reference Figures 1-4 This is one embodiment of the present invention, which provides a machine vision-based method for detecting packaging defects, including the following steps: S1. Collect the original light field data of the flexible packaging to be tested, use a depth probability prediction network to generate a global full-focus image that eliminates local defocusing, and obtain a high-resolution depth map characterizing the micro-undulations of the packaging surface from the original light field data; the original light field data includes spatial location information, angular direction information and light intensity information.
[0026] Raw light field data is acquired using a light field camera, and an initial depth probability distribution map and uncertainty heatmap are obtained through a depth probability prediction network. A set of non-uniformly sampled refocusing image patches is then generated.
[0027] The specific process includes: during the acquisition process, the light field camera simultaneously records spatial location information, angular direction information, and light intensity information to form raw light field data; the raw light field data is input into a depth probability prediction network, which, based on the parallax information between multi-view geometric relationships and light field sub-aperture images, derives the probability of each pixel position under different depth assumptions through probabilistic modeling, thereby outputting an initial depth probability distribution map, and simultaneously acquiring the confidence level of each pixel depth estimate to generate an uncertainty heatmap; based on this, according to the position of high-probability depth layers in the initial depth probability distribution map and the guidance of low-uncertainty regions in the uncertainty heatmap, a selective refocusing operation is performed on the raw light field data, and non-uniform sampling is performed on different depth planes according to the probability density, generating a set of non-uniformly sampled refocused image patches.
[0028] Furthermore, the training process of the depth probability prediction network adopts a supervised learning approach, using labeled raw light field data for end-to-end training. Each sample contains raw light field data acquired by a light field camera and the corresponding ground truth depth map. During training, the raw light field data is input into the depth probability prediction network, which outputs an initial depth probability distribution map and an uncertainty heatmap through convolutional layers and a probabilistic inference structure. The loss function consists of two parts: a negative log-likelihood loss between the initial depth probability distribution map and the ground truth depth map, used to optimize the probabilistic accuracy of depth estimation; and an adaptive regularization term between the uncertainty heatmap and the depth estimation error, used to ensure that the uncertainty heatmap reflects the true prediction confidence. The parameters of the depth probability prediction network are updated through backpropagation, iterating repeatedly until the initial depth probability distribution map and the uncertainty heatmap achieve stable convergence on the validation set.
[0029] The refocused image patch set is interactively fused with the initial depth probability distribution map to generate aggregated features and focus confidence for each pixel location.
[0030] The specific process includes: selecting refocusing image blocks of the corresponding depth layer from the set of refocusing image blocks as candidate focus responses for the pixel based on the depth probability distribution of each pixel location in the initial depth probability distribution map; at each pixel location, weighting the local image features of the selected refocusing image block with the depth probability value in the initial depth probability distribution map, with the weight determined by the depth probability magnitude, to form the aggregated features of the pixel location; and simultaneously, obtaining the focus confidence of the pixel location based on the sharpness response intensity of the refocusing image block in the depth layer and the probability concentration of the corresponding depth in the initial depth probability distribution map.
[0031] Using aggregation features and focus confidence as conditions, a global full-focus image with local defocusing is obtained by optimizing the composite loss function.
[0032] The specific process includes optimizing the image reconstruction process using a composite loss function, conditioned on aggregated features and focus confidence. The composite loss function includes a focus region fidelity term and an out-of-focus region smoothing constraint term. The focus region fidelity term uses focus confidence as a weight to impose a stronger reconstruction consistency constraint on high-confidence pixels in the aggregated features. The out-of-focus region smoothing constraint term encourages gradual intensity changes in adjacent pixels in low focus confidence regions, thereby suppressing blur or artifacts caused by local defocusing. During the optimization process, the pixel values of the global image are adjusted to minimize the composite loss function, ultimately obtaining a global full-focus image with local defocusing eliminated.
[0033] It should be noted that a global full-focus image refers to an image that is clear in all depth regions through light field refocusing. Its purpose is to provide a high-quality texture reference without local defocusing interference for subsequent defect detection and 3D deformation analysis.
[0034] Based on a set of global all-focus images and refocused image patches, an initial matching cost volume is constructed, and the spatial propagation of the initial matching cost volume is guided by the focus confidence to obtain an initial depth estimate.
[0035] The specific process includes obtaining pixel-level similarity metrics between the global all-focus image and each refocused image patch under multiple depth assumptions, based on a global all-focus image and a set of refocused image patches, to form an initial matching cost volume. Each voxel corresponds to a matching cost at a specific depth for a spatial location. Using focus confidence as a guiding weight, the initial matching cost volume is spatially propagated. In high focus confidence regions, the original matching cost is retained to maintain depth details, while in low focus confidence regions, reliable matching costs are propagated from neighboring high-confidence regions to fill in uncertain areas. Through this focus confidence-guided spatial propagation process, noise and ambiguity in the initial matching cost volume are corrected. Finally, an initial depth estimate is obtained by performing a depth selection operation on the propagated matching cost volume.
[0036] Based on the initial depth estimate, the optimal sub-pixel level depth correction is calculated, expressed as: ; in, This represents the optimal sub-pixel level depth correction. This represents the depth correction amount to be optimized. Indicates spatial frequency, Representation of spatial frequency Related weighting functions, Indicates spatial frequency Multi-view phase difference, Represents pi (π). This indicates the equivalent baseline length of the light field camera. This represents the initial depth estimate.
[0037] It should be noted that, Representation of spatial frequency The relevant weighting function is constructed based on the signal-to-noise ratio and phase reliability statistical characteristics of the light field image at different spatial frequencies. It is usually set by analyzing and normalizing the spectral energy distribution of multi-view images. The equivalent baseline length of the light field camera refers to the maximum horizontal distance between the virtual cameras corresponding to the microlens array or view sampling. It is obtained by combining the optical parameters of the light field camera (such as the microlens focal length, sensor pixel spacing and main lens aperture) with the geometric relationship of the view arrangement.
[0038] The specific process includes solving the initial depth estimate to obtain the optimal sub-pixel level depth correction. This process is based on the principle of multi-view phase consistency. It calculates the phase difference of image signals under different views in the spatial frequency domain and combines a weight function related to spatial frequency to weight the phase error. By minimizing the sum of squares of the weighted phase error, the phase change observed from multiple views is made as consistent as possible with the theoretical phase change caused by depth, thereby determining the optimal sub-pixel level depth correction.
[0039] It should be noted that the multi-view phase consistency principle refers to the fact that image signals of the same spatial point under different viewpoints should have a deterministic phase shift in the frequency domain determined by parallax. When the assumed depth is consistent with the true depth, the phase difference of the signals from each viewpoint is minimized, thus high-precision depth estimation can be achieved by optimizing phase consistency.
[0040] Global optimization by multi-cue fusion of the optimal sub-pixel level depth correction, aggregated features, and focus confidence yields a high-resolution depth map characterizing the micro-undulations of the packaging surface.
[0041] The specific process includes using the optimal sub-pixel level depth correction as a geometric constraint, aggregating features to provide texture and structural information, focusing confidence to reflect local focusing reliability, and constructing a global objective function under a unified energy minimization framework. The depth smoothing term adaptively adjusts the regularization intensity based on the focusing confidence, preserving detail changes in high focusing confidence regions and enhancing smoothness in low focusing confidence regions. Meanwhile, the data fidelity term refines the optimal sub-pixel level depth correction guided by aggregating features. By iteratively optimizing the global objective function, a high-resolution depth map characterizing the micro-undulations of the packaging surface is obtained.
[0042] S2. Analyze the high-resolution depth map and extract the feature parameter vector describing the current three-dimensional deformation state of the packaging.
[0043] High-resolution depth maps are converted into 3D point clouds, and a physical perception map structure is constructed by combining geometric continuity constraints.
[0044] The specific process includes mapping each pixel position in the high-resolution depth map to three-dimensional spatial coordinates based on camera intrinsic parameters and depth values to form a three-dimensional point cloud. Then, the normal changes and curvature continuity between adjacent points are analyzed on the three-dimensional point cloud. Geometric continuity constraints are used to filter and weight the connection relationships of the local neighborhood of the point cloud, retaining adjacent edges that satisfy surface smoothness and structural consistency, and eliminating discontinuous connections caused by noise or abnormal fluctuations, thereby constructing a physical perception map structure.
[0045] It should be noted that geometric continuity constraints refer to the requirement in 3D surface reconstruction that the changes in normal and curvature between adjacent points or patches are smooth, so as to conform to the physical continuity characteristics of the surface of a real object.
[0046] The physical perception graph structure is input into a multi-layer deformable graph convolutional network, and the deformation features of the nodes are updated through a deformation perception attention mechanism.
[0047] The specific process includes inputting the physical perception graph structure into a multi-layer deformable graph convolutional network. In each layer, the multi-layer deformable graph convolutional network aggregates features of the nodes and their adjacency relationships of the physical perception graph structure. At the same time, it uses a deformation-aware attention mechanism to dynamically allocate attention weights based on the geometric differences and local curvature changes between nodes, so as to enhance the response to high deformation regions and suppress noise interference, thereby updating the node deformation features layer by layer.
[0048] Furthermore, the training process of the multilayer deformable graph convolutional network adopts a supervised learning approach. It utilizes a physical perception graph structure labeled with real node deformation features for end-to-end training. The physical perception graph structure is used as input, and the multilayer deformable graph convolutional network aggregates neighborhood information layer by layer and updates node deformation features by combining deformation-aware attention mechanism. Finally, the predicted node deformation features are output. During training, the mean squared error loss function is used to measure the difference between the predicted node deformation features and the real node deformation features. The parameters of the multilayer deformable graph convolutional network are adjusted through the backpropagation algorithm, and the process is iterated repeatedly until the predicted node deformation features reach stable convergence on the validation set.
[0049] It should be noted that the deformation-aware attention mechanism is a method that dynamically adjusts the weights of neighboring node information during graph convolution, measuring deformation sensitivity based on the differences in geometric deformation and local curvature changes between nodes. The deformation-aware attention mechanism assigns higher attention weights to high-deformation regions to enhance the response to subtle structural changes, while suppressing interference from flat or noisy regions.
[0050] Adaptive graph pooling and physical mode decomposition are performed on the updated node deformation features to generate the feature parameter vector of the current packaging 3D deformation state.
[0051] The specific process includes: when performing adaptive graph pooling on the updated node deformation features, dynamically merging neighboring nodes based on the response intensity and local geometric importance of the node deformation features, retaining the structural information that plays a key role in the overall deformation characterization, and performing physical mode decomposition on the pooled graph structure to decouple the three-dimensional deformation state into several orthogonal components according to the inherent physical vibration or deformation mode of the current packaged three-dimensional deformation state, generating feature parameter vectors including the direction and depth distribution of the main folds, the overall surface curvature, and the flatness offset of key regions.
[0052] It should be noted that the response intensity of node deformation features refers to the magnitude of the feature vector output by the node after being weighted by the deformation-aware attention mechanism in the multi-layer deformable graph convolutional network. It reflects the degree of deformation significance in the region where the node is located and is obtained by the attention-weighted aggregation of geometric differences and curvature changes during the forward propagation of the network. Local geometric importance refers to the degree of contribution of a node to the overall geometry in the physical perception graph structure based on its neighborhood curvature changes, normal differences, and connection structure. It is obtained by analyzing the geometric relationship between the node and its neighboring nodes in the 3D point cloud obtained by converting the high-resolution depth map.
[0053] S3. Using the characteristic parameter vector as boundary conditions, construct a packaging physical simulation model using the finite element method, and perform a fast static equilibrium solution on the packaging physical simulation model to generate a defect-free ideal dynamic template image.
[0054] Physical parameter inversion is performed on the characteristic parameter vector to obtain the material property field and boundary load conditions.
[0055] The specific process includes, when performing physical parameter inversion on the feature parameter vector, taking the feature parameter vector as input, and using the inverse problem solution method based on the finite element principle, under the premise of known packaging geometry, iteratively adjusting the material property field and boundary load conditions so that the deformation response generated by the forward physical simulation driven by the material property field and boundary load conditions matches the current three-dimensional deformation state of the packaging represented by the feature parameter vector, thereby obtaining the material property field and boundary load conditions consistent with the observed deformation.
[0056] It should be noted that the material property field refers to the spatially varying distribution of mechanical parameters within the packaging geometry, including elastic modulus, Poisson's ratio, etc., used to describe the stiffness and deformation characteristics of the packaging material at different locations; boundary load conditions refer to external forces, displacement constraints, or contact pressures acting on the packaging surface or edges, used to characterize the mechanical excitation or support constraints experienced by the packaging in actual working conditions; both serve as inputs to the physical simulation, determining the deformation response of the packaging under stress, and are the foundation for constructing a high-fidelity packaging physical simulation model.
[0057] The material property field and boundary load conditions are assigned to the reference geometric model of the packaging, and the nonlinear constitutive relations and constraints of the reference geometric model are defined to form a parameterized finite element description.
[0058] The specific process includes assigning material property fields and boundary load conditions to the reference geometric model of the packaging, enabling the reference geometric model of the packaging to have spatially varying mechanical parameters and external action conditions, and defining the nonlinear constitutive relation and constraint conditions of the reference geometric model of the packaging based on the actual mechanical behavior of the packaging material. The nonlinear constitutive relation describes the nonlinear dependence between stress and strain, and the constraint conditions limit the displacement or degree of freedom of the reference geometric model of the packaging at a specific position or direction, thereby forming a parameterized finite element description.
[0059] It should be noted that the parameterized finite element description refers to the integration of the packaging's baseline geometric model, material property field, boundary load conditions, nonlinear constitutive relations, and constraints into a finite element discrete framework in the form of variable parameters. Its function is to provide a structured and adjustable mechanical calculation basis for the packaging physical simulation model, and to support the efficient generation of corresponding deformation responses under different material properties or load conditions.
[0060] A physical simulation model of packaging is constructed based on parametric finite element description.
[0061] The specific process includes integrating the packaging's baseline geometric model, material property field, boundary load conditions, nonlinear constitutive relations, and constraints into the finite element solution framework based on parametric finite element description. By discretizing the spatial domain and establishing the relationship between nodal degrees of freedom and element stiffness matrices, a packaging physical simulation model capable of simulating the mechanical response behavior of packaging under actual working conditions is formed.
[0062] A fast static equilibrium solution is performed on the packaging physical simulation model to obtain a high-fidelity deformation displacement field, expressed as follows: ; in, This represents a high-fidelity deformation displacement field. This represents the displacement field variables that change during the optimization process. In the displacement field The internal force vector calculated below, This represents the external load vector applied to the packaging.
[0063] It should be noted that, In the displacement field The internal force vectors calculated below are nodal internal forces derived from the element stiffness matrix and material constitutive relation in the finite element method based on the current displacement field, and obtained from the derivative of each element strain energy with respect to displacement. The external load vector applied to the packaging is obtained by measuring the physical excitations such as forces, pressures, or contact constraints acting on the packaging surface under actual working conditions and converting them into equivalent concentrated forces or distributed loads on the finite element nodes through mechanical sensors.
[0064] The specific process includes: when performing rapid static equilibrium solution for the packaging physical simulation model, the displacement field variable is used as the optimization variable. Based on the current displacement field variable, the internal force vector at each finite element node is calculated through the material property field and nonlinear constitutive relation. At the same time, the external load vector applied to the packaging is used as a known input. The force balance equation is constructed, requiring the internal force vector to be equal to the external load vector. A Newton-Raphson iterative algorithm is used to calculate the unbalanced force, i.e., the difference between the internal force vector and the external load vector, in each step. The displacement correction is solved by combining the tangent stiffness matrix to update the displacement field variable. This process is repeated until the norm of the unbalanced force is less than the equilibrium convergence threshold. The corresponding displacement field variable at this time is the high-fidelity deformation displacement field.
[0065] It should be noted that the Newton-Raphson iterative algorithm is a numerical method for solving nonlinear equation systems. It performs a first-order Taylor expansion of the nonlinear function near the current solution and iteratively corrects the solution. Its purpose is to efficiently approximate the displacement field under static equilibrium in finite element analysis. The equilibrium convergence threshold is preset based on the scale of the packaging physics simulation model, material stiffness characteristics, and engineering accuracy requirements. It is typically set by analyzing the magnitude of the internal forces per unit volume or nodal residual forces, ensuring a balance between computational efficiency and deformation accuracy. An exemplary value range is 10. -4 Up to 10 -6 The value is in the Newton range, but the specific value will be adjusted based on the packaging size and the hardness of the material.
[0066] The high-fidelity deformation displacement field is compressed using the nonlinear eigenorthogonal decomposition method to obtain low-dimensional eigencoordinates characterizing the current deformation state.
[0067] The specific process includes extracting nonlinear modal basis functions from multiple high-fidelity deformation displacement field samples when compressing the high-fidelity deformation displacement field using the nonlinear intrinsic orthogonal decomposition method. These nonlinear modal basis functions are constructed by maximizing the cumulative contribution rate characterizing the deformation energy. The current high-fidelity deformation displacement field is then projected onto the low-dimensional subspace spanned by the nonlinear modal basis functions to obtain a set of coefficients. This set of coefficients quantifies the activation degree of the current deformation in each dominant mode, forming a low-dimensional intrinsic coordinate characterizing the current deformation state.
[0068] It should be noted that nonlinear modal basis functions are a set of orthogonal space functions extracted from high-fidelity deformation displacement field samples that can characterize the dominant modes of nonlinear deformation. Their role is to provide a low-dimensional subspace basis for nonlinear intrinsic orthogonal decomposition, thereby achieving efficient compression and reconstruction of complex deformations.
[0069] Based on low-dimensional intrinsic coordinates and pre-calibrated detection viewpoint conditions, a defect-free ideal dynamic template image is generated through differentiable rendering.
[0070] The specific process includes inputting the low-dimensional intrinsic coordinates into the differentiable rendering process based on the low-dimensional intrinsic coordinates and the pre-calibrated detection viewpoint conditions. The differentiable rendering process determines the camera pose and lighting configuration according to the pre-calibrated detection viewpoint conditions, and uses the geometric surface and material properties of the packaging physical simulation model to map the ideal deformation state represented by the low-dimensional intrinsic coordinates into the corresponding surface displacement, thereby generating a defect-free ideal dynamic template image that meets the observation conditions under that viewpoint.
[0071] It should be noted that the pre-calibrated detection viewing angle conditions are pre-calibrated based on the installation position, imaging angle and lighting environment of the light field camera in the actual detection scenario. This is achieved by collecting observation images of standard defect-free packaging samples from different viewing angles and recording the corresponding camera pose parameters. The ideal dynamic template image refers to a defect-free, high-fidelity rendered image generated based on the current three-dimensional deformation state of the packaging and the detection viewing angle conditions. Its purpose is to serve as a comparison benchmark for detecting whether there are abnormal defects on the surface of the actual packaging.
[0072] S4. Perform high-order continuous deformation registration on the global full-focus image and the ideal dynamic template image, and perform pixel-by-pixel difference calculation and adaptive threshold filtering in the color space to generate a pure difference image that excludes normal deformation interference.
[0073] High-order continuous deformation registration is performed on the global full-focus image and the ideal dynamic template image to generate a geometrically aligned deformable template image.
[0074] The specific process includes: when performing high-order continuous deformation registration between the global full-focus image and the ideal dynamic template image, a non-rigid registration method based on thin plate splines or high-order displacement fields is adopted. By optimizing an energy function that includes an image gray-level similarity term and a deformation smoothing regularization term, the ideal dynamic template image undergoes high-order continuous deformation while maintaining topological continuity, so as to match the geometric structure in the global full-focus image to the greatest extent and generate a geometrically aligned deformed template image.
[0075] Based on the deformed template image, global full-focus image, high-resolution depth map, and focus confidence, a consistency weighted threshold is calculated using a physical consistency verification function and then filtered to obtain a preliminary difference mask, expressed as: ; in, Indicates the position of a pixel in the image. Consistency weighted threshold at the point, Represents the global base threshold. This represents the natural exponential function. Represents the deep gradient modulation coefficients. Represents a high-resolution depth map At image pixel location The spatial gradient vector at that point, Indicates based on high-resolution depth map The normalized statistic obtained from the gradient magnitude, This represents the focus confidence modulation coefficient. Indicates the position of a pixel in the image. Focus confidence level at the point, Represents the modulation coefficients based on structural similarity. Indicates the position of a pixel in the image. The structural similarity index between the deformed template image and the global full-focus image.
[0076] It should be noted that the global baseline threshold is a baseline tolerance determined by statistical analysis of the difference between the global full-focus image of the defect-free packaging sample and the ideal dynamic template image, with an exemplary value range of 0.02 to 0.08; the depth gradient modulation coefficient is a parameter used to adjust the degree of influence of local depth changes on the consistency weighted threshold, which is obtained by normalizing the gradient statistical characteristics of the high-resolution depth map and combining it with defect detection performance optimization on the empirical or validation set; Represents a high-resolution depth map At image pixel location The spatial gradient vector at a given location is a two-dimensional vector obtained by taking partial derivatives of the high-resolution depth map in the horizontal and vertical directions, respectively, reflecting the direction and rate of change of depth values in the neighborhood of that pixel; the focus confidence modulation coefficient is a parameter used to adjust the influence of focus confidence on the consistency weighted threshold. It is determined by evaluating the accuracy and false detection rate of defect detection under different coefficients on the validation set, and by making empirical adjustments or optimizations based on the focus confidence distribution characteristics. Indicates the position of a pixel in the image. The focus confidence score is obtained by analyzing and normalizing the sharpness index (such as gradient magnitude or Laplacian variance) of the local sharpness response at that location in multi-view or light field images, and is used to measure the reliability of the pixel depth estimation. The structural similarity modulation coefficient is a parameter used to adjust the influence of structural similarity on the consistency weighting threshold. It is determined by analyzing the relationship between the structural similarity index and the difference response on verification samples containing normal deformation and real defects, and optimized according to the detection performance index. The structural similarity index is obtained by calculating the comprehensive measure of the similarity of brightness, contrast and structural information between the deformed template image and the global full-focus image within a local window, and is used to quantify the structural consistency of the two images at the perceptual level.
[0077] The specific process includes calculating a consistency weighted threshold based on the deformed template image, the global full-focus image, the high-resolution depth map, and the focus confidence, and then filtering it using a physical consistency verification function. This consistency weighted threshold is dynamically adjusted according to the severity of local depth changes, the level of focus confidence, and the structural similarity between the deformed template image and the global full-focus image. In regions with gentle depth gradients and high focus confidence, a stricter threshold is used to preserve the true defect response, while in regions with drastic depth changes or low focus confidence, the threshold is relaxed to avoid misjudging normal deformation as a defect, thereby generating a preliminary difference mask.
[0078] A Markov random field optimization with multi-source information constraints is performed on the initial difference mask to generate a pure difference map that excludes interference from normal deformation.
[0079] The specific process includes optimizing the initial difference mask using a Markov random field with multi-source information constraints. Each pixel in the initial difference mask is treated as a node in the Markov random field, and the label is defined as a defect or non-defect state. An energy function containing a data term and a smoothing term is constructed. The data term is determined by the response intensity of the initial difference mask, while the smoothing term is constrained by the geometric continuity of the high-resolution depth map, the spatial reliability of the focus confidence, and the texture consistency of the global full-focus image. In regions with high geometric continuity and high focus confidence, the consistency of adjacent pixel labels is encouraged. In regions with abrupt depth changes or low confidence, the smoothing constraints are relaxed to preserve the true defect boundaries. The energy function is minimized through methods such as graph cut or confidence propagation to generate a pure difference map that excludes the interference of normal deformation.
[0080] It should be noted that multi-source information constraints refer to combining data or information from multiple different sources (such as spectral information, spatial information, or prior knowledge of images) to guide and limit the optimization process of Markov random fields in order to generate more accurate pure difference maps. These are usually extracted or fused from multi-source data depending on the specific application scenario.
[0081] S5. Through a multimodal feature interaction mechanism, pixel-level feature alignment and fusion are performed on the clean difference map and the high-resolution depth map to obtain the segmentation mask, category and severity information of the defects.
[0082] The clean difference map and the high-resolution depth map are input into the multimodal feature interaction mechanism to perform pixel-level feature alignment and fusion, generating multi-scale fused features.
[0083] The specific process includes inputting the clean difference map and the high-resolution depth map into a multimodal feature interaction mechanism. The multimodal feature interaction mechanism extracts the appearance anomaly features of the clean difference map and the geometric undulation features of the high-resolution depth map at multiple scales. Within each scale, pixel-level spatial alignment is used to ensure the semantic correspondence of the two modalities at the same spatial location. The appearance and geometric information are dynamically weighted and fused using cross-attention or channel interaction strategies, so that the difference response is enhanced in the depth abrupt region and suppressed in the smooth region. The alignment and fusion results of different scales are aggregated layer by layer to generate multi-scale fused features.
[0084] It should be noted that the multimodal feature interaction mechanism is a method for fusing features from different modalities. Through pixel-level spatial alignment and dynamic weight allocation, appearance anomaly information and geometric structure information enhance each other at corresponding positions, generating complementary joint feature representations.
[0085] Multi-scale fused features are input into a hierarchical decoder, and preliminary segmentation masks, categories, and severity information of defects are obtained through multi-task joint inference.
[0086] The specific process includes inputting multi-scale fused features into a hierarchical decoder. The hierarchical decoder recovers spatial resolution through progressive upsampling and simultaneously performs semantic segmentation, classification, and regression tasks at each decoding level. Specifically, the semantic segmentation branch outputs a preliminary segmentation mask of the defect, the classification branch predicts the category to which the defect belongs, and the regression branch estimates the severity information of the defect. The three tasks share encoded features and constrain each other. Through multi-task joint inference, the consistency optimization of defect region localization, type discrimination, and severity assessment is achieved, thereby obtaining the preliminary segmentation mask, category, and severity information of the defect.
[0087] It should be noted that the hierarchical decoder is a neural network structure that gradually recovers spatial details by upsampling and fusing multi-scale features at different resolution levels, and is used to simultaneously perform multiple related tasks at different resolution levels.
[0088] Spatial context refinement is performed on the initial segmentation mask, category, and severity information to obtain the segmentation mask, category, and severity information of defects.
[0089] The specific process includes refining the spatial context of the initial segmentation mask, category, and severity information by using a non-local context aggregation method to model long-range dependencies between pixels while preserving the details of the defect boundaries. The boundaries of the initial segmentation mask are corrected based on local appearance consistency, geometric continuity, and category semantic distribution. The category prediction results are smoothed for regional consistency, and the severity information is adaptively calibrated within the spatial neighborhood to obtain the segmentation mask, category, and severity information of the defect.
[0090] It should be noted that nonlocal context aggregation is a feature enhancement method that models long-distance dependencies by obtaining the feature similarity between any two locations in an image. Its role is to integrate global semantic information into the representation of each pixel without relying on local neighborhood limitations, thereby improving the ability to characterize structural consistency and semantic coherence in segmentation, detection, or recognition tasks.
[0091] S6. Perform quality grading decisions and sorting control on the segmentation mask, category, and severity information of defects, and associate them with feature parameter vectors to generate a visual quality traceability report; the visual quality traceability report includes defect root cause inference and process adjustment suggestions.
[0092] Based on the segmentation mask, category, severity information and feature parameter vector of the defect, an intelligent decision state representation is constructed.
[0093] The specific process includes: based on the segmentation mask, category, severity information and feature parameter vector of the defect, the spatial distribution range of the defect is provided by the segmentation mask, the semantic type of the defect is identified by the category, the degree of harm of the defect is quantified by the severity information, and the three-dimensional deformation characteristics of the overall packaging are represented by the feature parameter vector. The four together constitute a structured high-dimensional vector. This high-dimensional vector integrates multi-dimensional information of geometry, semantics, degree and physical state to form an intelligent decision state representation.
[0094] The intelligent decision-making state representation is input into the deep reinforcement learning strategy to execute multi-objective optimization decisions and generate sorting control instructions.
[0095] The specific process includes inputting the intelligent decision state representation into a deep reinforcement learning policy. The deep reinforcement learning policy uses the intelligent decision state representation as the current environment state and evaluates the long-term cumulative reward of different sorting actions through a policy network. At the same time, it takes into account multiple optimization objectives such as defect rejection rate, production efficiency and resource consumption. During the training process, the network parameters are continuously adjusted using experience playback and policy gradient methods. Finally, the optimal action is selected and the corresponding sorting control instruction is generated in the inference stage.
[0096] It should be noted that the policy gradient method is a reinforcement learning algorithm that directly optimizes the policy parameters of an agent. It maximizes the expected cumulative reward by estimating the gradient of the reward signal with respect to the policy parameters and updating the parameters along its ascending direction. Its role is to achieve end-to-end policy learning in continuous or high-dimensional action spaces, and it is suitable for tasks where it is impossible to explicitly model the dynamics of the environment or where stochastic policies are required.
[0097] The system executes sorting control commands and collects real-time production line status data during the execution process. After fusing the real-time production line status data with feature parameter vectors, it performs causal inference and counterfactual analysis to generate a visualized quality traceability report.
[0098] The specific process includes: after executing the sorting control command, collecting real-time status data of the production line during the execution process. The real-time status data of the production line includes the execution time of sorting actions, the posture of the robotic arm, the speed of the conveyor belt, and operating parameters such as ambient temperature and humidity. The real-time status data of the production line is fused with the feature parameter vector under a unified spatiotemporal alignment framework to form a joint representation that includes packaging deformation characteristics and production line operation context. Based on this joint representation, the causal inference method is applied to identify key intervention variables in the cause of defects. Counterfactual analysis is used to simulate whether defects can be avoided under different operating conditions or material properties, and a visualized quality traceability report is generated.
[0099] In summary, this invention achieves the following: First, by generating a global full-focus image using a depth probability prediction network, it ensures strict temporal and spatial synchronization and pixel-level alignment between 2D texture information and 3D geometric information, enabling intelligent expansion of the imaging depth of field and providing a high-quality 2D information foundation. Second, through deformation compensation and difference extraction using physical simulation and dynamic template generation, it achieves abstraction from apparent geometry to intrinsic physical parameters, providing precise input conditions for physical simulation. It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for detecting packaging defects based on machine vision, characterized in that: include, The original light field data of the flexible packaging to be tested is collected, and a global full-focus image that eliminates local defocus is generated using a depth probability prediction network. A high-resolution depth map characterizing the micro-undulations of the packaging surface is obtained from the original light field data. The high-resolution depth map is analyzed and the feature parameter vector describing the current three-dimensional deformation state of the packaging is extracted. Using the feature parameter vector as boundary conditions, a physical simulation model of packaging is constructed using the finite element method, and a fast static equilibrium solution is performed on the physical simulation model of packaging to generate a defect-free ideal dynamic template image. High-order continuous deformation registration is performed on the global full-focus image and the ideal dynamic template image, and pixel-by-pixel difference calculation and adaptive threshold filtering are performed in the color space to generate a pure difference image that excludes normal deformation interference. A multimodal feature interaction mechanism is used to perform pixel-level feature alignment and fusion between the clean difference map and the high-resolution depth map to obtain the segmentation mask, category and severity information of defects; The system performs quality grading decisions and sorting controls based on the segmentation mask, category, and severity information of defects, and associates feature parameter vectors to generate a visual quality traceability report.
2. The packaging defect detection method based on machine vision as described in claim 1, characterized in that: The raw light field data includes spatial location information, angular direction information, and light intensity information.
3. The packaging defect detection method based on machine vision as described in claim 2, characterized in that: The process involves acquiring the raw light field data of the flexible packaging to be inspected and using a depth probability prediction network to generate a global full-focus image that eliminates local defocusing. The specific steps are as follows: Raw light field data is acquired using a light field camera, and an initial depth probability distribution map and uncertainty heatmap are obtained through a depth probability prediction network. A set of non-uniformly sampled refocusing image patches is then generated. Interactively fuse the refocused image patch set with the initial depth probability distribution map to generate aggregated features and focus confidence for each pixel location; Using aggregation features and focus confidence as conditions, a global full-focus image with local defocusing is obtained by optimizing the composite loss function.
4. The packaging defect detection method based on machine vision as described in claim 3, characterized in that: The specific steps for obtaining a high-resolution depth map characterizing the micro-undulations of the packaging surface from the original light field data are as follows. Based on the set of global full-focus images and refocused image patches, an initial matching cost volume is constructed, and the spatial propagation of the initial matching cost volume is guided by the focus confidence to obtain the initial depth estimate. Calculate the optimal sub-pixel level depth correction based on the initial depth estimate; Global optimization by multi-cue fusion of the optimal sub-pixel level depth correction, aggregated features, and focus confidence yields a high-resolution depth map characterizing the micro-undulations of the packaging surface.
5. The packaging defect detection method based on machine vision as described in claim 4, characterized in that: The specific steps for parsing the high-resolution depth map and extracting the feature parameter vector describing the current three-dimensional deformation state of the packaging are as follows. High-resolution depth maps are converted into 3D point clouds, and a physical perception map structure is constructed by combining geometric continuity constraints. The physical perception graph structure is input into a multi-layer deformable graph convolutional network, and the node deformation features are updated through a deformation perception attention mechanism. Adaptive graph pooling and physical mode decomposition are performed on the updated node deformation features to generate the feature parameter vector of the current packaging 3D deformation state.
6. The packaging defect detection method based on machine vision as described in claim 5, characterized in that: The specific steps for constructing a physical simulation model of packaging using the finite element method with characteristic parameter vectors as boundary conditions are as follows. Physical parameter inversion is performed on the characteristic parameter vector to obtain the material property field and boundary load conditions; The material property field and boundary load conditions are assigned to the reference geometric model of the packaging, and the nonlinear constitutive relations and constraints of the reference geometric model are defined to form a parameterized finite element description. A physical simulation model of packaging is constructed based on parametric finite element description.
7. The packaging defect detection method based on machine vision as described in claim 6, characterized in that: The specific steps for rapidly solving the static equilibrium of the packaging physical simulation model to generate a defect-free ideal dynamic template image are as follows. A fast static equilibrium solution is performed on the packaging physical simulation model to obtain a high-fidelity deformation displacement field; The high-fidelity deformation displacement field is compressed using the nonlinear eigenorthogonal decomposition method to obtain low-dimensional eigencoordinates characterizing the current deformation state. Based on low-dimensional intrinsic coordinates and pre-calibrated detection viewpoint conditions, a defect-free ideal dynamic template image is generated through differentiable rendering.
8. The packaging defect detection method based on machine vision as described in claim 7, characterized in that: The specific steps for generating a pure difference map that excludes interference from normal deformation are as follows. High-order continuous deformation registration is performed on the global full-focus image and the ideal dynamic template image to generate a geometrically aligned deformable template image; Based on deformable template images, global full-focus images, high-resolution depth maps, and focus confidence, a consistency weighted threshold is calculated using a physical consistency verification function and then filtered to obtain a preliminary difference mask. A Markov random field optimization with multi-source information constraints is performed on the initial difference mask to generate a pure difference map that excludes interference from normal deformation.
9. The packaging defect detection method based on machine vision as described in claim 8, characterized in that: The specific steps for obtaining the segmentation mask, category, and severity information of the defects are as follows: The pure difference map and the high-resolution depth map are input into the multimodal feature interaction mechanism to perform pixel-level feature alignment and fusion, generating multi-scale fused features; Multi-scale fused features are input into a hierarchical decoder, and preliminary segmentation masks, categories, and severity information of defects are obtained through multi-task joint inference. Spatial context refinement is performed on the initial segmentation mask, category, and severity information to obtain the segmentation mask, category, and severity information of defects.
10. The packaging defect detection method based on machine vision as described in claim 9, characterized in that: The specific steps for generating the visualized quality traceability report are as follows: Based on the segmentation mask, category, severity information and feature parameter vector of the defect, an intelligent decision state representation is constructed; The intelligent decision-making state representation is input into a deep reinforcement learning strategy to execute multi-objective optimization decisions and generate sorting control instructions. The system executes sorting control commands and collects real-time production line status data during the execution process. After fusing the real-time production line status data with feature parameter vectors, it performs causal inference and counterfactual analysis to generate a visualized quality traceability report.