Footwear product defect detection method and device based on Gaussian splashing and wavelet transformation
By constructing a 3D Gaussian model based on Gaussian splashing and wavelet transform, and combining frequency domain feature extraction and adaptive strategies, the problems of blurred high-frequency texture reconstruction and difficulty in detecting minute defects are solved, achieving efficient defect detection results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUAQIAO UNIVERSITY
- Filing Date
- 2026-04-20
- Publication Date
- 2026-05-19
AI Technical Summary
Existing 3D reconstruction technology is prone to oversmoothing when processing high-frequency details, which leads to blurred defect detection of high-frequency texture features such as mesh uppers and fine weave patterns on soles, making it difficult to achieve high-sensitivity defect detection.
An anisotropic 3D Gaussian model is constructed using a method based on Gaussian splashing and wavelet transform. The model is trained by a frequency domain feature extractor. The camera pose is optimized by combining a hybrid optimization target loss function and an adaptive strategy, using a gradient descent iterative optimization algorithm for photometric error. Then, rasterization rendering and wavelet residual calculation are performed to generate a defect mask for defect detection.
It improves the ability to restore high-frequency textures, achieves high-fidelity detail reconstruction and accurate defect detection, and is suitable for millisecond-level detection in large-scale industrial production lines, with ultra-high detection accuracy.
Smart Images

Figure CN122066701A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing, specifically to a method and apparatus for detecting defects in footwear products based on Gaussian splashing and wavelet transform. Background Technology
[0002] In industrial appearance inspection, defect detection on surfaces with high-frequency texture features, such as shoe uppers and soles with subtle weave patterns, has always been a challenge. Traditional 2D vision methods are easily affected by changes in lighting, while existing 3D reconstruction techniques often exhibit "oversmoothing" when processing high-frequency details, causing dense mesh or weave patterns to become blurry in the reconstructed model. This blurring can mask subtle physical defects (such as broken yarns or minor wear), making defect detection based on the reconstructed model difficult to implement. Therefore, how to enhance the ability of 3D models to reproduce high-frequency textures and utilize them for highly sensitive defect detection remains a problem that needs to be solved. Summary of the Invention
[0003] The purpose of this application is to propose a method and device for detecting defects in footwear products based on Gaussian splashing and wavelet transform, which addresses the aforementioned technical problems and solves the issues of blurred high-frequency texture reconstruction and difficulty in detecting minute defects.
[0004] In a first aspect, the present invention provides a method for detecting defects in footwear products based on Gaussian splashing and wavelet transform, comprising the following steps:
[0005] An anisotropic 3D Gaussian model and a frequency domain feature extractor are constructed. The 3D Gaussian model is trained using the frequency domain feature extractor to obtain a trained 3D Gaussian model. During the training process, a hybrid optimization objective loss function is adopted and an adaptive strategy is executed. The adaptive strategy includes an adaptive compaction strategy based on frequency domain gradient and a frequency domain energy-driven dynamic adjustment strategy for the order of spherical harmonic functions.
[0006] The image to be detected is acquired, and the camera pose of the image to be detected is optimized using a gradient descent iterative optimization algorithm based on photometric error to obtain the optimized camera pose. The optimized camera pose is then used to perform rasterization rendering on the trained 3D Gaussian model to obtain a standard reference image. Wavelet residual calculation and statistics are performed based on the standard reference image and the image to be detected to obtain a wavelet residual map and an adaptive threshold. A defect mask is generated using the adaptive threshold and the wavelet residual map, and the defect detection result is obtained based on the defect mask detection.
[0007] As a preferred method, the construction process of the 3D Gaussian model is as follows:
[0008] Images of defect-free footwear products captured under different camera poses are acquired and sparse point cloud data is generated. An initial Gaussian primitive is initialized for each point in the sparse point cloud data. Each Gaussian primitive contains a set of learnable attribute parameters. For the first point in the sparse point cloud data... The corresponding point Gao Siyuan The attribute parameters include center position. Opacity , order of spherical harmonic function and the three-dimensional covariance matrix The three-dimensional covariance matrix is decomposed into rotation matrices. and scaling matrix As shown in the following formula:
[0009] ;
[0010] Where T represents the transpose of the matrix;
[0011] All Gaussian elements are used to construct a 3D Gaussian model.
[0012] Preferably, the frequency domain feature extractor includes a two-dimensional discrete wavelet transform module, which uses Daubechies wavelets as basis functions and includes a low-pass filter and a high-pass filter.
[0013] The image is convolved and decomposed using a low-pass filter and a high-pass filter to obtain low-frequency components, horizontal high-frequency components, vertical high-frequency components, and diagonal high-frequency components, as shown in the following formula:
[0014] ;
[0015] ;
[0016] ;
[0017] ;
[0018] in, Represents an image. The function representing the low-pass filter, This represents the function corresponding to a high-pass filter. This represents the low-frequency components of the image. This represents the horizontal high-frequency components corresponding to the image. This represents the vertical high-frequency component corresponding to the image. This represents the diagonal high-frequency components of the image;
[0019] By concatenating the horizontal, vertical, and diagonal high-frequency components of the image, the high-frequency feature tensor corresponding to the image is obtained, as shown in the following equation:
[0020] ;
[0021] in, Indicates splicing, Representing an image The corresponding high-frequency feature tensor.
[0022] As a preferred method, the rasterization rendering process is as follows:
[0023] The three-dimensional covariance matrix of each Gaussian element in the 3D Gaussian model is obtained by using a local affine approximation method. The projection is a two-dimensional covariance matrix. As shown in the following formula:
[0024] ;
[0025] in, The world-to-view transformation matrix for the current camera. Let T be the affine approximation Jacobian matrix of the projective transformation at the center of the Gaussian element, and let T denote the matrix transpose.
[0026] Based on the camera's view frustum, a culling operation is performed on all Gaussian primitives in the 3D Gaussian model. The center positions of the remaining Gaussian primitives after culling are then sorted from front to back according to their depth values in the camera coordinate system, generating an ordered list of Gaussian primitive indices. ;
[0027] According to Gaussian meta-index list The actual opacity corresponding to the opacity of each Gaussian element is calculated as follows:
[0028] ;
[0029] in, Represents the Gaussian meta-index list The first in The coordinates of the center position of each Gaussian element projected onto the two-dimensional projection area; Indicates the first Any pixel in a two-dimensional projection plane, represented by a Gaussian element. Represents the Gaussian meta-index list The first in The opacity of a Gaussian unit. Represents the Gaussian meta-index list The first in The actual opacity of the Gaussian unit;
[0030] By analyzing the Gaussian meta-index list The rendering color is obtained by performing alpha blending on each Gaussian element, as shown in the following formula:
[0031] ;
[0032] in, Indicates the rendered image at any pixel The rendered color at that location. Represents the Gaussian meta-index list The first in The actual opacity of the Gaussian unit. Represents the Gaussian meta-index list The first in The view-related color is calculated by the order of a spherical harmonic function for a given Gaussian element in the current camera pose.
[0033] As a preferred method, the training process for the 3D Gaussian model is as follows:
[0034] The initial 3D Gaussian model is used as the first 3D Gaussian model in the first iteration cycle to start the iterative training process;
[0035] The 3D Gaussian model will be rasterized and rendered at each iteration step within the current iteration cycle to obtain the corresponding rendered image. Using images of defect-free footwear products as real images, and combining them with rendered images, spatial domain loss and structural similarity loss are constructed, as shown in the following equation:
[0036] ;
[0037] ;
[0038] in, Indicates spatial domain loss, Describing the L1 norm, Represents a real image; Represents structural similarity loss. Represents the structural similarity function;
[0039] The rendered image and the real image are respectively input into the frequency domain feature extractor to extract high-frequency feature tensors for different texture directions, and the frequency domain perceptual loss is calculated as shown in the following formula:
[0040] ;
[0041] Where j represents the texture parameters corresponding to different texture directions, including horizontal, vertical and diagonal directions, and j corresponds to LH, HL and HH respectively; This represents the high-frequency features of the high-frequency feature tensor corresponding to the rendered image along the texture direction under texture parameter j. This represents the high-frequency features of the high-frequency feature tensor corresponding to the real image in the texture direction under texture parameter j. The adaptive weights corresponding to texture parameter j;
[0042] A hybrid optimization objective loss function is constructed based on spatial domain loss, structural similarity loss, and frequency domain sensing loss, as shown in the following equation:
[0043] ;
[0044] in, Denotes the hybrid optimization objective loss function. and These represent similarity weights and frequency domain perception weights, respectively.
[0045] The adaptive compaction strategy based on frequency domain gradient is as follows:
[0046] The mean view space gradient norm of the center position of each Gaussian element in the Gaussian element index list is calculated by backpropagation in each iteration cycle, as shown in the following formula:
[0047] ;
[0048] in, Represents the first element in the Gaussian meta-index list. Mean of the view space gradient norm corresponding to each Gaussian element. This represents the total number of iterations within one iteration cycle. Indicates the first Number of iterations, Indicates the first The hybrid optimization objective loss function calculated by forward propagation at the number of iterations. Represents the first element in the Gaussian meta-index list. The central position of each Gaussian element In the The two-dimensional coordinate vector that falls into the two-dimensional projection plane after the projection transformation in the nth iteration step. Represents the L2 norm;
[0049] Calculate the frequency domain response factor within the effective projected pixel set corresponding to each Gaussian element in the Gaussian element index list. As shown in the following formula:
[0050] ;
[0051] in, Represents the first element in the Gaussian meta-index list. The frequency domain response factor corresponding to each Gaussian element. Represents the first element in the Gaussian meta-index list. The effective projected pixel set of a Gaussian element on a two-dimensional projection plane, where the effective projected pixel set is the set of pixels whose actual opacity exceeds an opacity threshold, is represented as: , Indicates the opacity threshold. This represents the total number of pixels in the pixel set. Represents a two-dimensional projection plane; This represents the high-frequency error map constructed based on the high-frequency feature tensors corresponding to the rendered image and the high-frequency feature tensors corresponding to the real image at any pixel. The high-frequency error value at that location, and the expression for the high-frequency error map are: ; Represents the first element in the Gaussian meta-index list. Each Gaussian unit corresponds to any pixel in a two-dimensional projection plane. The actual opacity at the location;
[0052] The corrected gradient for each Gaussian element is calculated based on the mean viewspace gradient norm and frequency domain response factor of each Gaussian element in the Gaussian element index list, as shown in the following formula:
[0053] ;
[0054] in, Represents the first element in the Gaussian meta-index list. The corrected gradient corresponding to each Gaussian element, where β is the upper limit of gain and γ is the sensitivity coefficient;
[0055] The Gaussian meta-index list is updated based on the corrected gradient, as follows:
[0056] In response to determining the first If the corrected gradient corresponding to the th Gaussian element is greater than the gradient threshold, and the maximum value of the element in the corresponding scaling matrix is greater than the scaling threshold, then for the th... Each Gaussian unit performs a splitting operation;
[0057] In response to determining the first If the corrected gradient corresponding to the th Gaussian element is greater than the gradient threshold, and the maximum value of the corresponding element in the scaling matrix is less than or equal to the scaling threshold, then for the th... Each Gaussian unit performs a cloning operation;
[0058] In response to determining the first If the corrected gradient corresponding to the nth Gaussian element is less than or equal to the gradient threshold, then the original nth Gaussian element is retained. A high-level unit;
[0059] The updated Gaussian meta-index list of the current iteration period is used as the Gaussian meta-index list of the next iteration period for forward and backward propagation. The updated Gaussian meta-index list of the last iteration period is used to construct a trained 3D Gaussian model.
[0060] Triggering a frequency-domain energy-driven dynamic adjustment strategy for the order of spherical harmonics in one iteration step within the current iteration cycle includes the following steps:
[0061] Calculate the high-frequency energy features corresponding to each pixel in the real image to extract the high-frequency energy feature map. As shown in the following formula:
[0062] ;
[0063] in, , and These represent the horizontal, vertical, and diagonal components of the real image at any pixel. The component value at the location; Represents any pixel in a real image High-frequency energy characteristics at the location;
[0064] The normalized high-frequency energy distribution probability of each Gaussian element within the effective projected pixel set is calculated as follows:
[0065] ;
[0066] in, Indicates the first Each Gaussian unit corresponds to any pixel in a two-dimensional projection plane. The normalized high-frequency energy distribution probability at that location. Represents any pixel in a real image High-frequency energy characteristics at that location Indicates the first Each Gaussian unit corresponds to any pixel in a two-dimensional projection plane. The actual opacity at that location, Represent one of the positive constants;
[0067] The local frequency domain entropy of each Gaussian element is calculated based on the normalized high-frequency energy distribution probability, as shown in the following equation:
[0068] ;
[0069] in, Indicates the first The local frequency domain entropy of a Gaussian element, Represent another positive number;
[0070] The adjusted order of the spherical harmonic function is determined based on the mapping relationship between local frequency domain entropy and the order of the spherical harmonic function, as shown in the following equation:
[0071] ;
[0072] in, This indicates the number of the current iteration step. The order of the adjusted spherical harmonic function corresponding to each Gaussian element. Low-frequency threshold; For high-frequency thresholds, Let the order of the maximum spherical harmonic function be the number of iterations in one of them. The adjusted spherical harmonic function order corresponding to the _ Gaussian elements is used as the _th _ in the current preset iteration interval. The order of the spherical harmonic function corresponding to each Gaussian element is determined and propagated forward and backward.
[0073] Preferably, a gradient descent iterative optimization algorithm based on photometric error is used to optimize the camera pose of the image to be detected, resulting in an optimized camera pose, specifically including:
[0074] Using the current camera pose The trained 3D Gaussian model is rasterized to obtain the current rendered image. ;
[0075] Calculate the current rendered image With the image to be detected The photometric error loss between them is shown in the following formula:
[0076] ;
[0077] in, Indicates photometric error loss, These are the gradient weight coefficients. This represents the spatial gradient extracted from the current rendered image using the edge extraction operator. This represents the spatial gradient extracted from the image to be detected using the edge extraction operator. Describing the L1 norm, Represents the L2 norm;
[0078] The gradient of the photometric error loss with respect to the camera pose is calculated using a multi-level chain rule, as shown in the following equation:
[0079] ;
[0080] in, This represents the gradient of the photometric error loss with respect to the camera pose. To calculate the derivative of pixel-level error on a two-dimensional projection plane; To calculate the actual center position of the Gaussian pixel for pixel color changes The partial derivatives; The core projection Jacobian matrix connecting 2D and 3D;
[0081] The camera pose is updated using the Adam optimizer. The process stops when the preset number of iterations is reached or the photometric error loss converges, resulting in the optimized camera pose. ;
[0082] Wavelet residual calculation and statistics are performed based on a standard reference image and the image to be detected to obtain a wavelet residual map and an adaptive threshold. A defect mask is generated using the adaptive threshold and the wavelet residual map. The defect detection results are obtained based on the defect mask detection, specifically including:
[0083] The standard reference image and the image to be detected are input into a low-pass filter and a high-pass filter for convolution decomposition to obtain the corresponding horizontal high-frequency components, vertical high-frequency components, and diagonal high-frequency components. The wavelet residual map is then calculated using the following formula:
[0084] ;
[0085] in, , and These represent the horizontal, vertical, and diagonal high-frequency components of the image to be detected. , and These represent the horizontal, vertical, and diagonal high-frequency components corresponding to the standard reference image, respectively. Represents the wavelet residual plot;
[0086] Calculate the mean of the wavelet residual plot and standard deviation The adaptive threshold is calculated using the following formula:
[0087] ;
[0088] in, Represents the determination coefficient. Indicates the defect determination threshold;
[0089] The defect mask is determined based on the wavelet residual image and the adaptive threshold, as shown in the following formula:
[0090] ;
[0091] in, Represents any pixel in the defect mask The corresponding mask value;
[0092] If the mask value is 1, the corresponding pixel in the image to be detected is determined to be a defect point; if the mask value is 0, the corresponding pixel in the image to be detected is determined to be a normal texture point; the connected region formed by the defect points is taken as the defect detection result.
[0093] Secondly, the present invention provides a defect detection device for footwear products based on Gaussian splashing and wavelet transform, comprising:
[0094] The model reconstruction module is configured to construct an anisotropic 3D Gaussian model and a frequency domain feature extractor. The frequency domain feature extractor is used to train the 3D Gaussian model to obtain a trained 3D Gaussian model. During the training process, a hybrid optimization objective loss function is adopted and an adaptive strategy is executed. The adaptive strategy includes an adaptive compaction strategy based on frequency domain gradient and a frequency domain energy-driven dynamic adjustment strategy for the order of spherical harmonic functions.
[0095] The detection module is configured to acquire the image to be detected, optimize the camera pose of the image to be detected using a gradient descent iterative optimization algorithm based on photometric error, obtain the optimized camera pose, perform rasterization rendering on the trained 3D Gaussian model using the optimized camera pose to obtain a standard reference image, perform wavelet residual calculation and statistics based on the standard reference image and the image to be detected to obtain a wavelet residual map and an adaptive threshold, generate a defect mask using the adaptive threshold and the wavelet residual map, and obtain the defect detection result based on the defect mask detection.
[0096] Thirdly, the present invention provides an electronic device including one or more processors; and a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any implementation of the first aspect.
[0097] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method as described in any of the implementations of the first aspect.
[0098] Fifthly, the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the method as described in any of the implementations in the first aspect.
[0099] Compared with the prior art, the present invention has the following beneficial effects:
[0100] (1) The defect detection method for footwear products based on Gaussian splash and wavelet transform proposed in this invention not only calculates the spatial domain loss of RGB images during the training process of 3D Gaussian model, but also converts the rendered image and the real image to the frequency domain, separating low-frequency background and high-frequency texture information.
[0101] (2) The defect detection method for footwear products based on Gaussian splashing and wavelet transform proposed in this invention adopts an adaptive densification strategy based on frequency domain gradient during the training process of the 3D Gaussian model, and constructs a hybrid optimization loss function based on high-frequency subband differences. This hybrid optimization loss function is used to intervene in the splitting mechanism of Gaussian units. Specifically, the mean of the view space gradient norm of Gaussian units is weighted according to the magnitude of the high-frequency error, forcing Gaussian units to perform denser "splitting" operations in texture-rich regions, thereby increasing the sampling density. Multi-scale wavelet transform and adaptive gradient flow control are deeply coupled in the optimization loop of 3D Gaussian splashing to achieve high-fidelity detail reconstruction and accurate defect detection.
[0102] (3) The footwear product defect detection method based on Gaussian splashing and wavelet transform proposed in this invention uses a frequency domain energy-driven dynamic adjustment strategy of spherical harmonic function order to detect the high-frequency energy characteristics of Gaussian elements. For Gaussian elements with high-frequency energy characteristics greater than the threshold, the corresponding spherical harmonic function order is dynamically increased to enhance the detail expression. Therefore, it can accurately capture the high-frequency texture details and small difference information of the object surface in the frequency domain space, effectively solving the problem of difficult detection of small defects on the surface of footwear products. It has millisecond-level detection speed and ultra-high detection accuracy, and is suitable for large-scale industrial production line deployment. Attached Figure Description
[0103] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0104] Figure 1 This is a flowchart illustrating an embodiment of the footwear product defect detection method based on Gaussian splashing and wavelet transform, which is an example of this application.
[0105] Figure 2 This is a flowchart illustrating the training and inference phases of the footwear product defect detection method based on Gaussian splashing and wavelet transform, as an embodiment of this application.
[0106] Figure 3 This is a flowchart illustrating the frequency domain energy-driven dynamic adjustment strategy for the spherical harmonic function order in the footwear product defect detection method based on Gaussian splashing and wavelet transform, which is an embodiment of this application.
[0107] Figure 4 The standard reference image generated in the footwear product defect detection method based on Gaussian splashing and wavelet transform in the embodiments of this application;
[0108] Figure 5 This is an illustration of the underlying spatial structure of footwear products, which is represented and modeled using Gaussian splash and wavelet transform, as an embodiment of the present application.
[0109] Figure 6 This is a schematic diagram of a footwear product defect detection device based on Gaussian splashing and wavelet transform, which is an embodiment of this application.
[0110] Figure 7 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0111] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0112] Figure 1 This application illustrates an embodiment of a method for detecting defects in footwear products based on Gaussian splashing and wavelet transform, comprising the following steps:
[0113] S1. Construct an anisotropic 3D Gaussian model and a frequency domain feature extractor. Use the frequency domain feature extractor to train the 3D Gaussian model to obtain a trained 3D Gaussian model. During the training process, a hybrid optimization objective loss function is adopted and an adaptive strategy is executed. The adaptive strategy includes an adaptive compaction strategy based on frequency domain gradient and a frequency domain energy-driven dynamic adjustment strategy for the order of spherical harmonic functions.
[0114] In a specific embodiment, the construction process of the 3D Gaussian model is as follows:
[0115] Images of defect-free footwear products captured under different camera poses are acquired and sparse point cloud data is generated. An initial Gaussian primitive is initialized for each point in the sparse point cloud data. Each Gaussian primitive contains a set of learnable attribute parameters. For the first point in the sparse point cloud data... The corresponding point Gao Siyuan The attribute parameters include center position. Opacity , order of spherical harmonic function and the three-dimensional covariance matrix The three-dimensional covariance matrix is decomposed into rotation matrices. and scaling matrix As shown in the following formula:
[0116] ;
[0117] Where T represents the transpose of the matrix;
[0118] All Gaussian elements are used to construct a 3D Gaussian model.
[0119] Specifically, an industrial camera is used to acquire multi-view images of a standard reference sample, obtaining a calibrated image set of defect-free footwear products. Subsequently, the Structure for Motion Restoration (SfM) algorithm is used to extract and match features from this image set, calculating the camera's intrinsic and extrinsic parameters, and simultaneously resolving the sparse point cloud data of the defect-free footwear product in three-dimensional space. Sparse dotted clouds Each point in the graph is used to initialize a Gaussian element, and each Gaussian element contains a set of learnable attribute parameters: center position. Opacity , order of spherical harmonic function and the three-dimensional covariance matrix .
[0120] To achieve anisotropic modeling and ensure the positive semidefiniteness of the covariance matrix, instead of directly optimizing the covariance matrix, it is decomposed into rotation matrices. and scaling matrix These three coefficients are used to characterize the 3D rotation of the Gaussian element and its scaling along the three principal axes, respectively. Independent optimization of these three coefficients imparts anisotropic geometric characteristics to the Gaussian element, typically based on the point-to-point relationships between adjacent points in the point cloud. Set the average distance as the initial scaling value. From unit quaternion Obtained through conversion. During initialization, It is usually set as the identity matrix.
[0121] To ensure the actual opacity of Gaussian primitives projected onto a two-dimensional plane Always strictly bound by Within the effective physical range, to meet the needs of actual rasterization rendering, embodiments of this application introduce corresponding unconstrained parameters. . This represents the opacity representation value that actually participates in spatial color blending during forward rendering, while the unconstrained parameter... This means that values are truly maintained in video memory for a long time, directly updated by the optimizer, and their range is the entire real number domain. The underlying differentiable variables. The dynamic mapping relationship between the two runs through the entire lifecycle of the 3D Gaussian model: in the initialization phase of the 3D Gaussian model, the inverse Sigmoid function is used for inverse solution, based on the preset initial actual opacity. Calculate backwards and assign unconstrained parameters The corresponding initial value; in subsequent training iterations, the optimizer directly applies the error gradient to the unbounded... Perform unconstrained free updates, and call in real time each time forward rasterization rendering is performed. This positive activation function will update the underlying layer. Dynamically and securely mapped back Actual opacity within the range .
[0122] To represent the specular and reflective characteristics of an object's surface as the viewing angle changes, the spherical harmonic function SH is used to encode the colors of the Gaussian elements. For the Gaussian elements... Assign a set of spherical harmonic order : 0th order represents the basic color, initialized to the midpoint of the sparse point cloud data. The built-in RGB color values represent pure diffuse colors; the higher-order coefficients are used to represent view-related details, initialized to zero, and activated and optimized in subsequent training iterations based on changes in spatial perspective.
[0123] In a specific embodiment, the rasterization rendering process is as follows:
[0124] The three-dimensional covariance matrix of each Gaussian element in the 3D Gaussian model is obtained by using a local affine approximation method. The projection is a two-dimensional covariance matrix. As shown in the following formula:
[0125] ;
[0126] in, The world-to-view transformation matrix for the current camera. Let T be the affine approximation Jacobian matrix of the projective transformation at the center of the Gaussian element, and let T denote the matrix transpose.
[0127] Based on the camera's view frustum, a culling operation is performed on all Gaussian primitives in the 3D Gaussian model. The center positions of the remaining Gaussian primitives after culling are then sorted from front to back according to their depth values in the camera coordinate system, generating an ordered list of Gaussian primitive indices. ;
[0128] According to Gaussian meta-index list The actual opacity corresponding to the opacity of each Gaussian element is calculated as follows:
[0129] ;
[0130] in, Represents the Gaussian meta-index list The first in The coordinates of the center position of each Gaussian element projected onto the two-dimensional projection area; Indicates the first Any pixel in a two-dimensional projection plane, represented by a Gaussian element. Represents the Gaussian meta-index list The first in The opacity of a Gaussian unit. Represents the Gaussian meta-index list The first in The actual opacity of the Gaussian unit;
[0131] By analyzing the Gaussian meta-index list The rendering color is obtained by performing alpha blending on each Gaussian element, as shown in the following formula:
[0132] ;
[0133] in, Indicates the rendered image at any pixel The rendered color at that location. Represents the Gaussian meta-index list The first in The actual opacity of the Gaussian unit. Represents the Gaussian meta-index list The first in The view-related color is calculated by the order of a spherical harmonic function for a given Gaussian element in the current camera pose.
[0134] Specifically, to achieve efficient training and view rendering of 3D Gaussian models, embodiments of this application construct an end-to-end differentiable rasterization pipeline. Specifically, for any given camera viewpoint, the differentiable rasterization pipeline is used to project a set of Gaussian primitives in space onto a two-dimensional projection plane to generate a rendered image. This process specifically includes:
[0135] Before projection, all Gaussian primitives are first culled based on the camera's view frustum, discarding those completely outside the view frustum or with extremely low opacity to save computational resources. Then, to correctly handle the occlusion relationships of semi-transparent objects, the center positions of the remaining Gaussian primitives are used... The depth values (Z-axis coordinates) in the camera coordinate system are sorted from front to back to generate an ordered list of Gaussian meta-indexes. .
[0136] To calculate Gaussian elements Given the shape and distribution on the two-dimensional projection plane from a given camera viewpoint, the three-dimensional covariance matrix is approximated using a local affine approximation method. The projection is a two-dimensional covariance matrix. For physical plausibility, after calculation... After that, the third row and third column are usually ignored, and a tiny diagonal term is added to act as a low-pass filter to prevent aliasing in screen space.
[0137] For any pixel on a two-dimensional projection plane Its final rendered color It is done by using an ordered list of Gaussian meta-indexes covering that pixel. This was obtained through Alpha fusion calculation.
[0138] The above-described projection and alpha mixing calculation process is fully differentiable. After each forward rendering and calculation of the mixing optimization loss function using the real image, the analytical gradient of the mixing optimization loss function with respect to the attribute parameters of each Gaussian unit is calculated using the chain rule. In particular, the frequency domain-aware loss introduced in the embodiments of this application will be backpropagated through the above rendering formula to accurately guide the updating and adaptive splitting of the attribute parameters of Gaussian units in 3D space.
[0139] A differentiable rasterizer is used to efficiently project a set of Gaussian primitives in 3D space onto a 2D imaging plane. First, by combining the camera's intrinsic and extrinsic parameters, the center positions of the Gaussian primitives within the camera's view frustum and the 3D covariance matrix are projected and transformed into a 2D Gaussian distribution on the 2D projection plane. Then, to accelerate rendering computation, the imaging plane is divided into multiple independent pixel blocks, and the Gaussian primitives intersecting each block are strictly ordered from front to back according to their depth information relative to the camera. For each pixel in the image, along the ray projection direction, an alpha blending mechanism is used to progressively accumulate and blend the color and opacity parameters of the overlapping Gaussian primitives ordered by depth, ultimately generating a 2D rendered image from the current viewpoint. The entire rasterization process maintains strict mathematical differentiability, ensuring that the subsequently generated error gradients can be effectively backpropagated to iteratively update the various property parameters of the Gaussian elements.
[0140] In a specific embodiment, the frequency domain feature extractor includes a two-dimensional discrete wavelet transform module, which uses Daubechies wavelet as the basis function and includes a low-pass filter and a high-pass filter.
[0141] The image is convolved and decomposed using a low-pass filter and a high-pass filter to obtain low-frequency components, horizontal high-frequency components, vertical high-frequency components, and diagonal high-frequency components, as shown in the following formula:
[0142] ;
[0143] ;
[0144] ;
[0145] ;
[0146] in, Represents an image. The function representing the low-pass filter, This represents the function corresponding to a high-pass filter. This represents the low-frequency components of the image. This represents the horizontal high-frequency components corresponding to the image. This represents the vertical high-frequency component corresponding to the image. This represents the diagonal high-frequency components of the image;
[0147] By concatenating the horizontal, vertical, and diagonal high-frequency components of the image, the high-frequency feature tensor corresponding to the image is obtained, as shown in the following equation:
[0148] ;
[0149] in, Indicates splicing, Representing an image The corresponding high-frequency feature tensor.
[0150] Specifically, in the embodiments of this application, a two-dimensional discrete wavelet transform (2D-DWT) module is introduced as a frequency domain feature extractor during the training iteration of the 3D Gaussian model, and Daubechies wavelet is selected as the basis function. A low-pass filter L and a high-pass filter H are defined. Four frequency band components are obtained using the low-pass filter L and the high-pass filter H: low-frequency component, horizontal high-frequency component, vertical high-frequency component, and diagonal high-frequency component. A high-frequency feature tensor is constructed by cascading the horizontal high-frequency component, vertical high-frequency component, and diagonal high-frequency component. This high-frequency feature tensor explicitly represents high-frequency edge information such as shoe sole texture and fabric texture. Simultaneously, the following is defined: To extract from this high-frequency feature tensor The high-frequency features of the texture direction corresponding to the texture parameter j extracted from the middle slice, where .
[0151] The rendered image generated at the current iteration step. And the corresponding ground truth image. Simultaneously, the input is fed into the 2D-DWT module. For each local region in the image, the wavelet transform decomposes it into the following four frequency band components: the low-frequency component (LL) representing the approximate contour information of the image, the horizontal high-frequency component (LH) responding to edge and texture changes in the horizontal direction, the vertical high-frequency component (HL) responding to edge and texture changes in the vertical direction, and the diagonal high-frequency component (HH) responding to detail changes in the diagonal direction.
[0152] Since minute defects typically manifest as abrupt changes in high-frequency signals, embodiments of this application discard the low-frequency component LL and extract three high-frequency components, LH, HL, and HH, which are then stacked along the channel dimension to construct a high-frequency feature tensor. For the rendered image, the corresponding high-frequency feature tensor is generated. For real images, generate corresponding high-frequency tensors. .
[0153] In a specific embodiment, the training process of the 3D Gaussian model is as follows:
[0154] The initial 3D Gaussian model is used as the first 3D Gaussian model in the first iteration cycle to start the iterative training process;
[0155] The 3D Gaussian model will be rasterized and rendered at each iteration step within the current iteration cycle to obtain the corresponding rendered image. Using images of defect-free footwear products as real images, and combining them with rendered images, spatial domain loss and structural similarity loss are constructed, as shown in the following equation:
[0156] ;
[0157] ;
[0158] in, Indicates spatial domain loss, Describing the L1 norm, Represents a real image; Represents structural similarity loss. Represents the structural similarity function;
[0159] The rendered image and the real image are respectively input into the frequency domain feature extractor to extract high-frequency feature tensors for different texture directions, and the frequency domain perceptual loss is calculated as shown in the following formula:
[0160] ;
[0161] Where j represents the texture parameters corresponding to different texture directions, including horizontal, vertical and diagonal directions, and j corresponds to LH, HL and HH respectively; This represents the high-frequency features of the high-frequency feature tensor corresponding to the rendered image along the texture direction under texture parameter j. This represents the high-frequency features of the high-frequency feature tensor corresponding to the real image in the texture direction under texture parameter j. The adaptive weights corresponding to texture parameter j;
[0162] A hybrid optimization objective loss function is constructed based on spatial domain loss, structural similarity loss, and frequency domain sensing loss, as shown in the following equation:
[0163] ;
[0164] in, Denotes the hybrid optimization objective loss function. and These represent similarity weights and frequency domain perception weights, respectively.
[0165] The adaptive compaction strategy based on frequency domain gradient is as follows:
[0166] The mean view space gradient norm of the center position of each Gaussian element in the Gaussian element index list is calculated by backpropagation in each iteration cycle, as shown in the following formula:
[0167] ;
[0168] in, Represents the first element in the Gaussian meta-index list. Mean of the view space gradient norm corresponding to each Gaussian element. This represents the total number of iterations within one iteration cycle. Indicates the first Number of iterations, Indicates the first The hybrid optimization objective loss function calculated by forward propagation at the number of iterations. Represents the first element in the Gaussian meta-index list. The central position of each Gaussian element In the The two-dimensional coordinate vector that falls into the two-dimensional projection plane after the projection transformation in the nth iteration step. Represents the L2 norm;
[0169] Calculate the frequency domain response factor within the effective projected pixel set corresponding to each Gaussian element in the Gaussian element index list. As shown in the following formula:
[0170] ;
[0171] in, Represents the first element in the Gaussian meta-index list. The frequency domain response factor corresponding to each Gaussian element. Represents the first element in the Gaussian meta-index list. The effective projected pixel set of a Gaussian element on a two-dimensional projection plane, where the effective projected pixel set is the set of pixels whose actual opacity exceeds an opacity threshold, is represented as: , Indicates the opacity threshold. This represents the total number of pixels in the pixel set. Represents a two-dimensional projection plane; This represents the high-frequency error map constructed based on the high-frequency feature tensors corresponding to the rendered image and the high-frequency feature tensors corresponding to the real image at any pixel. The high-frequency error value at that location, and the expression for the high-frequency error map are: ; Represents the first element in the Gaussian meta-index list. Each Gaussian unit corresponds to any pixel in a two-dimensional projection plane. The actual opacity at the location;
[0172] The corrected gradient for each Gaussian element is calculated based on the mean viewspace gradient norm and frequency domain response factor of each Gaussian element in the Gaussian element index list, as shown in the following formula:
[0173] ;
[0174] in, Represents the first element in the Gaussian meta-index list. The corrected gradient corresponding to each Gaussian element, where β is the upper limit of gain and γ is the sensitivity coefficient;
[0175] The Gaussian meta-index list is updated based on the corrected gradient, as follows:
[0176] In response to determining the first If the corrected gradient corresponding to the th Gaussian element is greater than the gradient threshold, and the maximum value of the element in the corresponding scaling matrix is greater than the scaling threshold, then for the th... Each Gaussian unit performs a splitting operation;
[0177] In response to determining the first If the corrected gradient corresponding to the th Gaussian element is greater than the gradient threshold, and the maximum value of the corresponding element in the scaling matrix is less than or equal to the scaling threshold, then for the th... Each Gaussian unit performs a cloning operation;
[0178] In response to determining the first If the corrected gradient corresponding to the nth Gaussian element is less than or equal to the gradient threshold, then the original nth Gaussian element is retained. A high-level unit;
[0179] The updated Gaussian meta-index list of the current iteration period is used as the Gaussian meta-index list of the next iteration period for forward and backward propagation. The updated Gaussian meta-index list of the last iteration period is used to construct a trained 3D Gaussian model.
[0180] Triggering a frequency-domain energy-driven dynamic adjustment strategy for the order of spherical harmonics in one iteration step within the current iteration cycle includes the following steps:
[0181] Calculate the high-frequency energy features corresponding to each pixel in the real image to extract the high-frequency energy feature map. As shown in the following formula:
[0182] ;
[0183] in, , and These represent the horizontal, vertical, and diagonal components of the real image at any pixel. The component value at the location; Represents any pixel in a real image High-frequency energy characteristics at the location;
[0184] The normalized high-frequency energy distribution probability of each Gaussian element within the effective projected pixel set is calculated as follows:
[0185] ;
[0186] in, Indicates the first Each Gaussian unit corresponds to any pixel in a two-dimensional projection plane. The normalized high-frequency energy distribution probability at that location. Represents any pixel in a real image High-frequency energy characteristics at that location Indicates the first Each Gaussian unit corresponds to any pixel in a two-dimensional projection plane. The actual opacity at that location, Represent one of the positive constants;
[0187] The local frequency domain entropy of each Gaussian element is calculated based on the normalized high-frequency energy distribution probability, as shown in the following equation:
[0188] ;
[0189] in, Indicates the first The local frequency domain entropy of a Gaussian element, Represent another positive number;
[0190] The adjusted order of the spherical harmonic function is determined based on the mapping relationship between local frequency domain entropy and the order of the spherical harmonic function, as shown in the following equation:
[0191] ;
[0192] in, This indicates the number of the current iteration step. The order of the adjusted spherical harmonic function corresponding to each Gaussian element. Low-frequency threshold; For high-frequency thresholds, Let the order of the maximum spherical harmonic function be the number of iterations in the current iteration step. The adjusted spherical harmonic function order corresponding to the _ Gaussian elements is used as the _th _ in the current preset iteration interval. The order of the spherical harmonic function corresponding to each Gaussian element is determined and propagated forward and backward.
[0193] For details, please refer to Figure 2 In order to overcome the traditional To address the high-frequency decay of loss during gradient backpropagation, and to balance overall structure recovery with the reconstruction of minute details, a frequency-domain weighted hybrid optimization loss function is constructed. An end-to-end training and optimization of an anisotropic 3D Gaussian model is performed. This hybrid optimization loss function consists of spatial domain loss, structural similarity loss, and frequency domain perceptual loss. The spatial domain loss ensures color accuracy, the structural similarity loss ensures structural integrity, and the frequency domain perceptual loss is defined as the weighted sum of L1 distances between the rendered image and the real image in the high-frequency subband. This is achieved by minimizing... Using the backpropagation mechanism, the gradient of the hybrid optimization loss function with respect to the attribute parameters of Gaussian elements (including center position, rotation matrix, scaling matrix, opacity, and SH order) is calculated.
[0194] In order to further improve the performance of the 3D Gaussian model in complex texture regions during training, two steps are introduced in the optimization loop: an adaptive densification strategy based on frequency domain gradient and a dynamic enhancement strategy of the spherical harmonic function order driven by frequency domain energy.
[0195] First, an adaptive compaction strategy based on frequency domain gradient is implemented using a hybrid optimization loss function. The spatial distribution of frequency domain error is used to apply a nonlinear gain to the gradient of the Gaussian elements, thereby guiding splitting and cloning. The specific process is as follows:
[0196] Within the preset iteration period Within, the actual center position of each Gaussian element is statistically analyzed. Mean of the gradient norm in the view space Specifically, the mean of the view space gradient norm is defined as the mean value of the gradient over a continuous period of time. In this optimization iteration, the hybrid optimization loss function The average cumulative L2 norm of the partial derivatives of the Gaussian element at the center position of the two-dimensional projection plane. This is used to plot the high-frequency error. Back projection is performed, and the frequency domain response factor of each Gaussian factor is calculated. A larger value indicates a greater loss of high-frequency details in the region covered by the 3D Gaussian model, necessitating densification. Since the support domain of the Gaussian function is theoretically infinite, in practical calculations, for engineering efficiency, it is... Defined as the actual opacity of the Gaussian element exceeding a set opacity threshold. The set of pixels. Its two-dimensional covariance matrix is typically used. of Confidence ellipse as truncation boundary to set opacity threshold .
[0197] Traditional methods rely solely on positional gradients for densification, easily neglecting texture details. Embodiments of this application introduce a frequency response factor, which uses a nonlinear mapping function to adjust the mean of the view space gradient norm. Perform weighted summation to generate corrected gradients. In this formula, Used to enhance amplitude; Used to control activation sensitivity; Used to normalize enhancement terms to a reasonable range.
[0198] Then use the corrected gradient Perform a split determination. If... And Gaussian scale Then for the first Each Gaussian unit performs a splitting operation. These represent the scaling ratios along the three principal axes of the scaling matrix; if and Then for the first Each Gaussian cell performs a cloning operation and updates the Gaussian cells in the Gaussian index list.
[0199] To improve detail while maintaining efficiency, the spherical harmonic function order is dynamically adjusted based on local texture complexity. In the frequency domain energy-driven dynamic improvement strategy for the spherical harmonic function order, a local frequency domain entropy is first defined to measure the complexity of the texture within the k-th Gaussian region. In the embodiments of this application, the Shannon entropy model from information theory is introduced, and the local frequency domain entropy is constructed based on the energy distribution of the wavelet high-frequency subbands. The specific calculation process is as follows: Extract the real image. High-frequency energy characteristic map For any pixel on a two-dimensional projection plane Its high-frequency energy characteristics are the superposition of the absolute values of the three high-frequency directional sub-bands, and then the calculation of the first... Each Gaussian element in its effective projected pixel set The normalized high-frequency energy distribution probability within the Gaussian-based projection region is calculated. This normalization operation transforms the energy intensity into a "probability distribution" in information theory. Finally, the local frequency domain entropy of this normalized high-frequency energy distribution probability is calculated. A larger local frequency domain entropy indicates a more uniform high-frequency energy distribution and denser, more complex texture details (such as the weave points of the shoe upper mesh) within the projection region. Conversely, a very small local frequency domain entropy indicates that the region is essentially a smooth surface or contains only isolated noise points. (Reference) Figure 3 Red represents high entropy (complex texture), and blue represents low entropy (smooth).
[0200] An order mapping rule is established using local frequency domain entropy, low-frequency threshold, and high-frequency threshold. This rule is then used to adjust the order of the spherical harmonic function in the current iteration cycle to obtain the order of the spherical harmonic function in the next iteration cycle. At lower levels, allocate low-level color shaders, model only the basic colors, and save video memory. When the value is higher, allocate higher-order SH (e.g., This is used to model complex view-dependent lighting effects.
[0201] S2. Acquire the image to be detected. Use a gradient descent iterative optimization algorithm based on photometric error to optimize the camera pose of the image to be detected, and obtain the optimized camera pose. Use the optimized camera pose to perform rasterization rendering on the trained 3D Gaussian model to obtain a standard reference image. Perform wavelet residual calculation and statistics based on the standard reference image and the image to be detected to obtain a wavelet residual map and an adaptive threshold. Use the adaptive threshold and the wavelet residual map to generate a defect mask. Detect the defect based on the defect mask to obtain the defect detection result.
[0202] In a specific embodiment, a gradient descent iterative optimization algorithm based on photometric error is used to optimize the camera pose of the image to be detected, resulting in an optimized camera pose, specifically including:
[0203] Using the current camera pose The trained 3D Gaussian model is rasterized to obtain the current rendered image. ;
[0204] Calculate the current rendered image With the image to be detected The photometric error loss between them is shown in the following formula:
[0205] ;
[0206] in, Indicates photometric error loss, These are the gradient weight coefficients. This represents the spatial gradient extracted from the current rendered image using the edge extraction operator. This represents the spatial gradient extracted from the image to be detected using the edge extraction operator. Describing the L1 norm, Represents the L2 norm;
[0207] The gradient of the photometric error loss with respect to the camera pose is calculated using a multi-level chain rule, as shown in the following equation:
[0208] ;
[0209] in, This represents the gradient of the photometric error loss with respect to the camera pose. To calculate the derivative of pixel-level error on a two-dimensional projection plane; To calculate the actual center position of the Gaussian pixel for pixel color changes The partial derivatives; The core projection Jacobian matrix connects 2D and 3D;
[0210] The camera pose is updated using the Adam optimizer. The process stops when the preset number of iterations is reached or the photometric error loss converges, resulting in the optimized camera pose. ;
[0211] Wavelet residual calculation and statistics are performed based on a standard reference image and the image to be detected to obtain a wavelet residual map and an adaptive threshold. A defect mask is generated using the adaptive threshold and the wavelet residual map. The defect detection results are obtained based on the defect mask detection, specifically including:
[0212] The standard reference image and the image to be detected are input into a low-pass filter and a high-pass filter for convolution decomposition to obtain the corresponding horizontal high-frequency components, vertical high-frequency components, and diagonal high-frequency components. The wavelet residual map is then calculated using the following formula:
[0213] ;
[0214] in, , and These represent the horizontal, vertical, and diagonal high-frequency components of the image to be detected. , and These represent the horizontal, vertical, and diagonal high-frequency components corresponding to the standard reference image, respectively. Represents the wavelet residual plot;
[0215] Calculate the mean of the wavelet residual plot and standard deviation The adaptive threshold is calculated using the following formula:
[0216] ;
[0217] in, Represents the determination coefficient. Indicates the defect determination threshold;
[0218] The defect mask is determined based on the wavelet residual image and the adaptive threshold, as shown in the following formula:
[0219] ;
[0220] in, Represents any pixel in the defect mask The corresponding mask value;
[0221] If the mask value is 1, the corresponding pixel in the image to be detected is determined to be a defect point; if the mask value is 0, the corresponding pixel in the image to be detected is determined to be a normal texture point; the connected region formed by the defect points is taken as the defect detection result.
[0222] Specifically, during the inference phase, a trained 3D Gaussian model is used as a "standard reference image generator," which is compared with the image to be tested to detect minute defects. To address the large number of "false positive" high-frequency residuals caused by minute mechanical positioning errors in the position of the product to be inspected in industrial assembly line scenarios, a pose optimization strategy is introduced. The specific steps are: freezing the trained 3D Gaussian model, and then... All attribute parameters Set to a non-differentiable state. Define the camera pose to be optimized. The camera pose is formed by concatenating the camera's rotation matrix and translation vector. This serves as a priori approximate pose for automated machine sensors or coarse-matching algorithms, representing the initial camera pose. This serves as the initial iteration starting point for the pose optimization strategy. Utilizing the current pose... Rasterization rendering is performed on the trained 3D Gaussian model to obtain the rendered image at the current pose. .calculate With real images Photometric error loss between This is used to measure the overall visual difference between the rendered image and the real image at the current pose.
[0223] Calculate photometric error loss Regarding camera pose gradient Specifically, Defined as the analytical gradient of the photometric error loss with respect to the camera extrinsic parameters, this gradient is calculated using a multi-level chain rule since the 3D Gaussian sputtering rasterization pipeline is fully differentiable. In this case, the gradient flow only propagates back to the camera parameters and does not update the Gaussian primitive's property parameters. The camera pose is updated using the Adam optimizer, which follows this formula: Update camera pose, where, Indicates the first Camera pose at the next iteration; This represents the result calculated after this gradient descent update. The camera pose in the next iteration; This represents the learning rate of the Adam optimizer. This indicates an iterative update (assignment) operation for parameters.
[0224] The process stops when the change in photometric error loss is less than a threshold or when the maximum number of iterations is reached, eventually converging to obtain the optimized camera pose. Using the optimized camera pose The rendered image generated by the trained 3D Gaussian model is used as the standard reference image. Figure 4 Through optimized training based on local frequency domain entropy and high-frequency residual energy, the rendered standard reference image not only achieves high-precision alignment with the real product in terms of macroscopic geometric contours, but also achieves extremely high-fidelity reproduction in terms of microscopic surface textures. Specifically, such as... Figure 4 As shown in the magnified area, the complex mesh weave of the shoe upper, the edge direction of the fabric fibers, and the sharp boundaries of the dark-colored trademark logo are all reconstructed with extremely clear detail, without the over-smoothing or blurring of details commonly seen in traditional 3D reconstruction methods. This high-precision reconstruction provides a perfect and reliable "flawless" filter for subsequent accurate extraction of spatial high-frequency residuals and detection of physical defects (such as minor scratches, foreign objects, and wear).
[0225] Optimized camera pose It utilizes the Adam optimizer based on gradient flow. For the initial camera pose After multiple rounds of iterative fine-tuning, stable camera pose parameters are extracted. Specifically, in the iterative alignment loop, when any of the following stopping conditions are met, the pose optimization process is considered to have converged, and the current camera pose is extracted. As :
[0226] Condition 1 (Error Minimization Convergence): The change in photometric error loss in the current iteration step. Less than the preset minimum tolerance threshold (e.g.) This indicates that the virtual camera has achieved precise alignment with the actual 3D physical position of the product under test on the industrial production line, and further iterations can no longer produce visual gain.
[0227] Condition 2 (Maximum computing power truncation): The number of iterations reaches the preset maximum iteration limit (e.g., 200 times). This mechanism aims to ensure the cycle time and real-time requirements of industrial online detection and avoid getting stuck in an infinite loop on abnormal samples.
[0228] Obtain Then, use the optimized camera pose. Rasterization rendering is performed on the 3D Gaussian model with frozen parameters, and the rendered image is the standard reference image. .because It absorbed mechanical positioning errors. Able to match the macroscopic outline with real images Achieve pixel-level seamless integration.
[0229] Using a standard reference image that is pixel-level aligned with the image to be detected. The defect detection process is as follows:
[0230] To accurately locate minute defects, instead of directly performing RGB difference, the image to be inspected is... Compared with standard reference diagram The images are input into a two-dimensional discrete wavelet transform module, and then pass through a low-pass filter L and a high-pass filter H to transform the image to be detected. Compared with standard reference diagram The data is then converted to the wavelet domain, and the differences between the two data points in the three high-frequency subbands (LH, HL, and HH) are calculated separately. These differences are then fused into a single-channel wavelet residual map. By directly superimposing the high-frequency differences in the three directions, the response signal to high-frequency defects such as minor scratches and foreign objects is significantly enhanced.
[0231] wavelet residual plot Perform statistical analysis and calculate the mean of its pixel values. and standard deviation To adapt to changes in lighting and materials, an adaptive threshold is employed. Using adaptive thresholds For residual plot Binarization segmentation is performed to obtain the defect mask. :
[0232] If one pixel in the wavelet residual map The corresponding residual value If a defect point is found, it is marked as 1; otherwise, it is marked as 0, indicating a normal texture point. Through this step, the algorithm can accurately segment the topological shape of physical defects such as broken yarn and wear from the continuous residual signal. Finally, the defect mask is... The corresponding connected regions are used as the defect detection results, and in the image to be detected. The defect location is outlined above, and the final visual inspection result is output. For example... Figure 5 As shown, the entire 3D Gaussian model of the tested footwear product is not composed of a traditional polygonal mesh, but rather a mixture of massive discrete and semi-transparent 3D Gaussian primitives. It is also evident that in areas with complex textures and sharp edges (such as shoelace edges and mesh areas), the Gaussian primitives exhibit smaller volumes and extremely dense distribution; while in areas with uniform color and smooth surfaces, the Gaussian primitives maintain a larger scale. This underlying structural distribution pattern fully verifies the effectiveness of the frequency-domain gradient-based adaptive densification strategy proposed in this invention, successfully and adaptively allocating limited computing power and Gaussian primitive parameters (such as center coordinates, covariance matrix, and higher-order spherical harmonic coefficients) to the areas with the densest high-frequency details. This explicit point cloud-level high-dimensional feature representation not only ensures exceptional rendering quality but also maintains the complete differentiability of the entire reconstruction pipeline, thereby supporting online high-precision reverse alignment of camera poses.
[0233] The key to this invention lies in applying multi-resolution wavelet transform and an adaptive densification strategy based on frequency domain gradients to a 3D Gaussian model. This allows for the precise capture of high-frequency texture details and subtle differences in the surface of an object in the frequency domain, thereby enhancing the model's ability to reconstruct and detect minute defects (such as scratches and foreign objects). Therefore, this invention can be widely applied to automated surface quality inspection systems in the footwear manufacturing industry.
[0234] Further reference Figure 6 As an implementation of the methods shown in the above figures, this application provides an embodiment of a footwear product defect detection device based on Gaussian splashing and wavelet transform. This device embodiment is similar to... Figure 1 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0235] This application provides a defect detection device for footwear products based on Gaussian splashing and wavelet transform, including:
[0236] Model reconstruction module 1 is configured to construct an anisotropic 3D Gaussian model and a frequency domain feature extractor. The frequency domain feature extractor is used to train the 3D Gaussian model to obtain a trained 3D Gaussian model. During the training process, a hybrid optimization objective loss function is adopted and an adaptive strategy is executed. The adaptive strategy includes an adaptive compaction strategy based on frequency domain gradient and a frequency domain energy-driven dynamic adjustment strategy for the order of spherical harmonic functions.
[0237] Detection module 2 is configured to acquire the image to be detected, optimize the camera pose of the image to be detected using a gradient descent iterative optimization algorithm based on photometric error, obtain the optimized camera pose, perform rasterization rendering on the trained 3D Gaussian model using the optimized camera pose to obtain a standard reference image, perform wavelet residual calculation and statistics based on the standard reference image and the image to be detected to obtain a wavelet residual map and an adaptive threshold, generate a defect mask using the adaptive threshold and the wavelet residual map, and obtain the defect detection result based on the defect mask detection.
[0238] Figure 7 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present invention. For example... Figure 7 As shown, the electronic device of this embodiment includes a processor 701 and a memory 702; wherein the memory 702 is used to store computer execution instructions; and the processor 701 is used to execute the computer execution instructions stored in the memory to implement the various steps performed by the electronic device in the above embodiment. For details, please refer to the relevant descriptions in the foregoing method embodiments.
[0239] Alternatively, the memory 702 can be either standalone or integrated with the processor 701.
[0240] When the memory 702 is set up independently, the electronic device also includes a bus 703 for connecting the memory 702 and the processor 701.
[0241] This invention also provides a computer storage medium storing computer execution instructions, which, when executed by processor 701, implement the above method.
[0242] This invention also provides a computer program product, including a computer program that, when executed by a processor 701, implements the above-described method.
[0243] In the embodiments provided by this invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or modules, and may be electrical, mechanical, or other forms.
[0244] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to implement the solution of this embodiment according to actual needs.
[0245] Furthermore, the functional modules in the various embodiments of this invention can be integrated into one processing unit, or each module can exist physically separately, or two or more modules can be integrated into one unit. The unit formed by the above modules can be implemented in hardware or in the form of hardware plus software functional units.
[0246] The integrated modules implemented as software functional modules described above can be stored in a computer-readable storage medium. These software functional modules, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor 701 to execute some steps of the methods of the various embodiments of this application.
[0247] It should be understood that the processor 701 described above can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor, or the processor 701 can be any conventional processor 701. The steps of the method disclosed in this invention can be directly manifested as the hardware processor 701 executing the steps, or as a combination of hardware and software modules within the processor 701 executing the steps.
[0248] The memory 702 may include high-speed RAM memory, and may also include non-volatile memory NVM, such as at least one disk storage device, and may also be a USB flash drive, portable hard drive, read-only memory, disk or optical disc, etc.
[0249] Bus 703 can be an Industry Standard Architecture (ISA), a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Bus 703 can be divided into address bus, data bus, control bus, etc. For ease of illustration, the bus 703 in the accompanying drawings of this application is not limited to only one bus 703 or one type of bus 703.
[0250] The aforementioned storage medium can be implemented from any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The storage medium can be any available medium accessible to general-purpose or special-purpose computers.
[0251] An exemplary storage medium is coupled to a processor 701, enabling the processor 701 to read information from and write information to the storage medium. Alternatively, the storage medium can be an integral part of the processor 701. The processor 701 and the storage medium can reside in an application-specific integrated circuit (ASIC). Alternatively, the processor 701 and the storage medium can exist as discrete components in an electronic device or a host device.
[0252] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0253] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for detecting defects in footwear products based on Gaussian splashing and wavelet transform, characterized in that, Includes the following steps: An anisotropic 3D Gaussian model and a frequency domain feature extractor are constructed. The 3D Gaussian model is trained using the frequency domain feature extractor to obtain a trained 3D Gaussian model. During the training process, a hybrid optimization objective loss function is adopted and an adaptive strategy is executed. The adaptive strategy includes an adaptive compaction strategy based on frequency domain gradient and a frequency domain energy-driven dynamic adjustment strategy for the order of spherical harmonic functions. The image to be detected is acquired, and the camera pose of the image to be detected is optimized using a gradient descent iterative optimization algorithm based on photometric error to obtain the optimized camera pose. The optimized camera pose is then used to perform rasterization rendering on a trained 3D Gaussian model to obtain a standard reference image. Wavelet residual calculation and statistics are performed based on the standard reference image and the image to be detected to obtain a wavelet residual map and an adaptive threshold. A defect mask is generated using the adaptive threshold and the wavelet residual map, and the defect detection result is obtained based on the defect mask.
2. The method for detecting defects in footwear products based on Gaussian splashing and wavelet transform according to claim 1, characterized in that, The construction process of the 3D Gaussian model is as follows: Images of defect-free footwear products acquired under different camera poses are obtained and sparse point cloud data is generated. An initial Gaussian primitive is initialized for each point in the sparse point cloud data. Each Gaussian primitive contains a set of learnable attribute parameters. For the first point in the sparse point cloud data... The corresponding point Gao Siyuan The attribute parameters include center position. Opacity , order of spherical harmonic function and the three-dimensional covariance matrix The three-dimensional covariance matrix is decomposed into a rotation matrix. and scaling matrix As shown in the following formula: ; Where T represents the transpose of the matrix; All Gaussian elements are used to construct a 3D Gaussian model.
3. The method for detecting defects in footwear products based on Gaussian splashing and wavelet transform according to claim 2, characterized in that, The frequency domain feature extractor includes a two-dimensional discrete wavelet transform module, which uses Daubechies wavelet as the basis function and includes a low-pass filter and a high-pass filter. The image is convolved and decomposed using a low-pass filter and a high-pass filter to obtain low-frequency components, horizontal high-frequency components, vertical high-frequency components, and diagonal high-frequency components, as shown in the following formula: ; ; ; ; in, Represents an image. The function representing the low-pass filter, This represents the function corresponding to a high-pass filter. This represents the low-frequency components of the image. This represents the horizontal high-frequency components corresponding to the image. This represents the vertical high-frequency component corresponding to the image. This represents the diagonal high-frequency components of the image; The horizontal, vertical, and diagonal high-frequency components of the image are concatenated to obtain the high-frequency feature tensor of the image, as shown in the following equation: ; in, Indicates splicing, Representing an image The corresponding high-frequency feature tensor.
4. The method for detecting defects in footwear products based on Gaussian splashing and wavelet transform according to claim 2, characterized in that, The rasterization rendering process is as follows: The three-dimensional covariance matrix of each Gaussian element in the 3D Gaussian model is obtained by using a local affine approximation method. The projection is a two-dimensional covariance matrix. As shown in the following formula: ; in, The world-to-view transformation matrix for the current camera. Let T be the affine approximation Jacobian matrix of the projective transformation at the center of the Gaussian element, and let T denote the matrix transpose. Based on the camera's view frustum, a culling operation is performed on all Gaussian elements in the 3D Gaussian model. The center positions of the remaining Gaussian elements after culling are then sorted from front to back according to their depth values in the camera coordinate system, generating an ordered list of Gaussian element indices. ; According to Gaussian meta-index list The actual opacity corresponding to the opacity of each Gaussian element is calculated as follows: ; in, Represents the Gaussian meta-index list The first in The coordinates of the center position of each Gaussian element projected onto the two-dimensional projection area; Indicates the first Any pixel in a two-dimensional projection plane, represented by a Gaussian element. Represents the Gaussian meta-index list The first in The opacity of a Gaussian unit. Represents the Gaussian meta-index list The first in The actual opacity of the Gaussian unit; By analyzing the Gaussian meta-index list The rendering color is obtained by performing alpha blending on each Gaussian element, as shown in the following formula: ; in, Indicates the rendered image at any pixel The rendered color at that location. Represents the Gaussian meta-index list The first in The actual opacity of the Gaussian unit. Represents the Gaussian meta-index list The first in The view-related color is calculated by the order of a spherical harmonic function for a given Gaussian element in the current camera pose.
5. The method for detecting defects in footwear products based on Gaussian splashing and wavelet transform according to claim 4, characterized in that, The training process of the 3D Gaussian model is as follows: The initial 3D Gaussian model is used as the first 3D Gaussian model in the first iteration cycle to start the iterative training process; The 3D Gaussian model will be rasterized and rendered at each iteration step within the current iteration cycle to obtain the corresponding rendered image. The image of the defect-free footwear product is used as the real image and combined with the rendered image to construct spatial domain loss and structural similarity loss, as shown in the following formula: ; ; in, Indicates spatial domain loss, Describing the L1 norm, Represents a real image; Represents structural similarity loss. Represents the structural similarity function; The rendered image and the real image are respectively input into the frequency domain feature extractor to extract high-frequency feature tensors for different texture directions, and the frequency domain perceptual loss is calculated as shown in the following formula: ; Where j represents the texture parameters corresponding to different texture directions, including horizontal, vertical and diagonal directions, and j corresponds to LH, HL and HH respectively; This represents the high-frequency features of the high-frequency feature tensor corresponding to the rendered image in the texture direction under texture parameter j. This represents the high-frequency features of the high-frequency feature tensor corresponding to the real image in the texture direction under texture parameter j. The adaptive weights corresponding to texture parameter j; Based on the spatial domain loss, structural similarity loss, and frequency domain sensing loss, a hybrid optimization objective loss function is constructed, as shown in the following equation: ; in, Denotes the hybrid optimization objective loss function. and These represent similarity weights and frequency domain perception weights, respectively. The adaptive compaction strategy based on frequency domain gradient is as follows: The mean view space gradient norm of the center position of each Gaussian element in the Gaussian element index list is calculated by backpropagation in each iteration period, as shown in the following formula: ; in, Represents the first in the Gaussian meta-index list Mean of the view space gradient norm corresponding to each Gaussian element. This represents the total number of iterations within one iteration cycle. Indicates the first Number of iterations, Indicates the first The hybrid optimization objective loss function calculated by forward propagation at the number of iterations. Represents the first in the Gaussian meta-index list The central position of each Gaussian element In the The two-dimensional coordinate vector that falls into the two-dimensional projection plane after the projection transformation in the nth iteration step. Represents the L2 norm; Calculate the frequency domain response factor within the effective projected pixel set corresponding to each Gaussian in the Gaussian index list. As shown in the following formula: ; in, Represents the first in the Gaussian meta-index list The frequency domain response factor corresponding to each Gaussian element. Represents the first in the Gaussian meta-index list The effective projected pixel set of a Gaussian element on a two-dimensional projection plane, wherein the effective projected pixel set is a set of pixels whose actual opacity exceeds an opacity threshold, denoted as: , Indicates the opacity threshold. This represents the total number of pixels in the pixel set. Represents a two-dimensional projection plane; This indicates that the high-frequency error map constructed based on the high-frequency feature tensor corresponding to the rendered image and the high-frequency feature tensor corresponding to the real image is at any pixel. The high-frequency error value at the specified location, and the expression for the high-frequency error map are: ; Represents the first in the Gaussian meta-index list Each Gaussian unit corresponds to any pixel in a two-dimensional projection plane. The actual opacity at the location; The corrected gradient for each Gaussian element is calculated based on the mean viewspace gradient norm and frequency domain response factor corresponding to each Gaussian element in the Gaussian element index list, as shown in the following formula: ; in, Represents the first in the Gaussian meta-index list The corrected gradient corresponding to each Gaussian element, where β is the upper limit of gain and γ is the sensitivity coefficient; The Gaussian meta-index list is updated based on the corrected gradient, as follows: In response to determining the first If the corrected gradient corresponding to the th Gaussian element is greater than the gradient threshold, and the maximum value of the element in the corresponding scaling matrix is greater than the scaling threshold, then for the th... Each Gaussian unit performs a splitting operation; In response to determining the first If the corrected gradient corresponding to the th Gaussian element is greater than the gradient threshold, and the maximum value of the corresponding element in the scaling matrix is less than or equal to the scaling threshold, then for the th... Each Gaussian unit performs a cloning operation; In response to determining the first If the corrected gradient corresponding to the nth Gaussian element is less than or equal to the gradient threshold, then the original nth Gaussian element is retained. A high-level unit; The updated Gaussian meta-index list of the current iteration period is used as the Gaussian meta-index list of the next iteration period for forward and backward propagation. The updated Gaussian meta-index list of the last iteration period is used to construct a trained 3D Gaussian model. Triggering a frequency-domain energy-driven dynamic adjustment strategy for the order of spherical harmonics in one iteration step within the current iteration cycle includes the following steps: Calculate the high-frequency energy features corresponding to each pixel in the real image, and extract the high-frequency energy feature map. As shown in the following formula: ; in, , and These represent the horizontal, vertical, and diagonal components of the real image at any pixel. The component value at the location; Represents any pixel in the real image High-frequency energy characteristics at the location; The normalized high-frequency energy distribution probability of each Gaussian element within the set of effective projected pixels is calculated as follows: ; in, Indicates the first Each Gaussian unit corresponds to any pixel in a two-dimensional projection plane. The normalized high-frequency energy distribution probability at that location. Represents any pixel in the real image High-frequency energy characteristics at that location Indicates the first Each Gaussian unit corresponds to any pixel in a two-dimensional projection plane. The actual opacity at that location, Represent one of the positive constants; The local frequency domain entropy of each Gaussian element is calculated based on the normalized high-frequency energy distribution probability, as shown in the following equation: ; in, Indicates the first The local frequency domain entropy of a Gaussian element, Represent another positive number; The adjusted spherical harmonic function order is determined based on the mapping relationship between the local frequency domain entropy and the spherical harmonic function order, as shown in the following formula: ; in, This indicates the number of the current iteration step. The order of the adjusted spherical harmonic function corresponding to each Gaussian element. Low-frequency threshold; For high-frequency thresholds, Let the order of the maximum spherical harmonic function be the number of iterations in one of them. The adjusted spherical harmonic function order corresponding to the _ Gaussian elements is used as the _th _ in the current preset iteration interval. The order of the spherical harmonic function corresponding to each Gaussian element is determined and propagated forward and backward.
6. The method for detecting defects in footwear products based on Gaussian splashing and wavelet transform according to claim 3, characterized in that, An iterative optimization algorithm based on photometric error gradient descent is used to optimize the camera pose of the image to be detected, resulting in an optimized camera pose, specifically including: Using the current camera pose The trained 3D Gaussian model is rasterized to obtain the current rendered image. ; Calculate the current rendered image With the image to be detected The photometric error loss between them is shown in the following formula: ; in, Indicates photometric error loss, These are the gradient weight coefficients. This represents the spatial gradient extracted from the current rendered image using the edge extraction operator. This represents the spatial gradient extracted from the image to be detected using the edge extraction operator. Describing the L1 norm, Represents the L2 norm; The gradient of the photometric error loss with respect to the camera pose is calculated using a multi-level chain rule, as shown in the following equation: ; in, This represents the gradient of the photometric error loss with respect to the camera pose. To calculate the derivative of pixel-level error on a two-dimensional projection plane; To calculate the actual center position of the Gaussian pixel for pixel color changes The partial derivatives; The core projection Jacobian matrix connecting 2D and 3D; The camera pose is updated using the Adam optimizer. The process stops when the preset number of iterations is reached or the photometric error loss converges, resulting in the optimized camera pose. ; Based on the standard reference image and the image to be detected, wavelet residual calculation and statistics are performed to obtain a wavelet residual map and an adaptive threshold. A defect mask is generated using the adaptive threshold and the wavelet residual map. Based on the defect mask, the defect detection result is obtained, specifically including: The standard reference image and the image to be detected are input into a low-pass filter and a high-pass filter for convolution decomposition to obtain the corresponding horizontal high-frequency components, vertical high-frequency components, and diagonal high-frequency components. The wavelet residual image is then calculated using the following formula: ; in, , and These are the horizontal high-frequency component, vertical high-frequency component, and diagonal high-frequency component corresponding to the image to be detected. , and These are the horizontal high-frequency components, vertical high-frequency components, and diagonal high-frequency components corresponding to the standard reference image, respectively. Represents the wavelet residual plot; Calculate the mean of the wavelet residual plot and standard deviation The adaptive threshold is calculated using the following formula: ; in, Represents the determination coefficient. Indicates the defect determination threshold; The defect mask is determined based on the wavelet residual image and the adaptive threshold, as shown in the following formula: ; in, Represents any pixel in the defect mask. The corresponding mask value; If the mask value is 1, the corresponding pixel in the image to be detected is determined to be a defect point; if the mask value is 0, the corresponding pixel in the image to be detected is determined to be a normal texture point; the connected region formed by the defect points is taken as the defect detection result.
7. A defect detection device for footwear products based on Gaussian splashing and wavelet transform, characterized in that, include: The model reconstruction module is configured to construct an anisotropic 3D Gaussian model and a frequency domain feature extractor, train the 3D Gaussian model using the frequency domain feature extractor to obtain a trained 3D Gaussian model, and employ a hybrid optimization objective loss function and execute an adaptive strategy during the training process. The adaptive strategy includes an adaptive compaction strategy based on frequency domain gradient and a frequency domain energy-driven dynamic adjustment strategy for the order of spherical harmonic functions. The detection module is configured to acquire the image to be detected, optimize the camera pose of the image to be detected using a gradient descent iterative optimization algorithm based on photometric error, obtain the optimized camera pose, perform rasterization rendering on the trained 3D Gaussian model using the optimized camera pose to obtain a standard reference image, perform wavelet residual calculation and statistics based on the standard reference image and the image to be detected to obtain a wavelet residual map and an adaptive threshold, generate a defect mask using the adaptive threshold and the wavelet residual map, and obtain the defect detection result based on the defect mask.
8. An electronic device, comprising: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-6.