Nursing pad printing toilet paper positioning method and system based on machine vision
By using multi-source sensor fusion technology and visual feature matching, the problem of inaccurate visual positioning of printed sanitary paper for nursing pads in complex industrial environments has been solved, achieving high-precision positioning under lighting and vibration conditions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANDONG AISHULE HYGIENE PROD CO LTD
- Filing Date
- 2025-12-24
- Publication Date
- 2026-05-08
AI Technical Summary
In complex industrial settings, uneven ambient light, flickering, and equipment vibrations reduce the visual positioning accuracy and stability of printed sanitary napkins. Traditional methods struggle to achieve accurate identification and positioning under varying lighting conditions and mechanical vibrations.
By simultaneously capturing visual, light interference, and vibration information from multiple sensors, an enhanced visual feature image is generated. The grayscale gradient and texture continuity of the printed pattern are used for region expansion, geometric and spectral descriptors are extracted, and matching and spatial transformation calibration are performed to ultimately generate a precise positioning decision.
It improves the robustness and adaptability of the positioning system in dynamic environments, reduces the false judgment rate, and achieves high-precision matching and recognition under complex interference.
Smart Images

Figure CN121999194A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of industrial machine vision technology, specifically to a machine vision-based method and system for positioning printed sanitary napkins and toilet paper. Background Technology
[0002] In automated production lines for hygiene products, rapid and accurate visual positioning of printed sanitary napkins is crucial for subsequent automated gripping, cutting, or packaging. Current mainstream methods involve using industrial cameras to capture two-dimensional images from a single perspective, followed by edge detection, template matching, or color threshold-based image segmentation techniques to identify and locate the printed areas. These methods can achieve certain positioning results under controlled, stable lighting and ideal experimental conditions with minimal equipment vibration.
[0003] In complex, high-speed industrial environments, ambient light is uneven, flickering, or randomly changing, and the continuous operation of production equipment introduces significant high-frequency or low-frequency mechanical vibrations. These factors collectively affect image sensors, resulting in blurred, noisy, or locally overexposed / underexposed raw images, severely interfering with the stability and accuracy of subsequent feature extraction. Furthermore, for different products with similar textures and colors, or when the contrast between the printed pattern and the background is low, relying solely on traditional geometric or color features for matching is prone to mismatches or positioning drift, limiting the flexibility and reliability of the production line. A positioning technology is needed that can overcome dynamic environmental interference and utilize richer features for accurate identification. Summary of the Invention
[0004] The purpose of this invention is to provide a machine vision-based method and system for positioning printed sanitary paper in nursing pads, in order to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, the present invention provides a machine vision-based method for positioning printed sanitary napkins, the method comprising:
[0006] Through a multi-source sensing environmental perception process, the original visual field information stream, ambient light interference field information stream, and equipment vibration disturbance information stream containing the target nursing pad printed toilet paper are captured simultaneously to generate an enhanced visual feature image.
[0007] Seed point diffusion growth is performed on the printed pattern area on the enhanced visual feature image. The seed point diffusion growth expands the region according to the gray-scale gradient and texture continuity constraints of the printed pattern until it reaches the preset growth termination boundary, thereby outlining the candidate positioning area.
[0008] Extract the geometric shape description subset and spectral reflectance description subset of the candidate positioning region, and match them item by item with the standard nursing pad print shape spectral archive pre-stored in the reference template library;
[0009] Based on the matching results, calculate the spatial transformation parameters between the candidate positioning region and at least one standard nursing pad print morphology spectral file;
[0010] The coordinate frame of the candidate positioning area is rigidly adjusted by applying the spatial transformation parameters so that the candidate positioning area is spatially aligned with the selected standard nursing pad print morphology spectral file.
[0011] Within the aligned coordinate frame, the centroid coordinates of the candidate positioning region are calculated, and the centroid coordinates are mapped back to the original captured visual field space to generate the final positioning decision for the printed toilet paper pad.
[0012] Preferably, the process of simultaneously capturing the original visual field information stream, ambient light interference field information stream, and equipment vibration disturbance information stream of the target nursing pad printed toilet paper through multi-source sensing environmental perception includes:
[0013] The original visual field information stream is subjected to a spatial-spectral joint decomposition operation to extract the core visual feature bands related to the printed pattern and the background visual feature bands related to the toilet paper substrate.
[0014] The ambient light interference field information stream and the device vibration disturbance information stream are fused to generate a dynamic compensation vector field, and the dynamic compensation vector field is used to perform reverse disturbance calibration on the core visual feature band.
[0015] The core visual feature bands, which have undergone inverse perturbation calibration, are reprojected and synthesized with the background visual feature bands to generate an enhanced visual feature image.
[0016] The synchronization capture includes:
[0017] Deploy a composite vision sensor head that integrates a wide-angle microscope lens and a ring structured light projector, the composite vision sensor head scanning the printed sanitary napkins on the production line at a fixed frequency;
[0018] The ring structured light projector is activated, projecting a grid-shaped light spot pattern with specific encoded information onto the surface of the printed toilet paper pad during scanning.
[0019] The wide-angle microscope lens synchronously receives light spot pattern images reflecting back from the surface of the printed toilet paper of the nursing pad, which carry surface deformation information, and the light spot pattern images constitute the original visual field information stream;
[0020] An ambient light meter and a triaxial micro-vibration sensor are installed on the side of the composite vision sensor head to continuously record the light intensity change data and the equipment vibration acceleration data during the scanning process, respectively, forming the ambient light interference field information stream and the equipment vibration disturbance information stream.
[0021] Preferably, the spatial-spectral joint decomposition operation on the original visual field information stream includes:
[0022] The continuous multiple frames of images in the original visual field information stream are converted to the frequency domain, and the spectrum of each frame of the image is obtained by Fourier transform.
[0023] The periodic high-frequency components caused by the grid-like light spot pattern and the characteristic mid-to-low-frequency components caused by the printing pattern and base texture of the printed toilet paper pad were identified in the spectrum.
[0024] Design a set of bandpass digital filters, use the bandpass digital filters to separate and extract the characteristic low-frequency components, and reconstruct the core visual feature bands after inverse Fourier transform of the characteristic low-frequency components.
[0025] The image data corresponding to the core visual feature band is subtracted from the original visual field information stream, and the remaining part is processed by removing the periodic high-frequency component noise to form the background visual feature band.
[0026] Preferably, generating a dynamic compensation vector field includes:
[0027] Time series analysis was performed on the ambient light interference field information stream to extract its low-frequency trend component that fluctuates over time;
[0028] Frequency analysis is performed on the vibration disturbance information stream of the equipment to identify its main vibration frequency components and corresponding amplitudes;
[0029] A two-dimensional mapping model is established with time and image pixel position as independent variables. The low-frequency trend component, the main vibration frequency component and amplitude are used as input parameters to calculate the expected position offset and brightness fluctuation of each pixel at each time point due to ambient light and vibration.
[0030] The expected position offset and the brightness fluctuation together constitute the dynamic compensation vector field, and each vector in the dynamic compensation vector field contains direction information and intensity information.
[0031] Preferably, the step of performing inverse perturbation calibration on the core visual feature band using the dynamic compensation vector field includes:
[0032] Read the time and position information corresponding to each pixel of the core visual feature band in the dynamic compensation vector field;
[0033] Based on the direction and intensity information provided by the dynamic compensation vector field, a reverse translation interpolation operation is performed on the pixel position of the core visual feature band in the spatial domain to offset the expected position offset.
[0034] Simultaneously, based on the brightness fluctuation provided by the dynamic compensation vector field, the pixel grayscale values of the core visual feature bands are adjusted compensatorily.
[0035] After processing all pixels, the core visual feature bands with corrected position and brightness are output.
[0036] Preferably, the seed point diffusion growth of the printed pattern region on the enhanced visual feature image includes:
[0037] On the enhanced visual feature image, a local gray-level variance maximum point detection method is used to automatically select at least one point as an initial seed point;
[0038] Centered on the initial seed point, calculate the grayscale gradient values of the pixels in its eight neighboring areas and ensure consistency with the texture direction;
[0039] A dynamic growth threshold is set, which is determined by the average gradient of the edge pixels of the current growth region;
[0040] Pixels whose gray-level gradient values in their eight neighborhoods are less than the dynamic growth threshold and whose texture direction consistency is higher than the preset consistency threshold are included in the current growth region, and the edges of the growth region are updated.
[0041] The process of iteratively exploring the neighborhood, comparing thresholds, and updating the region continues until no new pixels are included in several consecutive iterations, or the edge of the growing region touches the physical boundary of the image. At this point, it is determined that the growth termination boundary has been reached.
[0042] The entire pixel region covered by the point where growth eventually stops is marked as the candidate localization region.
[0043] Preferably, the extraction of the geometric shape descriptor subset and spectral reflectance descriptor subset of the candidate localization region includes:
[0044] Contour tracking is performed on the candidate positioning region to obtain its closed boundary curve. The moment invariant, Fourier descriptor and curvature statistical features of the boundary curve are calculated and combined into the geometric shape descriptor subset.
[0045] On the enhanced visual feature image, for the pixels in the candidate localization region, the gray-level distribution of the pixels in multiple narrow bands is analyzed, and the mean, variance and higher-order moments of the spectral reflectance are calculated to form the spectral reflectance descriptor set.
[0046] Preferably, the step of matching it item by item with the standard nursing pad print morphology spectral file pre-stored in the reference template library includes:
[0047] The geometric shape descriptor subset and the spectral reflectance descriptor subset are encapsulated into a feature vector;
[0048] Calculate the Mahalanobis distance between the feature vector and the standard feature vector corresponding to each standard nursing pad print morphology spectral file in the reference template library;
[0049] The top few standard nursing pad print morphology spectral files with the smallest Mahalanobis distance were selected as matching candidates;
[0050] Using the iterative nearest point algorithm, the correspondence between local feature points between the candidate positioning region and each matching candidate is further calculated, and the optimal matching standard nursing pad print morphology spectral profile is selected.
[0051] Preferably, the step of calculating the spatial transformation parameters between the candidate positioning region and at least one standard nursing pad print morphology spectral profile based on the matching results includes:
[0052] The spatial transformation parameters include rotation, translation, and scale factor;
[0053] Based on the correspondence between local feature points between the candidate positioning region and the optimal matching standard nursing pad print morphology spectral file, the least squares estimation method is used to solve for the optimal rigid body transformation matrix.
[0054] The rotation about an axis perpendicular to the image plane, the translation in the image plane, and the overall scaling factor are decomposed from the rigid body transformation matrix.
[0055] Preferably, the present invention also includes a machine vision-based positioning system for printed sanitary pads, the system including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor, when executing the computer program, implements the steps of the machine vision-based positioning method for printed sanitary pads as described above.
[0056] Compared with the prior art, the beneficial effects of the present invention are:
[0057] By simultaneously acquiring and fusing raw visual information, ambient light interference, and equipment vibration disturbance information, a multi-source complementary environmental perception model was constructed. This design enables real-time modeling and collaborative compensation of complex lighting fluctuations and mechanical vibrations in the production line environment, suppressing interference such as image blurring, grayscale anomalies, and geometric distortions. This improves the quality and consistency of the input images, providing a stable and reliable visual data foundation for subsequent processing steps, and enabling the system to possess robustness and adaptability in dynamically changing real industrial scenarios that are difficult to achieve with traditional single-vision methods.
[0058] This method extracts and utilizes a subset of spectral reflectance descriptors from printed patterns, overcoming the limitations of conventional visual positioning that relies solely on shape and color features. By analyzing the reflectance characteristics of targets in specific spectral bands, it obtains more physically unique and environmentally invariant feature dimensions. In the matching stage, the extracted spectral reflectance features are compared in a high-dimensional manner with pre-stored standardized spectral files. Even when faced with challenges such as similar patterns, complex backgrounds, or significant changes in lighting conditions, accurate identification and high-precision matching can be achieved. This enhances the positioning system's ability to distinguish and improve decision-making reliability under complex interference, while reducing the false positive rate. Attached Figure Description
[0059] Figure 1 This is a schematic diagram illustrating the working principle of the machine vision-based method for positioning printed toilet paper in nursing pads according to the present invention.
[0060] Figure 2 A flowchart of the space-spectrum joint decomposition operation;
[0061] Figure 3 A flowchart for generating a dynamic compensation vector field;
[0062] Figure 4 A distribution analysis diagram of multispectral channel spectral reflectance descriptors for candidate localization regions;
[0063] Figure 5 A comparison chart of candidate template space transformation parameters. Detailed Implementation
[0064] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0065] Please see Figure 1This invention provides a machine vision-based method for locating printed sanitary napkins. The method includes: simultaneously capturing an original visual field information stream, an ambient light interference field information stream, and a device vibration disturbance information stream containing the target printed sanitary napkin; these information streams are fused to generate an enhanced visual feature image; seed point diffusion growth is performed on the enhanced visual feature image to expand the printed pattern region, strictly adhering to the grayscale gradient and texture continuity constraints of the printed pattern until a preset growth termination boundary is reached, thereby outlining a candidate positioning region; a geometric shape descriptor subset and a spectral reflectance descriptor subset of the candidate positioning region are extracted, and this descriptor subset is matched item by item with a standard sanitary napkin printed shape spectral profile pre-stored in a reference template library; spatial transformation parameters between the candidate positioning region and at least one standard sanitary napkin printed shape spectral profile are calculated based on the matching results; and the coordinate frame of the candidate positioning region is rigidly adjusted using the spatial transformation parameters to spatially align the candidate positioning region with the selected standard sanitary napkin printed shape spectral profile. The centroid coordinates of the candidate positioning region are calculated within the aligned coordinate frame, and the centroid coordinates are mapped back to the original captured visual field space to generate the final positioning decision for the printed toilet paper pad.
[0066] Example 1: A composite vision sensor head integrating a wide-angle microscope lens and a ring structured light projector is deployed. This composite vision sensor head scans printed sanitary napkins on the production line at a fixed frequency. The ring structured light projector is activated, projecting a grid-like light spot pattern with specific coded information onto the surface of the printed sanitary napkins at the instant of scanning. The wide-angle microscope lens simultaneously receives images of the light spot pattern reflected from the surface of the printed sanitary napkins, carrying surface deformation information. These light spot pattern images constitute the original visual field information stream. An ambient photometer and a triaxial micro-vibration sensor are installed laterally on the composite vision sensor head to continuously record data on changes in illumination intensity and equipment vibration acceleration during the scanning process, respectively, forming the ambient light interference field information stream and the equipment vibration disturbance information stream. A spatial-spectral joint decomposition operation is performed on the original visual field information stream to extract the core visual feature bands related to the printed pattern and the background visual feature bands related to the sanitary napkin substrate. The ambient light interference field information stream and the device vibration disturbance information stream are fused to generate a dynamic compensation vector field, which is then used to perform inverse perturbation calibration on the core visual feature bands. The core visual feature bands that have undergone inverse perturbation calibration are then reprojected and synthesized with the background visual feature bands to generate an enhanced visual feature image.
[0067] In its implementation, a machine vision-based method for positioning printed sanitary napkins utilizes a multi-source sensing environment perception process achieved by deploying a composite vision sensor head integrating a wide-angle microscope and a ring structured light projector. This composite vision sensor head scans the printed sanitary napkins on the production line at a fixed frequency. The ring structured light projector is activated, projecting a grid-like light spot pattern with specific coded information onto the surface of the printed sanitary napkin at the instant of scanning. Simultaneously, the wide-angle microscope receives images of the light spot pattern, carrying surface deformation information, reflected back from the surface of the printed sanitary napkin. These light spot pattern images constitute the original visual field information stream. An ambient photometer and a triaxial micro-vibration sensor are installed laterally to the composite vision sensor head. The ambient photometer and the triaxial micro-vibration sensor continuously record data on changes in light intensity and equipment vibration acceleration during the scanning process, respectively, forming an ambient light interference field information stream and an equipment vibration disturbance information stream. In some embodiments, the fixed scanning frequency of the composite vision sensor head is set to 120 Hz, the grid pattern of light spots projected by the ring structured light projector is encoded with sinusoidal stripe phase, the sampling frequency of the ambient photometer is 1 kHz, and the sampling frequency of the triaxial micro-vibration sensor is 2 kHz, thereby achieving synchronous acquisition of multi-source information.
[0068] In specific implementation, the spatial-spectral joint decomposition operation on the original visual field information stream includes inputting the captured continuous image sequence into the processing unit, and extracting the core visual feature bands related to the printed pattern and the background visual feature bands related to the toilet paper substrate in the frequency domain. The process of fusing the ambient light interference field information stream and the equipment vibration disturbance information stream is based on a physical model to generate a dynamic compensation vector field, which is used to perform inverse perturbation calibration on the core visual feature bands. It can be understood that the inverse perturbation calibration operation is performed pixel by pixel, and the calibration process is based on the offset vector and brightness correction coefficient of each pixel at the corresponding time recorded in the dynamic compensation vector field. Optionally, the core visual feature bands and background visual feature bands after inverse perturbation calibration are synthesized by an image reprojection algorithm. The reprojection algorithm aligns and fuses the processed feature bands on the image plane based on the internal parameters and external pose parameters of the composite visual sensor head, and finally outputs an enhanced visual feature image for subsequent printed pattern positioning.
[0069] In some embodiments, the specific mapping relationship for generating the dynamic compensation vector field is defined by the following model, whereby the dynamic compensation vector field is used to quantify the combined impact of ambient light and device vibration on image pixels. The dynamic compensation vector field is located at pixel coordinates. and time Vector at It can be represented as:
[0070]
[0071] in: Indicates the pixel in The expected position offset in the direction. Indicates the pixel in The expected position offset in the direction, these two components together constitute the position offset. This indicates the expected pixel brightness fluctuation caused by ambient light fluctuations. and The calculation mainly relies on the frequency analysis results of the equipment vibration disturbance information flow, and maps the identified main vibration frequencies, amplitudes and phases to the geometric deformation of the image plane. The calculation is directly related to the low-frequency trend component extracted after time-series analysis of the ambient light interference field information flow, converting changes in light intensity into linear or nonlinear modulation of image grayscale values. Dynamic compensation vector field As the direct input for subsequent inverse perturbation calibration, each vector component is signed, with positive values indicating the need for compensation in the negative direction to counteract the original perturbation. It can be understood that the model parameters need to be determined during the initial calibration of the production system by fitting image sequences of known standard patterns under controlled perturbation. In specific implementation, inverse perturbation calibration of the core visual feature bands using the dynamic compensation vector field involves reading the coordinates and capture timestamp of each pixel in the core visual feature bands and querying the corresponding... , and The process of setting values. For pixel position correction, in the image spatial domain, the coordinates are... The grayscale values of the pixels are then repositioned to new coordinates using a bilinear interpolation algorithm. For brightness correction, the original pixel's grayscale value is... minus The corrected grayscale value is obtained. The calibration process traverses all pixels in the core visual feature bands, generating new image data with both position and brightness corrected.
[0072] Example 2: See Figure 2The continuous multi-frame images in the original visual field information stream are converted to the frequency domain, and the spectrum of each frame is obtained through Fourier transform. The periodic high-frequency components caused by the grid-like light spot pattern and the characteristic mid-to-low-frequency components caused by the printing pattern and base texture of the printed toilet paper are identified in the spectrum. A set of bandpass digital filters is designed, and the characteristic mid-to-low-frequency components are separated and extracted using these filters. These characteristic mid-to-low-frequency components are then reconstructed into the core visual feature bands after inverse Fourier transform. The image data corresponding to the core visual feature bands is subtracted from the original visual field information stream, and the remaining portion, after removing the noise from the periodic high-frequency components, constitutes the background visual feature bands.
[0073] In specific implementation, the spatial-spectral joint decomposition operation on the original visual field information stream involves sequentially feeding multiple consecutive frames of images from the original visual field information stream into a digital signal processor. Each frame is treated as a two-dimensional discrete function, and a Fast Fourier Transform is performed on this two-dimensional discrete function to transform the image from the spatial domain to the frequency domain, obtaining a complex-form spectrum corresponding to each frame. The spectrum reflects the distribution characteristics of image brightness changes in the frequency domain. Component identification in the spectrum is achieved by analyzing the amplitude spectrum, which exhibits significant energy concentration regions. The periodic high-frequency components caused by the grid-like light spot pattern appear as a series of regularly arranged bright spots, while the characteristic mid-to-low-frequency components caused by the printed pattern and base texture of the printed sanitary napkin appear as continuous bright areas distributed around the center of the spectrum. These two types of components can be distinguished by setting a frequency threshold. In some embodiments, the spatial frequency of the periodic high-frequency components has a direct conversion relationship with the physical period of the grid-like light spot pattern projected by the ring structured light projector. The frequency range of the characteristic mid-to-low-frequency components is predetermined by analyzing the spectral statistical characteristics of known samples; for example, the dominant frequency of the printed pattern texture is usually distributed in a lower frequency band.
[0074] In practical implementation, a set of bandpass digital filters is designed to separate and extract characteristic low-to-mid-frequency components. This set of bandpass digital filters directly acts on the image's spectrum in the frequency domain. The frequency response function of the bandpass digital filter determines which frequency components can pass through. For a Gaussian bandpass digital filter used to extract characteristic low-to-mid-frequency components, its frequency response function... Response at the location Defined by the following formula:
[0075]
[0076] in: This indicates that the bandpass digital filter operates at the normalized frequency. Complex gain at that point, This represents the center frequency of the bandpass digital filter, corresponding to the dominant frequency of the characteristic low-to-mid-frequency components. The bandwidth control parameters of a bandpass digital filter determine the width of the passband. This can be understood from the formula... It is the radial frequency in the two-dimensional frequency domain, and needs to be converted into discrete frequency coordinates corresponding to the image size during calculation. By performing pointwise complex multiplication with the original spectrogram, periodic high-frequency components are suppressed while characteristic mid-to-low-frequency components are preserved. The filtered spectrogram is then transformed back into a spatial domain image through an inverse Fourier transform, which represents the reconstructed core visual feature bands. In some embodiments, a set of bandpass digital filters may contain multiple bandpass filters with different... and A filter is used to cover all subbands of the characteristic low-to-mid frequency components, and the outputs of multiple filters are synthesized through an image fusion algorithm. Optionally, after the core visual feature bands are reconstructed, their image data is subtracted pixel by pixel from the image data corresponding to the original visual field information stream. The resulting differential image data mainly contains grid-like light spot pattern information and noise.
[0077] In practice, the image data corresponding to the core visual feature bands is subtracted from the original visual field information stream. The remaining portion contains periodic high-frequency noise components and some residual substrate information. Removing the periodic high-frequency noise components is achieved by applying a notch filter in the frequency domain. The notch filter sets a stopband for the regularly bright spot locations identified in the spectrogram, setting the amplitude of frequency components near these locations to zero. Then, an inverse Fourier transform is performed on the processed spectrogram to obtain the noise-removed image data, which constitutes the background visual feature bands. It can be understood that the background visual feature bands mainly reflect the uniform texture and possible wrinkles and shadows of the toilet paper substrate, while the printed pattern information has been largely separated into the core visual feature bands. Optionally, the final output of the spatial-spectral joint decomposition operation is two registered images: the core visual feature band image and the background visual feature band image. These two images will be used for subsequent synthesis to generate an enhanced visual feature image.
[0078] Example 3: See Figure 3A time-series analysis is performed on the ambient light interference field information stream to extract its low-frequency trend component that fluctuates over time. Frequency analysis is performed on the equipment vibration disturbance information stream to identify its main vibration frequency components and corresponding amplitudes. A two-dimensional mapping model is established with time and image pixel position as independent variables. The low-frequency trend component, the main vibration frequency components, and amplitudes are used as input parameters to calculate the expected position offset and brightness fluctuation of each pixel at each time point due to ambient light and vibration. The expected position offset and brightness fluctuation together constitute the dynamic compensation vector field, where each vector contains direction and intensity information. The reverse perturbation calibration of the core visual feature band using the dynamic compensation vector field includes the following steps: Reading the time and position information corresponding to each pixel of the core visual feature band in the dynamic compensation vector field. Based on the direction and intensity information provided by the dynamic compensation vector field, performing a reverse translation interpolation operation on the pixel position of the core visual feature band in the spatial domain to offset the expected position offset. Simultaneously, based on the brightness fluctuation provided by the dynamic compensation vector field, the pixel grayscale values of the core visual feature band are adjusted compensatorily. After processing all pixels, the core visual feature band, with both position and brightness corrected, is output.
[0079] In practical implementation, generating a dynamic compensation vector field involves preprocessing and modeling the ambient light interference field information stream and the equipment vibration disturbance information stream. Time series analysis of the ambient light interference field information stream uses a moving average algorithm to extract its low-frequency trend components that fluctuate over time; these low-frequency trend components reflect the slow change in ambient light intensity. Frequency analysis of the equipment vibration disturbance information stream applies a fast Fourier transform to the acceleration data collected by a triaxial micro-vibration sensor, identifying the main vibration frequency components and corresponding amplitudes corresponding to peak values in the acceleration power spectrum that exceed a threshold. A two-dimensional mapping model is established with time and image pixel position as independent variables. The low-frequency trend components obtained from time series analysis and the main vibration frequency components and amplitudes obtained from frequency analysis are used as input parameters. The mapping model calculates the expected positional offset and brightness fluctuation of each pixel at each time point due to ambient light and vibration using geometric optical relationships and kinematic equations. In some embodiments, the calculation of brightness fluctuation is directly related to the low-frequency trend component of the ambient light interference field information flow. Assuming that the image sensor response is linear, the brightness fluctuation is proportional to the change in light intensity. The calculation of the expected position offset depends on the device vibration disturbance information flow. The vibration acceleration is converted into displacement through a quadratic integral and then mapped to the offset of the image pixel coordinates through the imaging model of the composite vision sensor head.
[0080] In practical implementation, the mapping relationship between the expected positional offset and the brightness fluctuation can be expressed by a parameterized function. For a pixel on the image plane, its dynamic compensation vector field... The mathematical expression is as follows:
[0081]
[0082] in: Represents the pixel coordinates in the image and timestamp The dynamically compensated vector generated at that location, Indicates in Time Pixel In the image Expected position offset in the axial direction Indicates in Time Pixel In the image Expected position offset in the axial direction Indicates in Time Pixel The expected brightness fluctuation. This represents a mapping function from input parameters to an output vector. Indicates time The low-frequency trend component obtained from the ambient light interference field information stream This indicates the first identified from the equipment vibration disturbance information stream. angular frequencies of the main vibrational frequency components Phase and amplitude The set. It can be understood that the mapping function... The specific form is determined during the system calibration phase. During calibration, known ambient light changes and mechanical vibrations are actively applied, and image displacement and brightness variation data are collected to fit the function parameters. The expected positional offset and brightness fluctuation together constitute a dynamic compensation vector field. Each vector in the dynamic compensation vector field contains direction and intensity information, with the direction determined by… and The sign determines the strength. , and The absolute value reflects the magnitude.
[0083] In practice, using a dynamic compensation vector field to perform inverse perturbation calibration on the core visual feature band involves reading the time and position information corresponding to each pixel in the core visual feature band from the dynamic compensation vector field. Each pixel in the core visual feature band image is accompanied by its capture timestamp and original coordinates. Based on the direction and intensity information provided by the dynamic compensation vector field, a reverse translation interpolation operation is performed on the pixel positions of the core visual feature band in the spatial domain. This reverse translation interpolation operation calculates the theoretical original position of the pixel. Then, it is obtained from the original core visual feature band image through bilinear interpolation sampling. The grayscale value at the specified location is used to assign that value to the position in the output image. Meanwhile, based on the brightness fluctuation provided by the dynamic compensation vector field... The grayscale values of pixels in the core visual feature bands are adjusted compensatorily. This compensatory adjustment involves adjusting the grayscale values of pixels after position correction. Subtract the corresponding brightness fluctuation amount, which is the output grayscale value. .
[0084] Example 4: On the enhanced visual feature image, a local gray-level variance maximum point detection method is used to automatically select at least one point as an initial seed point. Centered on the initial seed point, the gray-level gradient value and texture direction consistency of pixels within its eight neighborhoods are calculated. A dynamic growth threshold is set, determined by the average gradient of the edge pixels of the current growth region. Pixels with gray-level gradient values less than the dynamic growth threshold and texture direction consistency higher than a preset consistency threshold in their eight neighborhoods are included in the current growth region, and the edges of the growth region are updated. The neighborhood exploration, threshold comparison, and region update process is iterated until no new pixels are included in several consecutive iterations, or the edge of the growth region touches the physical boundary of the image, at which point the growth termination boundary is determined to have been reached. The entire pixel region covered by the final stop of growth is marked as the candidate localization region. The extraction of the geometric shape descriptor subset and spectral reflectance descriptor subset of the candidate localization region includes the following operations: Contour tracking is performed on the candidate localization region to obtain its closed boundary curve; the moment invariant, Fourier descriptor, and curvature statistical features of the boundary curve are calculated and combined to form the geometric shape descriptor subset. On the enhanced visual feature image, for the pixels in the candidate localization region, the gray-level distribution of the pixels in multiple narrow bands is analyzed, and the mean, variance and higher-order moments of the spectral reflectance are calculated to form the spectral reflectance descriptor set.
[0085] In specific implementation, the seed point diffusion growth of the printed pattern region on the enhanced visual feature image includes running a local gray-level variance maxima detection algorithm on the enhanced visual feature image. Within a sliding window of a preset size, the gray-level variance of each pixel's neighborhood is calculated. Pixels with variance values exceeding a preset threshold and being maxima in their local neighborhood are marked as candidate points. These candidate points are then sorted according to their variance, and the top N points are selected as initial seed points for region growth. Centered on each initial seed point, the gray-level gradient values and texture direction consistency of pixels within the eight-neighborhood of the initial seed point are calculated. The gray-level gradient values are calculated using the Sobel operator, and the texture direction consistency is determined by calculating the angle histogram of the principal gradient direction within the local image patch. Figure 1 Consistency is achieved. A dynamic growth threshold is set, which is determined by the average gradient of the edge pixels of the current growth region. The specific calculation formula is as follows:
[0086]
[0087] in: Indicates the dynamic growth threshold. This represents a scaling factor between 0 and 1. This represents the average grayscale gradient magnitude of all edge pixels in the current growth region. In some embodiments, the scaling factor... The typical value range is 0.4 to 0.6, the number of initial seed points N can be set to 3 to 5, and the preset consistency threshold for texture direction consistency is usually set to 0.7 (normalized value).
[0088] In practice, the iterative process of seed point diffusion growth involves treating each initial seed point as a growth region and checking in each iteration whether the eight-neighbor pixels of all edge pixels in the growth regions meet the inclusion criteria. The inclusion criteria require that the grayscale gradient value of the eight-neighbor pixels is less than the dynamic growth threshold of the current region. Furthermore, the consistency of its texture direction with the main texture direction of the growth region exceeds a preset consistency threshold. When a pixel meets the inclusion criteria, it is marked as a new member of the same growth region, and the boundary contour and edge pixel set of the growth region are updated. It can be understood that each time a new pixel is included, the edge pixel set and its average gradient of the growth region need to be recalculated. Thus, dynamically updated The value of this mechanism allows the growth process to expand rapidly in low-gradient regions (within the printed pattern) and slow down in high-gradient regions (edges of the pattern). Iterative processes of neighborhood exploration, threshold comparison, and region update continue until no new pixels are included in several consecutive iterations, or the edge of the growth region touches the physical boundary of the image. At this point, the growth termination boundary is determined. The entire pixel region covered by the final stop of growth is marked as a candidate localization region. If multiple initial seed point growth regions merge, the merged connected region is considered a single candidate localization region.
[0089] In practical implementation, extracting the geometric shape descriptor subset and spectral reflectance descriptor subset of the candidate localization region includes contour tracking of the candidate localization region to obtain its closed boundary curve. Contour tracking employs boundary tracking algorithms, such as the Moore neighborhood tracking algorithm. Calculating the moment invariants of the boundary curve involves calculating the Hu moments of the region; Hu moments are a set of moment features that remain invariant under translation, scaling, and rotation. Calculating the Fourier descriptor involves converting the coordinate sequence of the boundary curve into a complex sequence, then performing a discrete Fourier transform, and taking the amplitude spectrum of the transformed coefficients as the shape descriptor. Calculating curvature statistics involves performing discrete differentiation on the boundary curve, calculating the curvature of each contour point, and statistically analyzing the mean, variance, skewness, and kurtosis of the curvature of all contour points. The moment invariants, Fourier descriptors, and curvature statistics together form the geometric shape descriptor subset. On the enhanced visual feature image, for pixels within the candidate localization region, the grayscale distribution under multiple narrow bands is analyzed. Narrow bands are achieved by using multispectral light sources during image acquisition or by separating different color channels during image processing. The mean, variance, and higher-order moments of its spectral reflectance are calculated to form a set of spectral reflectance descriptors. The calculation process involves statistically analyzing the grayscale values of all pixels within the region in each band. Refer to Table 1, which shows the structure of a simplified set of geometric shape descriptors and spectral reflectance descriptors.
[0090] Table 1: Candidate Location Region Description Subset Table
[0091]
[0092] In some embodiments, higher-order moments include third-order and fourth-order central moments. After extracting the descriptor sets, the geometric topography descriptor set and the spectral reflectance descriptor set are encapsulated into feature vectors for subsequent matching. Optionally, before calculating the geometric topography descriptor, the contour curve can be resampled to ensure consistent point counts, facilitating comparison with the template. It is understood that the calculation of the spectral reflectance descriptor depends on the color or multispectral information contained in the enhanced visual feature image.
[0093] See Figure 4The study presents the distribution characteristics of spectral reflectance descriptors (mean, variance, third central moment, and fourth central moment) of candidate positioning regions in multiple spectral channels, including red, green, blue, near-infrared, and ultraviolet: the mean spectral reflectance (red bars) is significantly higher in the red channel than in other channels, reflecting the region's advantage in reflectance intensity in the red band; the variance of spectral reflectance (blue bars) shows a fluctuating decreasing trend from red to ultraviolet, reflecting the difference in the dispersion of reflectance intensity in different bands; the third central moment (green curve, representing skewness) is negatively skewed in the green and blue channels, turning positively skewed in the near-infrared and ultraviolet channels, indicating differences in the asymmetric characteristics of reflectance distribution in each band; the fourth central moment (purple curve, representing kurtosis) is steeply distributed in the red channel (kurtosis is significantly positive), turning flatly distributed in the ultraviolet channel (kurtosis is negative), reflecting the change in the steepness of reflectance distribution in different bands. The multi-channel distribution characteristics of the descriptor are the core basis for spectral matching between candidate localization regions and standard templates, and can support the subsequent solution of spatial transformation parameters and localization calibration.
[0094] Example 5: The geometric shape descriptor subset and the spectral reflectance descriptor subset are encapsulated into a feature vector. The Mahalanobis distance between the feature vector and the standard feature vector corresponding to each standard nursing pad print shape spectral file in the reference template library is calculated. The top few standard nursing pad print shape spectral files with the smallest Mahalanobis distance are selected as matching candidates. Using the iterative nearest point algorithm, the correspondence between local feature points between the candidate positioning region and each matching candidate is further calculated, and the optimal matching standard nursing pad print shape spectral file is selected.
[0095] The process of calculating the spatial transformation parameters between the candidate positioning region and at least one standard nursing pad print morphology spectral file based on the matching results includes the following steps. The spatial transformation parameters include rotation, translation, and scaling factor. Based on the correspondence between local feature points between the candidate positioning region and the optimally matched standard nursing pad print morphology spectral file, the optimal rigid body transformation matrix is solved using a least-squares estimation method. The rotation about an axis perpendicular to the image plane, the translation within the image plane, and the overall scaling factor are decomposed from the rigid body transformation matrix.
[0096] In practice, the process involves matching the extracted geometric morphology and spectral reflectance descriptors with the standard nursing pad print morphology spectral profiles pre-stored in the reference template library. This includes concatenating the extracted geometric morphology descriptors and spectral reflectance descriptors in a predetermined order and encapsulating them into a multidimensional feature vector. Each standard nursing pad print morphology spectral profile in the reference template library stores a corresponding standard feature vector. Calculate the eigenvectors With each standard feature vector Mahalanobis distance between Mahalanobis distance The calculation formula is:
[0097]
[0098] in: Indicates the query feature vector and the first The Mahalanobis distance scalar value between the template feature vectors This represents the query feature vector extracted and encapsulated from the candidate location regions. This indicates the first one stored in the reference template library. The standard feature vector corresponding to the spectral profile of the printed morphology of a standard nursing pad. The superscript represents the inverse of the covariance matrix of the set of all standard eigenvectors. This represents the transpose operation of a vector. It can be understood that Mahalanobis distance considers the correlation and variance between different dimensions of features, making it a more effective similarity measure than Euclidean distance. The first vector with the smallest Mahalanobis distance is selected. A standard nursing pad print morphology spectral profile was used as a matching candidate. It is a preset positive integer.
[0099] In specific implementation, the Iterative Closest Point Algorithm (IBPA) is used to further calculate the correspondence between local feature points of the candidate positioning region and each matching candidate, and to select the optimal matching standard nursing pad print morphology spectral archive. The IBPA is input to the boundary point set of the candidate positioning region and the standard boundary point set stored in the matching candidate archive. The algorithm iteratively finds the closest point pair between the two point sets and calculates a rigid body transformation that minimizes the sum of squared distances between corresponding point pairs. The point correspondence and transformation are updated through multiple iterations until convergence or the maximum number of iterations is reached. In some embodiments, after calculation by the IBPA, the average distance residual of all corresponding point pairs is used as a measure of matching accuracy. The standard nursing pad print morphology spectral profile with the smallest average distance residual is selected as the optimal match from the candidate matching pairs. Optionally, before executing the iterative nearest point algorithm, the boundary point set of the candidate positioning region and the standard boundary point set can be resampled to make the number of points in both sets the same and normalized to the same scale, thereby improving the convergence speed and stability of the algorithm.
[0100] In practice, spatial transformation parameters between the candidate positioning region and at least one standard nursing pad print morphology spectral profile are calculated based on the matching results. These spatial transformation parameters include rotation. Translation amount and scale factor Based on the correspondence between local feature points of candidate positioning regions and the optimal matching standard nursing pad print morphology spectral archive, provided by the iterative nearest point algorithm or other feature point matching algorithms, a series of coordinate pairs of source points (from candidate positioning regions) and target points (from template archives) are obtained. The optimal rigid body transformation matrix is solved using the least squares estimation method. The objective is to minimize the sum of squared position errors of all corresponding point pairs after transformation. ,in: These are the coordinates of the source point. These are the coordinates of the corresponding target point. It is a rotation matrix. It is a translation vector. It is the scaling factor.
[0101] See Figure 5 In the candidate template matching stage for positioning printed sanitary paper pads, the quantitative analysis of spatial transformation parameters relies on rigid body transformation matrix decomposition technology. Specifically, the spatial transformation parameters (rotation, translation, and scaling factor) of each candidate template are represented as a multidimensional feature space composed of discrete parameter values. Each candidate template (candidate templates 1 to 5) corresponds to a parameter set consisting of rotation (°), translation X (pixels), translation Y (pixels), and scaling factor. The quantification of parameter differences among different candidate templates is achieved through multi-dimensional numerical comparison using a bar chart: for each candidate template, the angle value of its rotation, the pixel offset of translation X / Y, and the magnitude of its scaling factor are recorded. The degree of difference in parameter distribution is reflected by the absolute value and fluctuation range of each dimension's values. This parameter set serves as the core basis for evaluating matching accuracy and is used to screen the optimal candidate template. During parameter presentation, the numerical range of each dimension's parameters covers [-8, 8] (pixels / °), with the translation Y exhibiting the largest fluctuation (e.g., the translation Y of candidate template 5 reaches -8 pixels), while the scaling factor's value is generally at a low level.
[0102] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0103] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A machine vision-based method for positioning printed sanitary napkins and toilet paper, characterized in that, The method includes: Through a multi-source sensing environmental perception process, the original visual field information stream, ambient light interference field information stream, and equipment vibration disturbance information stream containing the target nursing pad printed toilet paper are captured simultaneously to generate an enhanced visual feature image. Seed point diffusion growth is performed on the printed pattern area on the enhanced visual feature image. The seed point diffusion growth expands the region according to the gray-scale gradient and texture continuity constraints of the printed pattern until it reaches the preset growth termination boundary, thereby outlining the candidate positioning area. Extract the geometric shape description subset and spectral reflectance description subset of the candidate positioning region, and match them item by item with the standard nursing pad print shape spectral archive pre-stored in the reference template library; Based on the matching results, calculate the spatial transformation parameters between the candidate positioning region and at least one standard nursing pad print morphology spectral file; The coordinate frame of the candidate positioning area is rigidly adjusted by applying the spatial transformation parameters so that the candidate positioning area is spatially aligned with the selected standard nursing pad print morphology spectral file. Within the aligned coordinate frame, the centroid coordinates of the candidate positioning region are calculated, and the centroid coordinates are mapped back to the original captured visual field space to generate the final positioning decision for the printed toilet paper pad.
2. The machine vision-based method for positioning printed sanitary napkins and toilet paper as described in claim 1, characterized in that, The process of simultaneously capturing the original visual field information stream, ambient light interference field information stream, and equipment vibration disturbance information stream, including the target nursing pad printed toilet paper, through multi-source sensing environmental perception includes: A spatial-spectral joint decomposition operation is performed on the original visual field information stream to extract the core visual feature bands related to the printed pattern and the background visual feature bands related to the toilet paper substrate; The ambient light interference field information stream and the device vibration disturbance information stream are fused to generate a dynamic compensation vector field, and the dynamic compensation vector field is used to perform reverse disturbance calibration on the core visual feature band. The core visual feature bands, which have undergone inverse perturbation calibration, are reprojected and synthesized with the background visual feature bands to generate an enhanced visual feature image. The synchronous capture includes: Deploy a composite vision sensor head that integrates a wide-angle microscope lens and a ring structured light projector, the composite vision sensor head scanning the printed sanitary napkins on the production line at a fixed frequency; The ring structured light projector is activated, projecting a grid-shaped light spot pattern with specific encoded information onto the surface of the printed toilet paper pad during scanning. The wide-angle microscope lens synchronously receives light spot pattern images reflecting back from the surface of the printed toilet paper of the nursing pad, which carry surface deformation information, and the light spot pattern images constitute the original visual field information stream; An ambient light meter and a triaxial micro-vibration sensor are installed on the side of the composite vision sensor head to continuously record the light intensity change data and the equipment vibration acceleration data during the scanning process, respectively, forming the ambient light interference field information stream and the equipment vibration disturbance information stream.
3. The machine vision-based method for positioning printed sanitary napkins and toilet paper as described in claim 2, characterized in that, The spatial-spectral joint decomposition operation on the original visual field information stream includes: The continuous multiple frames of images in the original visual field information stream are converted to the frequency domain, and the spectrum of each frame of the image is obtained by Fourier transform. The periodic high-frequency components caused by the grid-like light spot pattern and the characteristic mid-to-low-frequency components caused by the printing pattern and base texture of the printed toilet paper pad were identified in the spectrum. Design a set of bandpass digital filters, use the bandpass digital filters to separate and extract the characteristic low-frequency components, and reconstruct the core visual feature bands after inverse Fourier transform of the characteristic low-frequency components. The image data corresponding to the core visual feature band is subtracted from the original visual field information stream, and the remaining part is processed by removing the periodic high-frequency component noise to form the background visual feature band.
4. The machine vision-based method for positioning printed sanitary napkins and toilet paper as described in claim 2, characterized in that, The generation of a dynamic compensation vector field includes: Time series analysis was performed on the ambient light interference field information stream to extract its low-frequency trend component that fluctuates over time; Frequency analysis is performed on the vibration disturbance information stream of the equipment to identify its main vibration frequency components and corresponding amplitudes; A two-dimensional mapping model is established with time and image pixel position as independent variables. The low-frequency trend component, the main vibration frequency component and amplitude are used as input parameters to calculate the expected position offset and brightness fluctuation of each pixel at each time point due to ambient light and vibration. The expected position offset and the brightness fluctuation together constitute the dynamic compensation vector field, and each vector in the dynamic compensation vector field contains direction information and intensity information.
5. The machine vision-based method for positioning printed toilet paper in nursing pads as described in claim 4, characterized in that, The reverse perturbation calibration of the core visual feature band using the dynamic compensation vector field includes: Read the time and position information corresponding to each pixel of the core visual feature band in the dynamic compensation vector field; Based on the direction and intensity information provided by the dynamic compensation vector field, a reverse translation interpolation operation is performed on the pixel position of the core visual feature band in the spatial domain to offset the expected position offset. Simultaneously, based on the brightness fluctuation provided by the dynamic compensation vector field, the pixel grayscale values of the core visual feature bands are adjusted compensatorily. After processing all pixels, the core visual feature bands with corrected position and brightness are output.
6. The machine vision-based method for positioning printed sanitary napkins and toilet paper as described in claim 1, characterized in that, The seed point diffusion growth of the printed pattern region on the enhanced visual feature image includes: On the enhanced visual feature image, a local gray-level variance maximum point detection method is used to automatically select at least one point as an initial seed point; Centered on the initial seed point, calculate the grayscale gradient values of the pixels in its eight neighboring areas and ensure consistency with the texture direction; A dynamic growth threshold is set, which is determined by the average gradient of the edge pixels of the current growth region; Pixels whose gray-level gradient values in their eight neighborhoods are less than the dynamic growth threshold and whose texture direction consistency is higher than the preset consistency threshold are included in the current growth region, and the edges of the growth region are updated. The process of iteratively exploring the neighborhood, comparing thresholds, and updating the region continues until no new pixels are included in several consecutive iterations, or the edge of the growing region touches the physical boundary of the image. At this point, it is determined that the growth termination boundary has been reached. The entire pixel region covered by the point where growth eventually stops is marked as the candidate localization region.
7. The machine vision-based method for positioning printed sanitary napkins and toilet paper as described in claim 6, characterized in that, The extraction of the geometric shape descriptor subset and spectral reflectance descriptor subset of the candidate localization region includes: Contour tracking is performed on the candidate positioning region to obtain its closed boundary curve. The moment invariant, Fourier descriptor and curvature statistical features of the boundary curve are calculated and combined into the geometric shape descriptor subset. On the enhanced visual feature image, for the pixels in the candidate localization region, the gray-level distribution of the pixels in multiple narrow bands is analyzed, and the mean, variance and higher-order moments of the spectral reflectance are calculated to form the spectral reflectance descriptor set.
8. The machine vision-based method for positioning printed toilet paper in nursing pads as described in claim 1, characterized in that, The step of matching it item by item with the standard nursing pad print morphology spectral files pre-stored in the reference template library includes: The geometric shape descriptor subset and the spectral reflectance descriptor subset are encapsulated into a feature vector; Calculate the Mahalanobis distance between the feature vector and the standard feature vector corresponding to each standard nursing pad print morphology spectral file in the reference template library; The top few standard nursing pad print morphology spectral files with the smallest Mahalanobis distance were selected as matching candidates; Using the iterative nearest point algorithm, the correspondence between local feature points between the candidate positioning region and each matching candidate is further calculated, and the optimal matching standard nursing pad print morphology spectral profile is selected.
9. The machine vision-based method for positioning printed sanitary napkins and toilet paper as described in claim 8, characterized in that, The calculation of spatial transformation parameters between the candidate positioning region and at least one standard nursing pad print morphology spectral file based on the matching results includes: The spatial transformation parameters include rotation, translation, and scale factor; Based on the correspondence between local feature points between the candidate positioning region and the optimal matching standard nursing pad print morphology spectral file, the least squares estimation method is used to solve for the optimal rigid body transformation matrix. The rotation about an axis perpendicular to the image plane, the translation in the image plane, and the overall scaling factor are decomposed from the rigid body transformation matrix.
10. A machine vision-based positioning system for printed sanitary pads and toilet paper, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the machine vision-based method for positioning printed toilet paper for nursing pads as described in any one of claims 1 to 9.