Method for determining boundaries of object based on digital correlation method of image processing

By applying controlled oscillations and digital image correlation with noise compensation, the method addresses the limitations of traditional boundary detection, achieving subpixel accuracy for objects with complex shapes and variable lighting, ensuring high precision and consistency with reference measurements.

RU2865664C1Active Publication Date: 2026-07-07FEDERALNOE GOSUDARSTVENNOE BYUDZHETNOE OBRAZOVATELNOE UCHREZHDENIE VYSSHEGO OBRAZOVANIYA SIBIRSKIJ GOSUDARSTVENNYJ UNIV PUTEJ SOOBSHCHENIYA SGUPS
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
RU · RU
Patent Type
Patents
Current Assignee / Owner
FEDERALNOE GOSUDARSTVENNOE BYUDZHETNOE OBRAZOVATELNOE UCHREZHDENIE VYSSHEGO OBRAZOVANIYA SIBIRSKIJ GOSUDARSTVENNYJ UNIV PUTEJ SOOBSHCHENIYA SGUPS
Filing Date
2025-12-11
Publication Date
2026-07-07

AI Technical Summary

Technical Problem

Traditional methods of boundary detection in static images are limited by pixel discreteness, leading to systematic errors and reduced accuracy, which is insufficient for modern industrial and scientific applications requiring subpixel precision, especially for objects with complex geometric shapes and variable lighting conditions.

Method used

Implementing controlled oscillations of objects with different frequencies and directions, using digital image correlation and subpixel interpolation, combined with noise compensation and morphological filtering to enhance boundary detection accuracy.

Benefits of technology

Achieves subpixel accuracy in boundary determination, reducing errors to less than 0.1 pixel, suitable for complex geometric shapes and variable lighting conditions, with high precision and consistency with reference measurements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000005
    Figure 00000005
  • Figure 00000019
    Figure 00000019
Patent Text Reader

Abstract

FIELD: digital image processing.SUBSTANCE: method consists of placing the object on a contrasting background, installing a video camera with a frame rate of at least 120 FPS, initiating controlled oscillations of the object in various directions with an amplitude of 0.3-2.0 mm and various frequencies in the range of 2-15 Hz, performing continuous video recording of the oscillating object, dividing each frame into overlapping subregions of 32×32 pixels with 75% overlap, calculating normalized cross-correlation between adjacent frames for each subregion, calculating subpixel interpolation using parabolic approximation, form a two-dimensional field of displacement vectors, calculating the gradient of the displacement field to identify areas with sharp changes corresponding to the object boundaries, performing adaptive thresholding using the Otsu method and morphological filtering to highlight contours, achieving boundary detection accuracy of less than 0.1 pixel due to temporal averaging of the results over multiple oscillation cycles.EFFECT: increase in the accuracy of determining the boundaries of objects with sub-pixel accuracy.2 cl, 1 tbl
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to the field of digital image processing and optical measurements and can be used to improve the accuracy of determining the boundaries of objects in computer vision systems, automated quality control, medical diagnostics and scientific research of dynamic processes, as well as to reduce measurement errors that arise during static image analysis.

[0002] Modern machine vision and quality control systems require high-precision object boundary detection with subpixel accuracy to ensure reliable inspection of product geometric parameters, material deformation analysis, and dynamic process monitoring. However, traditional methods of boundary detection in static images are limited by the discreteness of the pixel matrix and do not provide the required measurement accuracy.

[0003] Due to the lack of effective subpixel interpolation methods when analyzing stationary objects, it is possible to determine the position of boundaries only with an accuracy of up to a whole pixel, which leads to the accumulation of systematic errors, the influence of random matrix noise and lighting artifacts, reducing the overall measurement accuracy to 1-2 pixels, while modern industrial control and scientific research tasks require an accuracy of less than 0.1 pixel to ensure high-quality analysis of objects and compliance with international accuracy standards.

[0004] A known method for noise-immune gradient extraction of object contours in digital images (see patent RU No. 2589301, IPC G06K 9 / 36; G06K 9 / 40; G06K 9 / 56, published 10.07.2016, Bulletin No. 19), consists in the fact that a binary matrix of noise position estimates of the size of an image is first obtained, on the basis of which the values ​​of the weighting coefficients of differently oriented Prewitt gradient masks are changed, with the help of which the approximate value of the gradient components at each point of the image is calculated, after which the gradient modulus is determined, and by its threshold transformation, elements are highlighted in black on a new white matrix, the gradient modulus for which in the corresponding image coordinates exceeds the transformation threshold, characterized in that, in order to increase sensitivity to useful brightness differences in the image under conditions of pulsed noise, the values ​​of the weighting coefficients,and the form of differently oriented Prewitt gradient masks is changed as follows: if, during mask processing, all three non-zero coefficients of the same sign that form a negative or positive half mask are matched for multiplication by pixels of the image affected by impulse noise, as indicated by the non-zero elements of the noise position estimate matrix, then these three coefficients of the half mask are reset to zero, and the half mask itself is increased by the five nearest outer elements bordering them, while the weighting coefficients of the new five elements are assigned the value of 0.6 of the same sign as the zeroed coefficients, in the absence of noise in these five elements; or increased by four, or three, or two, or by one nearest element, the weighting coefficients of which are assigned the values ​​0.75, or 1, or 1.5, or 3, respectively, of the same sign in the presence of one, or two, or three,or four affected elements, respectively; or all six non-zero coefficients of the Prewitt mask are assigned zero values ​​if all five elements bordering the half-mask are affected by interference.

[0005] The disadvantage of this method is its high dependence on the quality of the preliminary determination of the noise position estimation matrix, since inaccurate noise identification can lead to distortion of gradient masks and deterioration in the quality of contour detection, and the complexity of the algorithm for adaptively changing the weighting coefficients and the shape of the Prewitt masks requires significant computing resources and can slow down the image processing process, in addition, the method is focused on working with impulse noise and may be ineffective with other types of noise, such as Gaussian noise or systematic distortions, and fixed values ​​​​of the weighting coefficients (0.6; 0.75; 1; 1.5; 3) do not provide optimal adaptation to various image characteristics and types of objects, which can lead to the loss of weak contours or the appearance of false boundaries in areas with variable contrast.

[0006] The closest to the proposed solution is a method for processing signals for detecting rectilinear boundaries of objects observed in an image (see patent RU No. 2522924, IPC G06T 7 / 00, published on 20.07.2014, Bulletin No. 20), which includes estimating the gradient field of the original image, processing it using the Radon transform and searching for local maxima in the resulting parametric space, characterized in that, based on the gradient field, three images are formed, which are then subjected to the Radon transform and combined into one image by means of pointwise weighted summation of three Radon transforms of the images.

[0007] The disadvantage of this method, adopted as a prototype, is that it is only applicable to detecting rectilinear boundaries of objects, which significantly limits the scope of its use in analyzing objects of complex geometric shapes with curved contours, as well as the high computational complexity of the Radon transform and the need for preliminary parameter adjustment for each type of analyzed objects. In addition, the method does not provide adaptability to changing lighting conditions and image contrast, which can lead to missing weakly defined boundaries or false detection of boundaries in areas with a high noise level, and the lack of mechanisms for verifying the detected boundaries through time analysis does not allow distinguishing the true boundaries of an object from image artifacts, which leads to the accumulation of measurement errors and a decrease in the overall accuracy of boundary determination.

[0008] The technical challenge is to improve the accuracy of determining the boundaries of objects by introducing controlled oscillations with adaptive selection of the optimal direction of movement and using digital correlation of images in a time sequence of frames, which ensures subpixel interpolation of the position of the boundaries through the analysis of multiple positions of the object relative to the pixel matrix and noise compensation, making it possible to achieve an accuracy of determining the boundaries of less than 1 pixel instead of the limitations of static gradient analysis on objects of complex geometric shapes.

[0009] The technical result is achieved by placing an object on a contrasting background, initiating controlled oscillations of the object in various directions with different frequencies, using a video camera with a frame rate of at least 120 frames per second (FPS) to continuously film the oscillating object, transmitting these images to a computing device, using special software installed on the hard drive of the computing device, processing the images: obtaining a set of images at various positions of the object, dividing each captured frame into a grid of overlapping square subregions, for each subregion calculating the normalized cross-correlation with the corresponding subregion in the next frame, comparing the pixel intensity distributions, determining the coordinates of the maximum of the correlation function for each subregion,subpixel interpolation is calculated using the parabolic approximation method in the vicinity of the found maximum, a two-dimensional field of displacement vectors is formed for the entire image, the gradient of the obtained displacement field is calculated to identify areas with sharp changes in displacements, threshold processing is performed on the gradient value to highlight the potential boundaries of the object, morphological filtering of the binary mask is performed to eliminate noise and fill gaps, the result is temporarily averaged over multiple cycles of oscillations in different directions, from which the final contours of the object boundaries are obtained, wherein controlled oscillations of the object in different directions are carried out with an amplitude of 0.3-2.0 mm and frequencies in the range of 2-15 Hz, continuous video recording of the oscillating object is performed by dividing each frame into overlapping subregions of 32x32 pixels in size with an overlap of 75%.

[0010] The proposed method for processing signals to detect object boundaries can be implemented using a general-purpose personal computer (PC) with specialized software. A video camera with a frame rate of at least 120 frames per second (FPS) and a resolution sufficient for detailed analysis of the object's boundaries is installed. The lighting system is adjusted to ensure uniform, flicker-free illumination. Using a vibration platform, controlled vibrations of the object are initiated in various directions with an amplitude of 0.3-2.0 mm and various frequencies in the range of 2-15 Hz.This amplitude range provides a displacement exceeding the recording system's noise level and enables subpixel displacement calculations using digital image correlation. The frequency range ensures a sufficient number of frames per oscillation period at a video recording rate of 120 FPS: with this setting, one oscillation cycle produces between 60 and 8 frames, respectively. This ensures the correct operation of image correlation algorithms while avoiding the excitation of high-frequency oscillation modes limited by the mechanical characteristics of the vibration platform and the object. Continuous video recording of the oscillating object is performed to obtain a set of images at various object positions for at least 10 complete oscillation cycles for each direction and frequency.

[0011] Divide each captured frame into a grid of overlapping square subregions measuring 32x32 pixels with 75% overlap, resulting in a regular rectangular grid of pixel coordinates with a horizontal and vertical pitch of 8 pixels, generated by the image processing module software. A normalized cross-correlation is calculated for each subregion of the resulting grid with the corresponding subregion in the next frame by comparing the pixel intensity distributions within a ±10 pixel search window.

[0012] Calculate the normalized cross-correlation using the formula:

[0013] , Where:

[0014] - F(x i ,y j ) - intensity of the reference image pixel;

[0015] - G(x i +u,y j ;+v) - pixel intensity in the deformed image at position (x i +u,y j +v);

[0016] - - average intensities in subregions of 32×32 pixels;

[0017] - u, v - displacement components in pixels;

[0018] - x i +u,y j +v - coordinates in the deformed image after applying the offset.

[0019] Determine the coordinates of the correlation function maximum for each subregion with integer precision. Calculate subpixel interpolation using parabolic approximation in the vicinity of the determined maximum to improve the accuracy of displacement determination to fractions of a pixel:

[0020] ;

[0021] , Where:

[0022] - c i,j - values ​​of the correlation function in the vicinity of the maximum in positions (i, j);

[0023] - dx, dy - subpixel corrections to the maximum position.

[0024] A two-dimensional field of displacement vectors is generated for the entire image. For each subregion, a displacement vector is programmatically determined based on the position of the maximum normalized cross-correlation between the current and subsequent frames. The resulting values ​​are entered into a two-dimensional array (field) generated by the image processing module of the computing device, where each vector corresponds to the displacement of the subregion center between adjacent frames. When using the digital image correlation method, work is performed with small subregions, where x, y is the center of the region, and Δx, Δy are the displacements from the center within the subregion. The relationship between the coordinates of the reference and deformed images is described by two-dimensional affine transformations:

[0025]

[0026] By multiplying this matrix by the vector [x, y, 1] we obtain the equations of coordinates on the deformed image:

[0027] ;

[0028] , Where:

[0029] - , - coordinates on the deformed image;

[0030] - x, y - coordinates on the reference image;

[0031] - u, v - translational movements;

[0032] - Δх, Δу - distances from the center of the subregion.

[0033] Calculate the gradient of the obtained displacement field to identify areas with sharp changes in displacement corresponding to the boundaries of the object:

[0034] ;

[0035] ;

[0036] Calculate the gradient value using the formula:

[0037] ;

[0038] Implement thresholding of the gradient value to identify potential object boundaries. The threshold is determined based on the statistical characteristics of the gradient field using the Otsu method:

[0039] , Where:

[0040] - σ 2 β (t) is the interclass variance for the threshold t;

[0041] - t - threshold value from 0 to the maximum gradient value;

[0042] - t* is the optimal threshold value found using the maximum interclass variance criterion.

[0043] Calculate the interclass variance:

[0044] , Where:

[0045] - ω1(t), ω2(t) - the proportions of pixels in classes below and above the threshold;

[0046] - μ1(t), μ2(t) - average gradient values ​​in each class.

[0047] Perform morphological filtering on the resulting binary mask using a circular structuring element with a radius of 2 pixels to remove noise and fill gaps in the contours. Morphological closing and opening operations are implemented.

[0048] Perform temporal averaging of results over multiple oscillation cycles in different directions to achieve a boundary detection accuracy of less than 0.1 pixel using the formula:

[0049] , Where:

[0050] - N - the total number of measurements at different directions and frequencies of oscillations;

[0051] - - position of the border in the i-th dimension;

[0052] - - the final average position of the boundary.

[0053] Extract the final contours of the object boundaries with an accuracy of less than 0.1 pixel.

[0054] The proposed method was experimentally tested at the St. Petersburg Moscow Sorting Depot of NVK LLC. Type 18-100 springs removed from the bogie of a freight car undergoing repairs were used as objects for determining the boundaries.

[0055] Each object was mounted against a contrasting background to ensure a clear distinction between the object and its surroundings. A uniform, matte white background with no pattern served as the background, providing maximum contrast against the spring surface. To generate controlled vibrations, a commercially available vibration platform—an ES-3 electrodynamic shaker (ASLi, China) with a 220 V power supply—was used. It is designed for a nominal sinusoidal excitation force of 3000 N, has a frequency range of 3-3500 Hz, a maximum two-way table travel of 25 mm, and a test specimen weight of up to 100 kg. The shaker has a rigid base and a movable table, moved by an electromagnetic drive and controlled by a digital controller, which ensures the generation of harmonic vibrations with a specified amplitude and frequency.During the experiments, the setup was operated in low-frequency vibration mode with an amplitude of 0.3-2.0 mm and frequencies of 2-15 Hz, significantly lower than the maximum capabilities of the setup, ensuring linear operation without overloading the drive. The platform is equipped with an object clamping system, designed as a universal mounting table with a set of adjustable clamps and clamping elements that are installed in threaded holes in the base and allow for rigidly securing objects of various shapes and sizes without shifting during the experiment. The clamps have soft contact pads to prevent damage to the object's surface and are equipped with screw-type locking mechanisms, ensuring reproducible sample positioning during repeated measurements.

[0056] Video was recorded using an ELP 1080P 120fps USB Camera with a resolution of 1920×1080 pixels and a frame rate of 120 fps to ensure detailed analysis of object movement. The camera was mounted on a rigid tripod at a fixed distance of 1 m from the object to eliminate unwanted vibrations in the recording system.

[0057] Controlled vibrations of objects were initiated in various directions with an amplitude of 0.3-2.0 mm and frequencies in the range of 2-15 Hz. For each vibration direction, video recording was performed lasting at least 10 complete vibration cycles.

[0058] All elements of the experimental setup, including the vibration platform, precision actuator, and digital video camera, were connected to a personal computer, which served as the central control and recording device. Control of the vibration modes (amplitude, frequency, direction) was performed using specialized software installed on the computer, which generated control signals for the actuator and synchronized these signals with the video recording process. The video data stream from the camera was transmitted in real time via a USB interface to the computer, where a hardware and software system received, buffered, and stored it as a sequence of frames with specified resolution and frequency parameters.Next, the same system implemented algorithmic data processing: preliminary noise filtering and brightness correction were performed, the region of interest (ROI) containing the object was identified, a series of images were generated for subsequent digital correlation, and the numerical parameters of the object's displacements and boundaries were calculated and stored in a format convenient for further use, accompanied by graphical representations of the measurement results. To convert the coordinates of the identified boundaries from pixels to millimeters, a calibration block with a known linear dimension was first placed in the camera's field of view, aligned with the spring being monitored. A personal computer with the installed software automatically determined the number of pixels, N, from the image of the calibration block. эт , corresponding to a segment of length L эт , the scaling factor k=L was calculated эт / N этin mm / pixel, after which, based on the vertical coordinates of the upper and lower selected boundaries of the spring, its height in pixels ΔN was determined and converted into millimeters using the ratio , which ensured the spring's geometric parameters were obtained in the metric system with subpixel accuracy. The resulting spring height was measured in millimeters.

[0059] A total of 100 type 18-100 springs of varying heights and diameters were tested, including both inner and outer springs of various sizes. For each spring, the goal was to identify three main boundaries—two lateral and one central—from which the spring height in millimeters was subsequently determined. When necessary, the outer and inner diameters were additionally calculated, i.e., the set of geometric parameters specified in the design and technical documentation. The identified contours were then used to determine the spring height in millimeters. For this purpose, it is critical to accurately determine the object's boundaries, as contour determination errors directly impact the accuracy of the spring's linear dimensions. Spring boundaries were determined using the proposed method. The measurement results for 20 springs are presented in Table 1.

[0060]

[0061] where:

[0062] - L standard - spring height measured by contact method using a micrometer;

[0063] - L by image - height determined by the proposed method of digital image correlation;

[0064] - ΔL - the difference between the reference and obtained value;

[0065] - ε - relative error modulus, %.

[0066] Analysis of the data presented in the table showed that for 20 springs of type 18-100 (external and internal), the difference between the reference height measured by the contact method and the height obtained from the image lies within the range of -0.11 to +0.14 mm, with the maximum deviation modulus not exceeding 0.14 mm. The average value for the sample is approximately 0.07 mm, and the mean value is close to +0.01 mm, indicating the absence of a significant systematic bias of the method relative to contact measurements.

[0067] The relative error for individual samples ranges from 0.01% to 0.06%, with an average value of approximately 0.3%. Thus, the proposed method ensures consistency of results with reference contact measurements at the fraction of a percent level, with absolute deviations not exceeding tenths of a millimeter, confirming its high accuracy and suitability for spring geometric parameter testing.

[0068] The advantage of the method is the high-precision determination of the object boundary by using controlled vibrations in different directions and applying the method of digital image correlation with subpixel interpolation.

Claims

1. A method for determining the boundaries of an object based on a digital correlation method of image processing, including estimating the gradient field of the original image, processing it and searching for local maxima in the resulting parametric space, characterized in that the object is placed on a contrasting background, controlled oscillations of the object are initiated in various directions with different frequencies, continuous video recording of the oscillating object is performed using a video camera with a frame rate of at least 120 frames per second (FPS), these images are transmitted to a computing device, and image processing is performed using special software: a set of images is obtained at different positions of the object, each captured frame is divided into a grid of overlapping square subregions, for each subregion a normalized cross-correlation is calculated with the corresponding subregion in the next frame, the pixel intensity distributions are compared,determine the coordinates of the maximum of the correlation function for each subregion, calculate subpixel interpolation using the parabolic approximation method in the vicinity of the found maximum, form a two-dimensional field of displacement vectors for the entire image, calculate the gradient of the resulting displacement field to identify areas with sharp changes in displacements, perform threshold processing to the gradient value to highlight potential boundaries of the object, perform morphological filtering of the binary mask to eliminate noise and fill in gaps, and temporarily average the result over multiple cycles of oscillations in different directions, from which the final contours of the object boundaries are obtained.

2. A method for determining the boundaries of an object using the digital image correlation method according to paragraph 1, characterized in that controlled oscillations of the object in various directions are initiated with an amplitude of 0.3-2.0 mm and frequencies in the range of 2-15 Hz, continuous video recording of the oscillating object is performed by dividing each frame into overlapping subregions measuring 32×32 pixels with an overlap of 75%.